Direct Preference Optimization (DPO) is a technique for aligning language models that removes the complexity of reinforcement learning by deriving a loss function directly from the Bradley-Terry model. Instead of using reinforcement learning algorithms like PPO, DPO transforms the constrained optimization problem of maximizing reward while preserving the language model's original behavior into a simple loss function that can be optimized via gradient descent. The key insight is that the optimal policy derived from the RL objective can be plugged into the Bradley-Terry model, causing the intractable terms to cancel out and producing a computationally feasible loss. This loss function compares the log probabilities of the chosen answer versus the rejected answer under both the optimized and reference language models, with a hyperparameter β controlling the trade-off between reward maximization and behavior preservation.
DPO Explained: Direct Preference Optimization Math for LLM Alignment
Added:hello guys welcome back to my Channel today we are going to talk about DPO which stands for direct preference optimization it's a new technique that came out in the middle of last year in 2023 uh to align language models let's review the topics of today I will start with a short introduction to language models as usual so we can review how language models work then we will introduce the topic of AI alignment so what we mean by AI alignment and then we will review reinforcement learning now may be wondering why are we reviewing reinforcement learning if the whole point of TPU is to remove REM reinforcement learning from language models well the reason is that actually even if DPO is does not use reinforcement learning algorithms they are still interconnected especially when we talk about the reward model and the bread literary model so in order to understand the reward model and the bread lary model we need to review reinforcement learning and how the reward model affected the the process in reinforcement learning from Human feedback in in the last part of the video we will see we will derive the DPO loss so we will understand how the where does it come from uh I will also give you uh the show you the code on how to compute the log probabilities so how we actually can use this log uh this this loss in practice now what are the prerequisite for watching this video well for sure that you're familiar with a little bit of probability and statistics not much uh for example conditional probability uh we are familiar with deep learnings so what we mean by gradient descent and loss functions you it's really great if you have watched my previous video on reinforcement learning from Human feedback in which I explain all the aspects of the reward model and the reinforcement learning framework and the p uh but it's not necessary for this video how because I will review most of the part that are needed to understand the DPO but it's really great if you have already watched that video so you can compare the two uh methods and also that you're familiar with the transform model because we will be using it in practice when we want to compute the log probabilities otherwise we don't know how to use the loss of the DPO let's start our journey so what is a language model well a language model is a probabilistic model that assigns probabilities to sequence of words in practice given a prompt for example a language model allow us let me use the laser so given a prompt for example Shanghai is a city in a language model tells us the probability of what is maybe the next token or word now in my videos I always make the simplification that a token is a word and a word is a token this is actually not the case in most language models but it's useful for explanation purposes so what is the probability that the next token is China or the next token is Beijing or the next token is cat or pizza given a particular prompt this is the only thing that a language model does and um the language model gives us this probability now you may be wondering how can we use this language model to generate text well we do it with an iterative uh process so we take a prompt for example where is Shanghai we give it to the language model the language model will give us a list of probabilities over what is the possible next word or token suppose that we choose the token with the most with the highest probability score so suppose it's Shanghai we take this token we select it and we put it back into the prompt and we ask again the language model what is the next token then the language model again will give us a list of probabilities over what is the possible next token we select the one that we think is the most uh relevant usually we select the one with that is most probable uh and we put it back into the um prompt and we ask again the language model etc etc until we reach a spe specified the number of generated tokens or we reach the end of sentence token which is a special token in this case after four tokens generated the language model will probably say Shanghai is in China which is the answer to our question what do we mean by AI alignment now when we train a language model we train it on a massive amount of data for example thousands of books billions of web pages and the entire Wikipedia Etc this gives the language model a vast knowledge um in to to to complete any prompt in a reasonable way however this does not teach the language model to behave in a particular way so for example we can this does not the pre-training does not teach the language model to be polite or to not use any offensive language or to not use any racist Expressions Etc because the language model will just behave based on the data that it has seen and if you feed the internet data the language model will behave very very very badly actually so we need to kind of align the language model to a desired Behavior so we don't want the language model to use any offensive language we don't want it to be racist we want the language model to be helpful to the user to so to answer questions like an assistant etc etc and this is the goal of AI alignment now let's talk about reinforcement learning so reinforcement learning is an area of AI that is concerned with training intelligent agents to perform actions in an environment in order to maximize a reward that they receive from this environment let me show you with a very concrete example I usually always use my cat oo for examples so let's talk about oo oo is the agent in this case in this reinforcement learning scenario and he lives in a very simple word let's call it a greed word that is made of cells in which the cat's position is indicated by two coordinates the X position and the Y position uh this can also be treated as the state of the agent because at every every positions corresponds to a particular state the agent when it is in a particular state it can take some actions in the case of the cat it can go right left up or down for every action that the agent takes it will receive some reward from the environment it will for sure change its state to a new one so for example when the cat moves down it will change to a new state to a new position and it will receive some reward according to a reward model that we specify in my case I have specified the following reward model so when the cat moves to an empty cell it receives a reward of zero If It Moves towards the broom it receives a reward of minus one if somehow it arrives to the btub it will receive a reinforce reward of minus 10 because my cat is very scared of water and if it arrives to the meat which is the cat's dream it will re receive a reward of plus 100 now what dictates what action the agent will take given a part particular state or position well it is the policy the policy indicates what is the probability of the next action among all the actions that are available that the agent can take can take given a particular State and we usually write it like this so that next action at time step T is distributed like the distribution uh induced by the policy according to the state the um the agent is in now what is the goal in reinforcement learning the goal in reinforcement learning is to select a policy or to optimize a policy in order for the agent to maximize the expected return when it acts according to this policy so imagine we have such a policy that uh is optimized while our policy for sure if the goal is to maximize the expected reward when using this policy for sure we will have a policy that will take us on average to the meet because that's one way to maximize the expected reward and for sure it will be a policy that will allow us to minimize the chance of ending up in the water here or to the broom here now you may be wondering okay the cat can be seen as a reinforcement learning agent as a physical agent that takes some action spot what is the connection between reinforcement learning and language models well as we saw before in reinforcement learning we have this thing called policy in which given State the policy tells us what is the the probability over the action space of the next action so what possible next action we can take and the probability of each action this is also something similar to what we do with language models because in language models we also have some kind of state which is our prompt and we ask the language model to give us the probability of the next token or we can consider it the next action that the language model can take and we want to reward this language model for selecting tokens in such a way that that they end up generating good responses and we don't want to reward the language model for selecting sequence of tokens that end up giving us bad responses now imagine we are trying to uh um train a language model that needs to act like an AI assistant so for sure we want the language model to be helpful to answer questions in a meaningful way so not just output garbage we want the language model to not be racist or not use any offensive language right so this is all good behaviors that we want from this language model so we may want to build a reward model that will treat good responses for example responses that actually answer the question asked by the user and we we will reward them with a high reward and we give a maybe zero reward or negative reward to all those answers that are not coherent with what we want so for example if the language uh model generates um uh dirty jokes or um racist jokes for example we want we can give zero reward to those um to those responses so the language model acts also as a policy because the policy is something that given a prompt tells you what is the pro probability over the action space or in this case the probability over the token space so we want to optimize this policy so we want to optimize the language model to maximize the probability to maximize the expected return or the expected reward that it receives receives from our reward model so we want to optimize our language model to generate good response because that's one way to obtain uh High reward from the reward model now you may be wondering okay but how to define the reward model for a language model well one way would be okay we can have a list of questions and answers generated by the language model and then we can give give a numeric reward to each one of them and then we can use some reinforcement learning algorithm to uh feed this reward model to the language model to optimize it but the problem is what kind of reward can we give to each of these pairs of questions and answer because for example let's look at the first question where is Shanghai the answer suppose it is generated by the language model is Shanghai is a city in China now in in my opinion this is a good response because it's short and up to the point but some other people maybe think that only the word China is enough because this the user just asked where is Shanghai so there is no need to repeat the word Shanghai but someone else maybe think that this response is too short and the assistant should say hello I think that the your qu the answer to your question is Shanghai is a c in China so different people will have different opinions on what reward to assign to this particular pair of question and answer because we humans are not very good at finding a common ground for agreement but unfortunately we are very good at comparing and we will exploit this fact so instead of building a data set set that is made of questions and answers and the rewards because we do not know what kind of reward to assign we will build a data set of questions and multiple answers and then we ask people to choose an answer that they that they like According to some preference that we have so we want generate for sure a language model that is helpful so we want the gener language model to give responses that are correct and we want the language model to be polite for example so for example imagine we have a list of questions and then we ask the language model by by using for example a high temperature to generate multiple answers and then we ask people to choose which one they like in this case for example where is shanai for sure the people most people will choose the answer number one because Shanghai is a city in China is the correct one uh for this question here for example also people will probably choose the this question here even if it's very short because the other one is probably wrong so using a data set like this we can actually train a model to trans transform a pair of question and answer into a numeric reward let's see how it is done if you have a pet you probably know that to teach a particular Behavior to your cat or to your dog you need to use biscuits or some treats so you ask the cat to do something and if the cat does it then you give it a treat so it will reinforce this memory in your cat and then the next time the cat is more likely to do it because it will remember that it received some treat and so it will again perform that action again so it can probably receive another treat this is exactly what we do in reinforcement learning we want to give some digital biscuits to our reinforcement learning agent so that it uh is it is more likely to perform that action or that series of actions again in order to receive more reward however the data set that we have built so far is made up made up of preferences so we have a question multiple answers and then we ask people to choose which answer they like we need to convert this data set of preferences into a numeric score that we can give as a reward to our language model to choose more likely the answer that was chosen by the people and to make it less likely to choose the answer that was not liked by the People by our annotators and this is can be done through a preference model in DPO and also in reinforcement learning from Human feedback we make use of the breadly model so the bread L model is a way of converting a data set of preferences into a numeric score called reward that is given for each pair of questions and answers our goal is to train a model that given a question and answer or a prompt and the text generated text to give a score that resembles the preferences that have been chosen by our not ators um this is the expression of the bread lary model so it is a model meaning that we choose to model our preferences like this and actually it makes sense because it is a probability so that's why for example we use exponentials because we want the probabil the probability has to be uh non- negative and also the probability that the assigned to the correct preference so the probability of choosing the correct answer of the wrong answer so the one that is chosen by the annotators over the one that was not chosen by the annotators here I call it winner and loser because also in the DPO paper they call it winner and loser it is modeled like this so it is proportional to the reward to the exponential of the reward that was assigned to the winning answer now how to train a model to convert a data set of preferences into a numeric reward we take this expression and we can use a maximum likelihood estimation now it doesn't matter if you don't know what is maximum likelihood estimation the point is we want to maximize the probability of assigning the correct ordering in our preferences so we want to maximize the probability of choosing the correct uh answer over the wrong answer and suppose that we are maximizing this expression here let's see how we can derive the loss to maximize this expression here if you look at the DPO paper you will see that they go from the Brad lary model which is this one directly to the loss here but they don't show you the derivation so I will I will show you how to derive the loss that maximizes this uh probability here uh the uh derivation is very simple actually so uh first of all as you can see in the loss you can see this function here it's a sigmoid function the expression of the sigmoid function is this one and this is the graph of the sigmoid so the expression of the sigmoid function is 1 / 1 + e to the^ of minus x uh the first step of the derivation is to prove that um two uh exponentials so a fraction of the this expression here so exponential divide divided by the sum of two exponentials can be written as a sigmoid of a minus B so here I call all this part here so let me use the pen I think it's easier so uh this part here so the reward assigned to the let's say the winning answer is we call it a and the reward assigned to the losing answer we call it B so this uh expression can be written as e to the power of a / e to the power of a plus e to the power of B and we will prove that it can be written as the sigmoid of a minus B through the following through the following step so first we can divide uh we take this expression which is basically this one we just replace the rewards with A and B because it makes it simpler to visualize uh we divide the numerator and denominator by the same quantity e to the power of a we can do it uh then uh we can um uh on at the numerator e to the power of a cancels out with e to the power over and becomes a one then in the denominator we add and subtract one we we we can do it because it's like adding zero and then we collect the minus one so we don't change anything we just put the parentheses this is possible through the associative property uh then we do the uh common denominator for these two expressions for these two expressions and we arrive to this one we can simplify e to the power of a with minus E to the^ of a so it becomes e power of bide by e to the power of a which thanks to the property of the exponentials can be written as e to ^ of B minus a then we can take a minus sign outside and this expression here is exactly the expression of the sigmoid function you can see here so it's 1 over 1 + e to the^ of minus something so it becomes the sigmoid of that something here A minus B and this is exactly the loss that you see here so it is the sigmoid of the reward assigned to the winning answer minus the reward assigned to the losing answer here we also see a log because usually we do not log model the probability directly but we model the log probability so we have also this log because we want to model the log probabilities it is something that we can do because it's the logarithm is a monotonic function and also you may be wondering why do we have this minus sign here uh this is because we want to maximize this expression but as you know in deep learning uh Frameworks like pytorch we have an Optimizer that is always minimizing a loss so instead of maximizing something we can minimize the negative expression of the objective function which is the same thing so basically we take this loss function and if we apply it to a reward model which is a neural network it will be trained to uh maximize the probability of giving the correct ordering to our preferences which can only happen when it assigns a high reward to the winning answer and a low reward to the losing answer because if you look at this expression here as you can see the probability is maximized when the in the numerator we have the uh reward assigned to the winning answer so the reward assigned to the winning answer is higher than the one assigned to the losing answer and um if you are wondering how to read an expression like this so let me cancel because we will use it a lot this kind of convention uh this one this basically means that we have a data set of preferences uh where we have a prompt a winning answer and a losing answer and they belong to our data set of preferences and we train a module with the gradient descent for each of these preferences we calculate this loss here this expression here and if we minimize this loss with the gradient descent we will have a neural network that is trained for the following the the bread literary model basically okay now that we have built a reward model which means that we have a model that given a question and answer can assign a numeric reward to the language model if the response is correct or looks good according to the behavior that we want from our language model or looks bad according to the behavior that we want from our language model now we can um we can train our language model so what as you recall what is the goal in reinforcement learning in reinforcement learning the goal is to optimize a language model which is also the policy of our reinforcement learning agent in order to maximize the cumulative reward when the agent acts according to this policy in other words if we let me use the pen so let's ignore for now this green part here Suppose there is no green part here so this doesn't exist imagine we have a language model let's call it Pi Theta because it's a policy and we want to uh optimize this policy so we want to optimize this language model in order to maximize the reward that it receives from the reward model it means that the language model will generate answers that give good reward and how they get good reward if the answers are looks good they for example are not racist they are not using any sexual jokes and they are actually answering the question that was asked um however and this is the goal in reinforcement learning from Human feedback for example uh that's why there is it's called a reinforcement Le from Human feedback now uh if we use a model if we use an objective um like this that is we only want to maximize the reward then the language model may become greedy and just output garbage that gives it good reward so imagine we have a reward model that rewards the language model for being polite the language model may just start saying a list of thank you thank you thank you or please please please and a lot of please or a lot of thank yous just to get high reward because probably the word the word thank you and please are highly rewarded by the reward model but we don't want the language model to just output garbage to get reward we want the language model to also um uh output something that was according to its training data so so it's a pre-training but we want to change it a little bit so that it also acts according to our reward model so to our data set of preferences so it is more polite but without forgetting what it has learned from the pre-training and this is why we add this uh KL Divergence in the objective so let me use the pen again so we we change the objective a little bit so we want the language model to maximize the reward it gets from the reward model but at the same time we add a constraint to the language model through a KL Divergence now the K Divergence can be thought of as a distance metric it is not a distance metric but can be thought of as a distance metric between two um distributions um in which we have a pre-train model so a language model that was not fine-tuned through reinforcement learning from Human feedback or DPO so it's just the language model that has been pre-trained on the Wikipedia on the books and on the internet web pages and then we have the language model that we are optimizing so this Pi Theta and we want them to be very similar so we want the language model to not change much compared to what it was before the reinforcement learning from Human feedback or before the DPO training tring and this is why we add this KL Divergence so we want the language model to maximize its reward but at the same time not forget or not change too much its output in getting this reward now what that now that we know the reinforcement learning objective uh which is basically also the same objective that we have in DPO because also in DPO we want to train a language model that maximizes a reward but at the same time does not for get its training data let's look at what does it mean to actually maximize an objective function because this is an objective function that we have and we want to maximize it but what does it mean to maximize an objective let's see um maximizing a function means to find the values of some variable such that the value of the function is maximized for example if I give you the following function f ofx is equal to Min - x - 3 + 4 whose graph is very simple it's just a parabola facing down uh to maximize these functions means to find the value of the X variable such that the function the y basically the Y of this function is maximized how to do that analytically well we calculate the derivative of this function here we set the derivative equal to zero and we find the values of X for which this derivative is Z and that is also the value for which the function will be maximized so the derivative of this simple function is - 2x + 6 and the value of x that makes this derivative zero is the value 3 x equal to 3 which is also the value of the uh as you can see in the graph that maximizes the function now the the objective function that we saw before so this one so in which we want to maximize a reward but at the same time we want it we want the the language model to not be too much different from the uh unaligned language model so the the language model that is not aligned through reinforcement learning from Human feedback or DPO it is called a a constrained optimization problem because we want to maximize the reward but at the same time we want to put some constraint on this um objective function we don't want the uh K Divergence to be uh too big we want it to be constrained in some limit now the the point is okay there are many techniques for constraint optimization and we will not be see them because there are univers entire phds on optimization but one thing you may notice is that okay this one here looks like the objective function looks like a loss function so why cannot we just use for example gradient decent to um optimize this objective function here such that we can train our language model to behave in a particular way to maximize this reward well we could but as you know in deep learning and especially with back propagation we need an objective function or a loss function that is differentiable the following this uh objective function is not differentiable why because as you can see from the expression here this is an estimation of over the all the prompts in our data set and then a output that is generated by the language model now to generate the output of the language model as we saw before we need to use an iterative process in which we feed one token at a time into the prompt uh we add uh we sample one token at a time from the language model we take this token and we put it back into the prompt feed it again to the language model Etc and we use many strategies for selecting the next token sometimes we use the gitty strategy sometimes you use the beam search sometimes you use the top case top P etc etc now this is sampling operation that we do on the language model to sample the answer of the language model is not differentiable that's why we cannot run reinforcement learning to maximize this objective or to minimize the negative objective in case we treat it as a loss and that's why in reinforcement learning we were forced to use algorithms like Po now let's see how DPO handles this in the DPO paper uh they start with a very simple uh introduction to the reinforcement learning objective so as we saw before the reinforcement learning objective is to select a policy so a policy that um maximizes the expected reward when using this policy so the polic is the language model and at the same time puts a constraint on how much this policy can change during this training this optimization and in the D paper they say okay there is an exact solution to this optimization problem and it is the following it is the equation for in the DPO paper and exact solution me I mean that there is an analytical solution to the this uh constrained optimization problem just like we had a analytical solution for the maximization problem of this Parabola so we could find through the derivative and setting the derivative equal to zero we could find the value of x such that this function here is maximized and for the using the same reasoning but but different technique um there we can we also have a analytical solution for the constrainted optimization problem that we saw before and this is the solution now you may be wondering okay great we have an exact solution just like the parabola so now we are all set right yes the problem is we we have an an exact solution but it's not easily computable so it's not easy to compute so mathematically it it is it exists it makes sense but it's not easy to compute why because we have this Z of X term here now this Z of X term here if you look at how it's defined it's the summation of all possible y's um that are generated by the reference model so the as you know when we do reinforcement learning from Human feedback or DPO we have two models one the one one is the language model that we are trying to optimize and one is the Frozen model that we don't optimize but we use it as a reference for the K Divergence so this is called the p ref so all the outputs generated by P ref multiplied by the exponential of the reward now the problem is this summation is done over all possible WIS it means that we need to sample all possible outputs through through um from our language model given all the prompts that we have in our data set of preferences now to generate all possible outputs is very very very expensive imagine you need to generate your language model can generate 2,000 tokens for each prompt it means that and you have a vocabulary size of 30,000 it mean that for the first position you have 30,000 possibilities for the second position you have 30,000 possibilities for the third position you have 30,000 possibilities and then you multiply all these possibilities so it becomes a lot a lot a lot of uh outputs that you need to generate to evaluate this Z of X term so the analytical solution to the constraint optimization problem that we so before exists but it's not easy to compute however one thing is interesting from this expression imagine that somehow magically we have um access to an optimal policy so this um solution to the optimization problem allow us to compute what is the optimal policy given the optimal reward model and the reference policy so the reference language model but imagine that for some reason some magically we have access to uh we have this term here so if we have this term here we can compute the optimal reward model with respect to the optimal policy how well we can just isolate this R of X and Y term from this expression here and it's very easy to compute because we can apply the logarithm on the left and the right side of this expression so let's do it step by step we can apply the the log on the left side and on the right side of this expression here so this expression here and we will get that the uh the log of a product as you know is the sum of the logs and the log of the ratio is the difference of the logs so this Z term is in the denominator so it becomes a minus log of Z of X this one is in the numerator and this one is in the numerator so they become sums of logs so this one plus this log here then the log and exponential can cancel out because they're inverse functions so this allow us to isolate this R of XY term with respect to all the other terms and we can write it like this so we can calculate R of X and Y with respect to an optimal policy that we think we have access to we do not have access to it but we pretend we have access to it okay so there are two things that we do not have in this expression we do not have the reward model the optimal reward model and we do not not have the optimal policy but we pretend that we have the optimal policy why let's see the next step that they do in the DPO paper the next step is they say okay do you remember the bread literary model as you remember the bread lary model is a reward model right is the model that given a data set of preferences allow us to compute a numeric score a numeric reward well this um bread literary model receives uh is based on a reward that we assign right so what if we plug the reward that we have uh computed uh from the constraint optimization problem into the bread literary model well we can do it so we have this reward that we obtain from the constraint optimization Problem by inverting the formula and we plug it inside the bread lary model so if you remember the bread literary model can also be written as a sigmoid and and we prove it before in the previous slide so what we do is okay the bread lary model can be written as a difference of rewards in the sigmoid function so if we plug the reward here so the reward uh obtained by the constraint optimization uh problem solution we will see that the two Z of X terms because this is a difference of rewards as you can see if we plug here for the reward assigned to the winning um response and here the reward assigned to the losing response we have these two Z of X terms so plus beta log of Z of X and minus beta log of Z ofx that will cancel out because they are one the opposite of the other this way we can obtain a formula that does not contain the Z of X term and it's now computable so basically if we use the loss of the pr lary model so as you remember the the pr literary model is a model that allow us to to train a language model to uh model the reward right and um if we use the loss of the bread literary model in which the reward is coming with respect to the um to the optimal uh policy we can use it to optimiz the policy to adhere implicitly to the reward model according to the bread ly model and this is the whole idea of the DPO paper so we can plug the exact solution of the constraint optimization of the reinforcement learning objective we can invert it to get the reward we plug it into the Brad literary model because the Brad literary model only depends on the difference of rewards assigned to the losing uh to the winning answer and to the losing answer the uncomputable term Z of X cancels out and then it becomes computable and now we can use it to train a language model that will act according to the reward model of the Brad literary model so to the U preference model um given by the bread literary model so it will favor good responses and at the same time it will um uh you will be less likely to output the preferences that were not chosen and at the same time it will put a constraint onto the K Divergence so at the same time it will put a restriction on how much the language model can change with respect to the uh reference model so the language model that was not optimized with reinforcement learning from Human feedback or DPO so basically with DPO we are doing kind of the same thing that we are doing in reinforcement learning but without using the reinforcement learning um algorithms so the goal in both of them is the same so we want to optimize a policy we want to optimize a language model to maximize a cumulative reward but at the same time we want to put a constraint on how much it can change using the K Divergence in the case of reinforcement learning from Human feedback we are using the pop algorithm to optimize this um objective to optimize this policy but in the case of DPO we do not have to use reinforc learning from Human feedback because we found a loss that implicitly is already um mapping this reward fun with this reward objective through this loss um let's see how to actually now uh compute the um the log probabilities so how to actually use this loss because first of all let's look at the expression of this loss this law says that if you have a data set of preferences in which X is the prompt the Y YW is the chosen answer and Y L is the not chosen answer because as you remember this data set is made up of preferences of question and two answers and then we asked some annotators to tell tell us which answer they prefer so if we have this data set we can run gradient descent using this data set over this loss here this loss here now to calculate this loss okay the logarithm we can always calculate it's just a function the sigmoid is a function we can calculate but the logarithm and the beta are the beta is a hyperparameter that indicates how much we want the language model to change with respect to the reference language model or how much we want to constraint it and then we have to compute these log probabilities so the log of the probability of generating this uh y w when the language moduel is prompted with the prompt X and also for the P ref so also for the language model that is not being optimized by DPO let's see how to practically uh compute this log probabilities so um when you run DPO it's very simple so imagine for example you are uh using a hugging phase it's just a matter of using this classes or DPO trainer in which you pass the language model that you're optimizing the Frozen version of the language model that you don't want to optimize but it's is the reference language model um that is used to compute the log probabilities to calculate the K Divergence then you can give some other training arguments you can have the list of them on the website of hugging phas and then this beta parameter which indicates the strength on how much you want the language model to change and also in the website of DPO of hugging pH they also give you what is the typical range for this um for this uh hyper parameter now what will happen inside the Library of the DPO trainer so inside the hugging face Library when you use the DPO trainer they compute this loss so they calculate this log probabilities you can see here so for example this log probability you can see here but how do they actually compute well as you know a language model usually most very most um in most cases it is a Transformer model and to compute these log probabilities as you know they use a prompt a question the answer that is chosen and the answer that is um not chosen so the winning and the losing answer suppose that we want to generate the log probabilities for the winning answer so this one so we have a language model we give it a prompt and the answer that was generated and we want to calculate this log probabilities what we can do is we can combine the question and the answer in the same uh string in the same uh input for the language model so imagine the question is where is Shanghai oops imagine the question is where is Shanghai question mark and the answer is Shanghai is in China now let me use the laser okay we can feed all of this to our language model so the pi Theta the language model is a Transformer model most of the cases and it will generate as you know the Transformer model generates some hidden States so it takes some input which which are uh embeddings and it outputs some embeddings that are contextualized also according to the self attention mask now if you don't know how this works I highly recommend you watch my previous video on the transformer in which I show the self attention mechanism but basically the Transformer model is a model that takes some embeddings and through the sortation mechanism output embeddings then we take these embeddings and we can project them into Logics using a linear layer and we can do that for all the tokens that we give to the input usually we only apply the linear layer to the last token when generating the tokens because we are interested in generating the next token but we we but we can do it for all the hidden States we are not forced to only use the last one because the each hidden State encapsulates information about itself and all the tokens that come before it so we can take this logits and then we can also convert them into probabilities but we do not want probabilities we want log probabilities because as you can see here we have this log function here so instead of applying the soft Max we can apply the log soft Max to each of these logits now when we apply the soft Max it will become a distribution over the entire vocabulary for one for each uh token in the vocabulary but we want only the probability corresponding to the Token that was actually chosen to generate this particular answer and we also know which token it was because we have the answer so to the question where is shanai we know what is the answer because it's in our data set of preferences so we know that the answer is Shanghai is in China so how to compute these log probabilities so we can compute the log probabilities over the entire dictionary over the entire vocabulary and then we only select the log probability corresponding to the Token that was actually selected in the answer so for this uh question for example we select the for example the last hidden State for the question which correspond to what should be the next token and we know what is the next token the next token is Shanghai so we take the Lo probability only corresponding to Shanghai for this prompt here so where is Shanghai question mark Shanghai we know that the next token should be is because it is already present so we take the lock probability only corresponding to the Token is etc etc and we do it for all the token that are in the answer and this gives us all the log probabilities of the tokens of the answer for this given question and then we can sum them up why we need to sum them up because it's a log probabilities usually if they are probabilities we multiply them but because they are log probabilities uh we sum them up because the the logarithm transforms transforms uh products into submissions and this is exactly what happens inside the hugging face Library so inside the hugging face library to compute the log probabilities to calculate this loss they actually do what I described so they take the logits and they use the labels what are the labels just the um the tokens corresponding to the answer uh for example to the winning answer or to the losing answer depending on which term you are Computing this one or this one and the model that we are using so the reference model or the model that we are trying to optimize and then they select here in this line here they check the the log probabilities only corresponding to the uh labels so to the next token that we that we already know what is it and then they sum them up as you can see here and here they also apply a loss because we know we don't want all the log probabilities but only the one corresponding to the tokens that are belonging to the answer not to the one that belong to the question and this is how DPO works thank you guys for watching my video I hope you learned a lot I tried to simplify as much as possible the math of DPO but the basic idea is that we want to remove reinforcement learning uh to align language models this makes the tract much simpler because it just becomes a simple loss in which you can run a gradient descent and you don't have to worry about training a separate reward model which is something that we did in reinforcement learning from Human feedback so if you watch my previous video as you remember the um the math is much more hard and much more topics to introduce on how to optimize the objective that we saw and um please come back to my channel for more videos like this I usually try to make videos that are very deep very um uh in depth for every topic sometimes they can be a little hard but I try to simplify as much as possible also depending on my knowledge and also depending on how much it is possible to simplify a difficult topic and and um if you have any questions please leave it in the comment and I will probably um keep publishing uh more videos like this but if you want videos that are more simpler please let me know and also let me know in the comments what kind of topics you would like me to explore for next thank you guys and have a nice day
Up Next

LLM Preference Tuning Explained: RLHF, PPO, DPO | Stanford CME295
@stanfordonline
22.2K views•2025-11-14

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

RAG Explained: Embeddings, Sentence BERT, and HNSW Vector Databases
@umarjamilai
84.9K views•2023-11-27

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence























![[인공지능,머신러닝,딥러닝] (심화) Direct preference optimization (DPO)](https://i.ytimg.com/vi/A80ue5nS_A4/maxresdefault.jpg)

















