This lecture covers preference tuning and reinforcement learning from human feedback (RLHF) for aligning large language models with human preferences. The process involves collecting pairwise preference data where humans rate which model response is better, then training a reward model using the Bradley-Terry formulation to score responses. The second stage uses reinforcement learning algorithms like Proximal Policy Optimization (PPO) with clipping and KL-divergence penalties to align the model with these preferences without deviating too far from the base model. An alternative approach called Direct Preference Optimization (DPO) provides a simpler supervised method that directly optimizes model weights using preference pairs, avoiding the complexity of maintaining separate reward and value functions.
LLM Preference Tuning Explained: RLHF, PPO, DPO | Stanford CME295
Added:Cool.
Hello everyone and welcome to lecture 5 of CME 295.
So first of all, thank you all for taking the time to take the midterm last week. So I hope it was reasonable for you. Um so for those of you who are auditing in case you are interested in taking the exam just know that the exam and the solutions are both posted on the websites and um so now you know a little bit how the exam looks like. So for the final we'll have the same format except that the content will be on lectures five. So this one up until 9.
So I'll go down that point.
Cool. Uh great. So with that we're going to start the lecture. So today we're going to talk about LLM tuning.
So as usual we're just going to recap what we saw last time. Um, so last time was already two weeks ago, but we talked about how to train an LLM. And in in particular, we've looked at two important steps. So the first one was called pre-training where you're basically taking a model that has been initialized and you're trying to teach the model about language, about code. So you have this step that is very time consuming, expensive um compute heavy that is happening on a lot of data. So we've seen training optimizations on how to make that happen. So we've seen uh techniques to parallelize that across GPUs. So we've seen data parallelism methods and in particular zero so the variant 0 1 2 3 and we've also seen very quickly what model parallelism was in this case. So at the end of this step what you obtain is a model that knows about the structure of language about codes basically all the text that it has been fed. But what this model can do is only predict the next token.
So it's a great autocompleter, but it's not a helpful model yet, which is why we have a second step. And we saw uh here it's typically called fine-tuning or SFT for supervised fine-tuning.
So here what we do is we take our pre-trained model and we train it for specific tasks.
So nowaday you have you know chat GPT and all these like uh chat assistants.
So this can be one application. So transforming your model into an assistant. And so here the goal is to teach the model how to behave.
So the model already knows what language is, what code is, etc. And you're just trying to make it behave like the use case that you're trying to tune it for.
So typically here what you have is a data set that is much smaller in scale but of much higher quality and you're basically tra uh taking your model your pre-trained model and teaching it exactly which tokens to predict with the next token prediction task.
And uh we've seen uh Laura which is parameter efficient method which does not tune all the weights but in a clever way introduces low rank matrices that are the ones that are being tuned and we had stopped there. So what we're going to see today is how to align the model to align with what we call human preferences.
And so here we're taking our model that has been fine-tuned for a specific task and trying to align our model to um make it more something that a human would like or that some metric that we're defining would be more aligned with that.
So as an example, let's suppose if you have an assistant at the end of step step two. So it's very much possible that your assistant is you know behaving the way you want but not let's say at the tone that you want. So it's not let's say friendly or it's not uh you know safe. So you want to tune those aspects in that third step.
Cool. So this step is called preference tuning and we're going to exactly see what that is. So here for context let's suppose that we have an SFT model and so by SFT model I mean a model that has gone through the pre-train stage and the finetuning stage. And so for instance we may ask our model uh to suggest a new activity we could do with our teddy bear. And here the model let's say responds with I would suggest you to not spend much time with your teddy bear at all. So this is the response of an assistant but it's not necessarily with aligned with what we want.
So the idea here is to take these quote unquote bad outputs and find an output or rewrite an output that we would want to have instead.
So this pair would be what what we call a preference pair.
So in other words given this prompt we have two responses.
One that we want to see which is this one I'm going to read in a second and then the other one that we do not want to see.
So the for instance in this example the answer we want to see is you know of course teddy bears not only make awesome companions for delightful sleep but also can also be great buddies for fun activities and you know just like suggest some activities.
Does the setup sound good?
Yeah. So long story short we want to align our model with human preferences.
So you may ask, you know, we have this fine-tuning step already. Why would we want to have a third step? Well, during the second step, which is the fine-tuning stage, if you remember, what we did was construct a very high quality data set of the kinds of prompts on which we want our model to behave in a certain way. So in order to compose such a data set, it's actually something that is very time consuming and very actually difficult because the data set must be of very high quality. And in this case, what we're doing is not really teaching the model exactly what it should generate, but rather telling the model what kind of output it should prefer.
So we're less in a you know please generate that kind of thing and more in a you know I prefer this option kind of thing.
And so typically, I don't know, if we asked you uh write a poem, write a great poem from scratch, it would typically be much more difficult to do rather than just showing you two poems, one bad poem and one great poem, and ask you to just say which one is better.
So in order to obtain the data sets, it's already much easier.
The second reason is so during the SFT stage when we compose our data set of very high quality there is one aspect that we really try to get right which is the distribution of prompts and what I mean by that is if we have too much of a given kind of prompt our model will be more biased towards responding in that particular way. So what people try to do is to be careful about the distribution of prompts that are in the SFT data.
And so here, let's say if our model misbehaves, if we were thinking about just adding one example in the SFT data set, well, we would have to be very careful about which prompt we're adding and whether it's not going to bias the model like too much in that direction.
So that's the second reason. And then uh the third reason is what I mentioned which is uh the SFT data is typically very high quality. So if you're um looking at all the missteps that your model is doing and trying to put that in the SFT data, you're just, you know, have a hard time takes a lot of time. Um but one note here is if your SFT data is misbehaving a lot, it may also be due to the fact that your SFT data set has some problem.
So preference tuning is not the answer to everything. Maybe it's better sometimes to just check your SFT data set for some issues.
That sound good? Yep.
Yeah. So to do that in the preference tuning stage. Okay. So the question is we saw Laura for SFT. What is the equivalent for preference tuning? uh so we'll see that later in the in the lecture but you can think of Laura as being some way to I guess reduce the number of parameters that you need to tune that is slightly different with the objective function that you're using to train your model and here preference tuning you can think of it more as a different objective function but it's it may very well be something that you also use LoRa for. So the two are not incompatible.
>> Yes, >> but it will become more clear later on.
But yeah, great question. Uh one last thing I will add here is um another difference with the SFT stage is that preference tuning allows us to inject some negative signal because SFT is all about teaching the model about like what it should predict but it does not teach the model about what it should not predict and we will see that perference tuning allows you to inject some negative signal Cool. So to start, of course, we need to have our preference pairs. And so we're going to look at this uh data collection step first.
So here's the setup.
You have a prompt. Let's say write a poem and you have a given response which is the poem that is being generated by the model.
You have a few ways to construct your preference data.
So either you start with kind of a p pointwise mindset where you score each proposed poem with some kind of pointwise score. And here pointwise score means a score just relative to one observation.
You could very well do that, but I will say it's kind of hard.
It's kind of tough as a human to say, okay, this one is, I don't know, 0.9, this one is like a 0 2. It's not super clear uh exactly how you would scale that.
The second idea is for you to get two observations at the time and for you to say which one is better. So this one is called pair-wise preference data and it's much easier.
And then the third one is listwise which is you know you get a list of let's say n poems and then you just rank you know which one is best which one is uh worse etc etc and I guess this one is easier than pointwise because you don't have to specify I guess how much better it is but I guess it's still a bit more complicated and that's the reason why people typically use parise So what they do is they collect pair-wise preference data meaning for each prompt they have two possible answers and then they just specify which one is better and so that's the one we'll continue um this lecture with. Okay so now you may ask okay great uh pair wise preference data but how do you get it?
Well, here is the recipe. So, in order to generate a pair of responses, so first you need a prompt and we've seen um you know, previously I think it's lecture probably three um that what you could do was to generate different answers if you have a temperature that's positive. So typically what people do is they put this prompt let's say twice into the model with a positive temperature and then they can get two different answers.
Um the prompt is typically something where we wanted to follow the distribution of what the users typically ask. So the prompt X can be something that we obtain from the logs or from let's say a desired set of prompts.
And then what we do is we have these two observations. So the first one is the prompt and the first response. The second one is the prompt and the second response. And what we do is we rate we compare them.
So we can compare them with of course human ratings but we can also compare them with some other metrics.
So I'll just list a few. Uh so we have LLM as a judge which you may have heard of which we have not seen yet which we will see in a few lectures. That is also typically used to just compare I guess how much better one observation is compared to another.
We can also use some other metrics rule-based one like blur rouge etc. Although it is not as used these days and the simplest way to compare these two observations is to have a binary setting where you say okay is response one better or worse than response two. But you could also think of a more nuanced scale. Meaning you can also say okay response one is much better, better, slightly better, slightly worse, worse or much worse. So this is also something that you can do. Um but there are some challenges with that approach because for instance if you uh I don't know consider human ratings a lot of tasks are a little bit subjective.
So in a lot of cases actually what people do is having a pair wise preference data set on the binary scale.
So only is it better or worse.
Does that sound good?
Yeah. Um okay. So another way to obtain that data is to have to find in your logs a response that you did not like to take that response and to rewrite it which is basically what we did here. So here when we have the response here what we did is take the response and rewrite rewrite a good one.
So this is also what people do but of course it is a bit more involved because you need to you know generate and I told you that generation was uh kind of costly and tough but this is also possible.
Does the data collection make sense?
Yeah. Cool. Okay. So now we have our preference data and what we want is to align our model to prefer responses that were preferred by the rating and I guess down downweight the responses that were not preferred.
And so in order to do that we will see a method that's called RLHF RLHF.
And uh I'll see we'll see that uh in more details in a second. Uh but as the name indicates RLHF relies on RL.
So I'm just going to start with some RL basics. So do we have any RL experts here?
Yeah. No. Okay. So no need to be an RL expert. So don't worry, we'll go really slowly on that.
So in the RL world world what people do is they have an agent and here I mean agent in the RL term which interacts with an environment.
So what does it do? It is at a given state at let's say time t. It can take an action at time t and it takes that action according to some policy.
The policy is typically noted pi pi of theta of the action given the state. So what that policy means is simply giving you the probability of you taking an action given a state.
Easy, right? And so given that the agent takes some action, it also receives some rewards.
So sometimes it's a good reward, sometimes it's bad rewards.
So what we're going to do is to leverage that mindset for our preference tuning exercise.
So let's see together how we can transpose these quantities in the LLM world.
So what is the agent?
The agent is the LLM.
So in terms of uh the state that it is in so it's simply the input that it has so far.
So the action that it wants to do is to predict the next token. So it basically always wonders okay what is the next token given this input and this action or this next token is among the set of tokens that there are out there. So if you want you know the environment can be the set of tokens of your vocabulary and in order to decide which token should come next.
The LLM determines that using the probability of next token which is obtained by you know when you do your forward pass and when you look at uh the probability distribution as an output which is basically our policy.
So our policy here is simply equal to the output of the LLM given the input in order to determine the next token.
So far so good. And now we're adding one additional thing. So I told you, you know, we're composing our preference data sets and we want to know which output is better than the other output.
And so we're going to use that for a reward. We're going to somehow use that for a reward and we will see how we will use that.
So I'm going to just recap that part. So we have our LLM take some inputs and then it wants to predict the next token. It wants to take this action. So in order to do that, it uses the probability distribution that it outputs and then the token that it predicts or the output that it generates then receives some reward that is going to feed back into tuning the agent. And here's the LLM. Yep.
So the question is wouldn't that be expensive if we do uh this for all pairs? Why would that be expensive?
Mhm.
Uh yeah. So yeah, the question is what do you do in terms of how expensive that is? So yeah, typically people take batches for instance. Um but I would not so it is indeed expensive but we're going to see a little bit some order of magnitude and how that works. But you can think of this as uh being um kind of a training procedure that could be seen as as expensive as some other training procedure. There's nothing that makes this more expensive.
>> Okay, >> but we're going to see exactly. So I think you're there are some points that you're you know touching correctly. Uh there's some parts that are added compared to the regular you know supervised let's say supervised fine tuning uh setting and we're going to see that in a second.
H so the question is uh does the reward that you're getting from here uh I guess powerful enough for the LM to change well you will see some expressions on the internet just characterizing this uh training procedure as you know like it's nearly not as many signals as what you would get for SFT let's say because for SFT you're literally always taking a partial input and making the LLM learn how to um generate the next token. But here you're only getting roughly one one signal per completion. So yeah, definitely it's more sparse and that's why you will see that RHF is seen as more an approach that has sparse signals.
Yeah. Yeah.
Great question. So the question is, do you apply reward for each token or for the whole thing? We're going to see this in more detail, but it's for the whole thing. It's for the whole thing, but we'll see that in uh in just a second.
Yeah.
So the AD and H. Oh, right. So the question is what is the AT and ST? So uh just as a reminder ST is the state that you're in and 80 is the action you want to take. So in the context of an LLM uh the state that you're in for an LLM is the input that you have that you have so far and then the action is which token you want to generate given that input. Yeah.
Cool. Yeah. So far so good.
Okay, perfect. So now that we know a little bit what a mental model could be for I guess LLM based RL, I just want to highlight once more that what we're trying to achieve is learn how to align this policy with rewards.
So we want to learn theta and theta is the parameters of our LLM such that pi of theta align with preferences and that's where RHF comes into play. So RHF stands for reinforcement learning from human feedback and it is typically composed of two stages. So the first stage is you figuring out how to distinguish good output from bad output.
So here you know all the preference pairs that you collected are actually used for you to learn what is good what is bad. So the input here is the concatenation of the prompt and the response and the output is a score.
So what you want to know given a prompt and a response how good that is.
And then the second step is the RL step, the reinforcement learning step. And this is where you use the rewards to align your model with the preferences. So here as input you have the prompt and what you somehow want to do is to be able to generate yhat that is more aligned with the rewards.
And by the way I just want to call out one thing. So RLHF is reinforcement learning from human feedback.
The human feedback part refers to the labels on which the reward model is trained.
So if the preference pairs are based on human ratings then we're relying on human preferences then we're in RHF because you will see out there there is also RL let's say Aif reinforcement learning from AI feedback and that one is relying on nonhuman preferences.
So yeah cool. Okay, so now that we know what RHF is and we know that there are two steps, we'll go through the first step naturally. So here the idea is we want to construct a model that knows which output is good, which output is bad. And of course here what we want is to not only consider the output but also the input because uh you need to somehow contextualize that answer.
So let's go with the our favorite example. Uh so let's suppose you have the following prompt. Suggest a new activity I could do with my teddy bear.
So what you want is to have a reward model that here we note RM that tells you that you know the answer that we rewrote into the good answer is good and we want to somehow have that model that tells us that the output that we constructed or that we did not construct the output that we saw was bad is bad.
So we want to somehow have a model that takes in the prompt and the good response and say it's you know good and we want to somehow have a model that uh you know takes the prompt and the bad response to say it's bad. So now the question is how do you construct such a model?
Well in order to do that we are using a formulation that's called the Bradley Terry formulation.
So that one is an important formula. So uh we just stay here for a little bit.
So what it what it says is that the probability bless you uh to have an output yi be better than an output yj is equal to an exponential of some score which is a score with respect to i over the exponential of that score with respect to i plus the exponential of some 4 respect to J.
So that is called the Bradley Terry formulation and this is the formulation we will use to build our model.
So here um it's also equal to sigma of RA I minus RJ. So who knows what sigma is?
>> Yes, exactly. So it's a sigmoid. So just as a reminder sigmoid is 1 / 1 + exponential of minus x. So this is how uh the graph looks like. So when it's minus infinity it's uh it tends towards zero and then when it's towards plus infinity it tends uh towards one.
So in other words, if I is better than J, what we want is for the input to sigma to be as high as possible because you want the probability to be as close to one.
So you want to somehow have R I be high if the output I is good.
and somehow our J to be low if the output is bad.
So far so good.
So here what we want to do is to somehow train a model that is able to output scores RA I and RJ using this formulation.
So this formulation involves two quantities because we have a pair-wise data set.
So far so good. Okay. So it will become more clear in a second. So here for training the idea here is you have some model that you initialize and you input on the one hand your prompt X and your winning output. So I'm saying winning. So it's yhat W.
You put it into the model.
It produces score R of X and Y W.
And then you have a second output. So X and Y L Y L which is the losing output.
So you put it into the model and it has a second score.
And now you somehow want to have a loss function that takes into account these two scores.
So based on this formulation, do you have like a some suggestion as to which loss function to use?
So here our loss function would be parise.
Yep.
Yeah. Uh so the uh the proposed answer is a binary cony. Um so can you explicit that a bit more?
>> Mhm.
>> Uh-huh. Okay. So um I guess in that case what would it be?
Yep.
Yeah.
Yeah. Yeah. Exactly. So I guess the Yes.
So great great answer. So I guess your answer is having uh some negative log likelihoods of this quantity which is uh so the cross entropy can be seen as kind of a special case of this but so just so that that we make sure that we uh get this part. Um so in order to just get that uh formulation which you mentioned again I guess one idea that we can have is given our data and given this formulation that we saw the Bradley Terry one to find parameters theta that maximize the probability of that data happening which basically leads leads to what you mentioned. Uh so here I'm just going to write what that means so that just we can um kind of be all aligned. So here it's so let's suppose you have a preference data set of let's say um you know winning example associated with losing example. So you have a bunch of pairs like this. So let's suppose that these pairs they happened independently of one another. So what you want is to find parameters that somehow maximize the probability of you seeing those examples.
Right? So here what that means is we want to somehow maximize the product of let's say these n preference pairs which is the probability of you know uh like some output uh w being uh better than some output l and we saw that with the Bradley Terry formulation. So we have this formulation which is basically the product of uh y I1 to n of this uh sigma of this reward which is basically a function of the input prompt and the output like this.
So what I'm doing here is to just reconstruct the loss from first principles.
So whenever you see a product of probabilities, the first thing you need to think about is yes log because this can become very small. So it can cause instabilities. So if you take the log, if you want to maximize a product is the same as maximizing the log of that. So if you take the log of that. So let's suppose I'm taking let's say the log of that.
Let's say I'm taking the log of that.
So it's basically equal to the sum of the log of sigma of r of the winning minus the one from the losing.
So we want to maximize that.
But people in ML they like to minimize things. So we're going to take like have a negative in front. So maximizing plus the sum of the log is the same as minimizing minus the sum of the log. And so this will be our loss function.
Does that make sense?
And this is exactly what you mentioned.
And typically we uh you know write the expectation and this is our loss function.
So our last function is minus the expectation of log of sigma r and x and y y w minus the reward of the losing.
Sounds good. So here the last function is pairwise but the reward model is it pair wise or is it pointwise like I guess do you need a pair to make a prediction or do you need just one?
>> Do you need a pair pair? Well that's the beauty of things.
You don't need a pair. So you're just training it pair-wise but it's actually pointwise.
So I think that's one thing I kind of realized. You know I look at this loss and you know this is a beautiful thing.
So you're training in a pair wise way, but at the end of the day, you have a reward model that takes in one prompt and its output and just outputs one number.
And we saw that it will try to output a high score for, you know, winning um examples and then a low score for losing examples.
So this is our reward model. And I'm looking at the time. I'm actually not on time, so I'll just move on. Um, so typically here what you would need is a set of data, which is typically on the tens of thousands or maybe even more.
Um, and here the label would be the preference rating. The preference here should come from humans if we're talking about RLHF.
And in terms of model, well, you have a bunch of different choices. Uh so as we know now uh models that are decoder only that predict the next token are very popular. So you could very well take such an LLM and just have some classification heads at the end of your sentence that you're uh taking into consideration and then use that to predict the reward. Or if you remember uh we've also seen encoder only models uh in the beginning of the class. So for instance BERS you could typically also project the embedding of the CLS token.
This could very well be an option.
So typically what people do nowadays is take the LLM route because everything is an LM these days. So it's like decoder only. You just put a classification head.
And so people have also come up with benchmarks to evaluate how good you're doing. So just put a reference here. Reward bench in case you're interested is a pretty popular one.
Um but yeah, so at the end of this you get a reward model that takes in a prompt and a given response and gives you a score.
Yep.
>> Exactly. So the question is uh in some tasks the human preference would be something but in other tasks it may be something else. So typically those rewards they are with respect to a given dimension.
So they can be let's say is the output useful or is the output I don't know friendly is the output safe. So all of that are different dimensions. So they can be different reward models.
Um you can also have like some holistic score as well.
uh but yes so you need to define a dimension across which you're actually quantifying how good your response is.
So the ones that I mentioned are uh the common ones that you would you would find. Um another thing that I will say as well uh while we're at it so human ratings are very sensitive to the guidelines that you're also exposing.
So, we're not going to go into details here, but one important aspect of human preference data is to make sure that the guidelines you're telling your raiders are as objective as it could get. I mean, sometimes you cannot have them be super objective, but you you have the I guess the task of just making sure they're clear enough so that your human preferences uh they're not noisy because they can be noisy, but that's also one challenge.
Yep.
So the question is uh is the reward here like a regression or a classification.
So let's look at the at the loss function and here.
So well it's kind of hard I guess I would say um because the reward can be kind of something that can be interpreted as a score but we have a probabilistic formulation um so I'll probably frame that as a classification task because you have preference data is it better worse so it's like one or zero but at the end you use the rewards so it's not I purely a classification task because at the end you have this um but one thing to note is that the scale that these rewards are at is typically something that is scaled itself.
>> So at inference time you have some normalization procedure that normalizes that across your batch but yeah I would probably characterize the formulation more as a probabilistic one I would say.
Yep.
Yep.
Yep. So the question is do we normalize the score? So there are many different methods to I guess normalize things. We typically do. We typically do normalize.
So I guess if you're in a regression setting, you're trying to predict some score on a given scale, which you're you're not doing here. So it's kind of free form here.
Um so yeah, so there is some rescaling that happens. Yeah. You had a question.
>> Yeah. The question is uh can you tell us more about what a the output of a reward model is? So we're going to see that a bit later but you can think of you know good outputs as being let's say one bad outputs as being a minus two minus three. So it's basically on the continuous scale if you want.
We're going to see an example in a little bit. So hopefully this can Do you have a question? Perfect. Uh we're going to see an example. So hopefully that will be clear.
Cool. Um so what we did is train a reward model.
Now the second step is to use that reward model to align our model or and by model I mean the LLM. The LLM that went through the pre-training stage and the SFT stage is the model that we want to align with human preferences.
We're going to do that using reinforcement learning and we're going to do that using the reward model that we just constructed.
So the reward model here is the model that we obtained in step one and it allows us to distinguish good outputs and bad outputs.
So here is a general recipe of how you would align your model.
So first you would take your prompt as input.
Your LLM will generate a completion.
So here a completion means like a full response from the model. So by the way completion people use also the term roll out. So it generates a completion, a roll out, a full response and that full response along with the prompt goes into the reward model.
So as we saw the reward model right now knows if it's a good, if it's bad. Let's say right now it's bad. So in practice it does not generate the the thumbs down. It generates some kind of score.
So if you if you want like minus two, let's suppose.
So we take that reward into consideration and then what we do is we tune the LLM with this information.
So it's more complicated than that and we'll see how we go about doing this but this is the general idea.
So just as a reminder, the reward model is the model that we trained as step one and it is a model that is frozen.
We're not training the reward model. The reward model has been trained.
The model that we're training is the the LLM.
So I think that's an important point to note.
And our goal is to optimize for higher rewards but then without going too far from the initial model.
So I think the you know optimizing for higher rewards I think everyone agrees with me right. But I I I just said a second statement which is but we don't want to go too far from the initial model.
So why would that be like why would we want to not deviate too much from the base model?
So one suggested answer is it will catastrophically forget what it has learned. Why would that be a problem?
great. Yeah, I think it's a great way to put it. So the the the suggested answer here is uh you have all this knowledge in your initial model which is pre-trained and then instruction tuned or tuned and you don't want to move away too much. This is exactly it. So that's one reason.
What could be another reason? So definitely one, it's a very very good reason. What could be another reason?
Yeah.
Yeah. So uh overfeeding on that data. So can you tell me more?
>> Yeah, great point. So the second suggested answer is uh you can overfeit on a data that is not super clean. And I will go a bit more into detail on what that means.
Yep. You want to add?
Yeah. Yeah. Exactly. So I guess your two points are I guess along the same lines which is the reward model can be noisy.
So that's exactly it. So, it's a phenomenon that's called reward hacking where if you are trying to optimize too much for the rewards, you're basically assuming that the reward is exactly quantifying what you want, but it may not be the case because the reward model itself is imperfect.
So just to make sure that everyone gets the point going to illustrate that with an an example. Um so let's suppose that I'm giving a lecture and my objective is for the lecture to be as informative as possible.
So I can take the choice to have something that I can quantify to know if what I'm saying is informative. So let's suppose if my reward is uh how loud the claps are at the end of the lecture if the clapping is very is a lot or not.
The problem is if I optimize too much for the volume of clapping at the end.
It may be the case that if I found out that you all like jokes and I start making jokes and then you clap at the end, you know, my reward is maximized but my objective is not fulfilled which is make the lecture informative.
So in practice, this reward hacking phenomenon is something that leads your model to optimize too much on a metric or a score that is imperfect. So you can very well maximize the rewards but not fulfill what you want. And this is exactly what the idea behind reward hacking is.
So yeah and another reason by the way.
So it's one reason is because uh your base model is already great. it knows a lot of things. Second one is because the reward model is imperfect. So you have this reward hacking phenomenon. The third reason is also because you can have some training instabilities.
Anyways, we have a bunch of reasons why we don't want to deviate too much.
And so in this setup, you typically have more observations than in the reward model case. So you have typically at least 100k observations and the label that you get here is from the reward model and uh so you start with the model that you obtain that's after SFT. So that's the one that you train. And the last function will be something that will try for you to maximize the rewards, but also to not have you deviate too much from the base model. So we're going to see exactly how we can do that. But that's the idea of how our loss function will look like. It will have these two parts.
Sounds good.
And as we saw uh so we don't want it to deviate too much because of reward hacking and instability and and so on.
So in order to do that you may have heard a very popular RL algorithm that's called PO.
So PO stands for proximal policy optimization.
So the reason why it's named proximal is because we do not want our model to deviate too much from the base model and we're going to see how we do that.
So this is roughly how our loss function will look like. So we'll have a part that is a reward model that we want to maximize and then some other part that is around uh measuring how far the distribution of the policy that we're training is to the reference model or the base model. So here the reference model would be your SFT model.
So you have these two parts.
Does the formulation of the loss roughly make sense so far? That make sense?
Cool. Uh okay. Kale divergence. Um so I'm just talking about Kale divergence, but I don't want to just assume you know what that is. Um so the Kale divergence is some measure. I don't want to say distance because distance has a lot of connotations to it. It's not a distance but it's some measure of how far two probability distributions are.
So you take a probability distribution let's say P and another one probability distribution Q and you compute this um formula which is the sum of PI log of PI over QI just gives you a number of how distant or how far these two uh probability distributions are.
So you want to minimize the loss.
So, I guess you want to minimize how far they are.
Um, does that make sense that we're using this? Yeah. Um, okay. So, just one little question. So, okay. Is the K KL divergence positive or negative or like what what can you tell me about the KL divergence?
Have you seen the Yeah, it's positive.
Why is it positive?
Yeah, but why?
So maybe we can come later. But great, great. I think the, you know, the result is here, but I guess I want to know why.
Yes.
Ah, yes. Exactly. So yeah, if you use Jensen's inequality because uh when you look at this formula, I mean luck can take like positive and negative values.
It's not super clear, but uh yeah, the the way you prove that it's always positive or equal to zero is with the Jensen's inequality.
Um and so yes, this quantity is equal to zero if and only if P and Q are equal.
Cool.
So with that, we're going to go a little bit more in depth into um the loss function. So I mentioned that the loss function tries to maximize for rewards but also make it so that the model does not deviate too much from the base model.
Well, there's something that I've like kind of lied a little bit about you, which is we don't want to maximize rewards. We actually want to maximize a quantity called advantage.
So what is the advantage? So you can think of the advantage as being how much better is your output compared to what you would expect.
So people typically use advantage to make the whole process be I guess more stable the training more stable and we estimate that advantage with a function called the value function which we will see in a little bit. Did you have a question? No, no question. So we we we have two quantities right? We have the reward that we saw previously how to compute which is given a prompt given a completion how good is this you know prompt and completion and then the value function is another estimation but this estimation is not at the completion level is actually at the token level.
So what we're doing here is we're taking the prompt along with the partial generation as input and we're trying to estimate what the reward would be if we were to continue to generate the output following the policy.
So I'm just going to repeat this one more time because it's charged in meaning. The value function is a token level estimation of the rewards that takes in a partial input and predicts the reward.
So if that predicts a reward if you were to follow the policy in your generation.
So this value function is something you will see in papers as something that people have to train when they do PO and it's typically because people need to have this value function when they compute the advantage that I mentioned above and this value function is typically trained jointly with the policy.
So in other terms you have your LLM which is initialized as the SFT model which you know generates the prediction.
So what you typically do is also have some head value head that also so this time is not a classification problem it's a regression problem. you try to estimate the final reward if you were to continue to generate according to policy.
So this is typically something that you also need to train jointly with the policy and I am intentionally not going too much into detail as to exactly how that advantage is formulated. It is just beyond the scope of this class. But in case you're interested, I would highly recommend reading this paper which is uh mathematically involved about the generalized advantage estimation methods which people typically use for this. So it's this highdimensional continuous control using generalized advantage estimation paper. If you if you're interested just feel free to take a look. But this is typically the method that people use to estimate these advantages.
So now we're going to see in a second where those advantages are used. But just as a high level, does that make sense to everyone? What we're trying to achieve here?
Um yeah. So um in other words, what we want is to maximize the rewards but not deviate too much from the base model.
But what we want is to actually find a way to make the rewards a bit more relative as to what you would actually expect.
So you want to compare the rewards that you're obtaining with kind of an average response if you want like the idea is okay if you are having a great reward what would be the reward kind of in the average case and you want to still maximize that and the reason why people use this reward minus baseline is because it reduces the variance of these estimates and it just makes the training just kind of go faster typically. So that's why people use that part.
But in order to compute the advantage, what people use is a method called generalized advantage estimation which you can think of it as like a big formula with some hyperparameters.
Uh that is in the paper below.
Just a general idea.
Cool. So, okay, looking at the time. So, we're going to look at two typical variants of the PO loss.
The first one is called PO clip.
So, the idea here is to prevent the updates to our model from being too large from one iteration to the next.
So what you have is this loss formulation which is the minimum of the ratio and we're going to see what that is times the advantage and some clipping of that ratio between two uh bands which we're going to see in a second times that advantage.
So before we go into details, I just want to call out a few things that may seem like it's confusing because it was confusing to me. Um so first of all that formula is expressed as L is equal to but it's actually not something we want to minimize something that we want to maximize.
So what people typically mean by L is loss.
Um so yeah just know that we want to maximize that we don't want to minimize.
The second thing is we're introducing a term called R of theta.
This is not the reward model.
So typically people use R to say it's the reward but here it's not the reward.
It's actually the ratio in the probabilities between our policy and the policy from the previous step which is denoted pi old pi theta old.
So it quantifies how different your policy at the current stage is to your old policy.
So here when we say old, it's not the SFT model. It's not the SFT model.
It's the model at the previous iteration during the RL stage.
Because here, what we want is for the policy to not have updates that are too wide.
So far so good.
Yeah.
What is a? Uh a is our advantage.
So you can think of it as you know this complicated formula that's a function of the reward and the value function that just tells you how good your uh your your piece is basically.
Okay. So this um complicated formula we're just going to see why it makes sense.
So if the advantage is positive, it means that we somehow want to reinforce whatever was generated.
And so here our objective function will look like this. So if a is positive then this L of clip will look like some linear function as a function of R because it's R * A.
So it's a it's a linear function as a function of R and then after 1 + epsilon it will just be flat.
So why why do we want such a such a shape for our graph? Well, we need to think again about what R means. So if A is positive, meaning it's something that we want to reinforce, we typically want the probability of whatever happened to be increased.
So we want R to be higher, but we don't want it to be too too high.
So that's why we have this clipping mechanism.
So that's why we have this uh this clipping here. We want to I guess maximize L. So we go uh towards increasing R. But we don't want to increase it increase R too much.
So that's the rational for this.
But when a is negative, it means that somehow what we've done is something that we don't want to do as much.
So in order to uh maximize L, what we want is to actually decrease the probability of just generating what we generated. So we want R to be smaller because it also increases L, but we don't want it to be much smaller. We don't want too big of an update, which is why we have this clipping mechanism here.
So in other words, this super complicated formula can be just intuitively explained with what I just mentioned, which is if you have positive advantages, you want to reinforce the fact of generating this output because this is something that the model liked. I mean the by model I mean the reward model. Um but you don't want it to update it too much which is why you have this clipping and same when the advantage is less than zero. Um I guess you want to downweight the probability of this output from happening. So you want to have a smaller R which is your ratio between the policy you're training with respect to the old one. But you don't want to be too small either again to not make too big of an update.
Does this make sense?
Does this roughly make sense? Okay, roughly makes cool. Just I will just um recommend just rereading this formula.
I'm just uh remembering what I mentioned and hopefully it would make a bit more sense. Um but yeah, happy to answer any questions if there are any. Yeah. Oh, yeah.
Yes.
Yeah. Uh so question is uh what happens uh where it's not continue. It's the same as ReLU, right? Like you you would basically use the same methods. Um so yeah, it's the same same thing. Yep.
So great question. question is why are you using from the previous iteration and not the base model? Because actually we don't want just to not be too far from the base model. We also want our updates to be not too extreme from one iteration to another. And this is to uh improve with uh stability of training.
So this is just like another additional constraint we want.
uh question is are we considering the base model? We will consider it and I will tell you exactly but u in this in this formulation we're not in this formulation we're not but in the formulation we're using typically for LLM based RL we are and we will see how Yep.
Yep.
Yeah. Yeah. Yeah. So the question is how do you quantify R like what is R? What does it mean? So it's typically u the probability of a token happening with respect to the input. So you can think of the probability distribution that you get at the output of your LM. So you would use that.
Cool. So this is the first variant that the PO paper mentions. So it's called PO clip.
So the second variant that we have is the one that uses KL divergence which is called KL penalty. And this one is a function of the ratio times the advantage minus the K divergence between the old policy which is the previous iteration and the current iteration.
So you had a great point which is why are we using the old? In fact, old means previous RL iteration. But nowadays, people use the reference. So, it's just a change because this paper has been published in 2017 where in 2025 LLMs, they kind of came more into like something that was popular in the 2020s.
So, um, so yeah, there's just been this like difference in what you put there.
So yeah, it's indeed the reference that you would typically put in a Kyle divergence.
Uh the other thing that I will say is that nowadays when someone wants to train uh a policy using PO, they would typically use the kale divergence and they would also maybe mix it with the clipping clipping thing with the old iteration.
they were basically mix the two.
There are many different formulations of what the loss can be. So what I'm saying is not always true, but this is what people sometimes do. They sometime mix the two.
Cool.
So, I talked to you about PO and how it allowed you to maximize rewards, maximize advantages by not making your model deviate too much from the base model and also from the previous iteration.
But the problem is that for PPU, you need a lot of models.
So you need your policy, you need your LLM that you're training. You also need the value function which I me which what I mentioned is is used for the advantage estimation.
You also need the reward model of course and then you need the base model. So the base model is frozen. It's the just the SFT model just to be able to compare with the current policy. So it's a lot of models.
So now the question that you may ask is do you really need to have all these models to perform well?
And the answer is maybe not. Maybe but maybe not. And nowadays there is a bunch of other methods that are becoming more popular. So you may have heard of GRPO in the context of reasoning models. So we will look at GRPO in more detail next week for lecture six. Um but just know that there are many variants. So PO is just one algorithm that has been very popular and even within PO there's been many variants and nowadays there are many other RL algorithms that are being used.
Yeah.
And uh yeah if you want to just uh anticipate next lecture um you can also take a look at this paper from uh the deepseek team. So deepse math limits of uh mathematical reasoning in open language models that introduces the gRPO algorithm. But if you don't it's fine we're going to talk about it next week.
So I also want to talk about some challenges when it comes to RLbased approaches.
So the first one is you need to have this two-stage process. You need to first train your reward model and then you need to use the reward model to train your policy.
But the problem is you do step one, you do step two and then you realize, oh wait, step one has a problem. then you need to redo everything. So there is like a bunch of dependencies that that can really you know make your life harder and um you know just the training recipe is not the simplest. So that's one downside.
The second downside is about the number of hyperparameters that you need to tune. So we saw beta from the Kyle divergence formulation. So that's one.
We have epsilon from the PO clip version. That's another one. The generalized advantage estimation. We have not seen the formula, but there's I think at least two hyperparameters that I can think of.
Anyways, you have a bunch of them.
So of course, you know, some hyperparameters perform better than others. And you know, if you need to retrain, then well, you need to do everything again.
um then you have some instability challenges. So you know you can constrain your iterations to not be too wides to kind of control for that but sometimes it's not uh you know sufficient.
So in terms of the metric that you can use to monitor the training, what what is the metric you would use?
So we're trying to maximize rewards in some sense. So typically the metric that you would use to monitor your training would be the average rewards.
But this is not like a great I mean this is a way to to you know monitor your training but this is not like the best way necessarily.
So at least when it came to pre-training and SFT you had this cross entropy loss that you were monitoring that was really telling you how your model was able to you know just reproduce the behavior you were trying to make it uh output.
But yes, so this is another challenge.
And then the other thing is in the RL world, in order for you to know which completion you should not do and which completion you should do more, you need to have some diversity whenever you generate. So if during the RL training loop, every time you have a prompt, you generate something and then you retry and you generate something else that's too close.
Well, it's not it's not great because you're not exploring the set of possible completions. So this is another challenge you need to have in mind which is you need to force your model to do some exploration to see I guess what can be good.
Um, and then the last point is I mean a lot of you were not familiar with RL. I was not and it's not super clear why we would necessarily need need RL to do this stage.
Um, also I will say one thing. So this RLbased training you will hear a lot the term on policy training. I just want to tell you what that means. So what is different between this and SFT is during SFT we had some prompt and some response that we wanted our model to generate to mimic.
But here what we're doing is at each iteration we're asking our model to generate an output and then we're using how good that output is to update our model. So there is a core difference between the two because in here what we're doing is we're asking the model to generate something whereas for SFT we're just using some you know some data that is not necessarily generated by our model.
So on policy training refers to training that involves the model generating from its current policy and optimizing on that.
And this is in contrast with off policy training which is relying on generations that are not done by the model that you're training.
Does this make sense?
Yeah. So this is a term you will see a lot which is why I'm telling you. Um and yeah so PO is an on policy algorithm.
Yeah.
So the question is why don't we do SFT on this preference data? Well SFT you you tell the model you should generate this given this input. then you should generate this given this input. But you're not telling it about things it should not generate.
Like let's suppose you have something that you don't want to happen unless you rewrite that into you know a perfect sentence. There is no way for you in a traditional SFT framework to tell your model to not generate that.
But that being said, it's a great prompt because later on in let's say 10 minutes, we will see a supervised way of doing what we did with RL.
But yeah, it will be in about 10 minutes.
But apart from that, does that roughly make sense?
Yeah, perfect. So in the next few minutes we will see a method that is actually used more often than not and this is for people who have a reward model but that don't want to do RL.
So why do you not want to do RL because you know it's expensive it's you need a lot of models. So to your points it's indeed expensive because you need all these like models and you know sometimes it's just too hard to tune that training. So you have this method called best of n or bun and the idea here is to leverage the reward model that you obtained in order to choose the output that you will return to the user.
So the idea here is you have a prompt and you tell your models okay I you actually want to have several um generations. So you generate let's say I don't know four times five times and what you do is you give a score to each prompt completion and you just return the best one the top rated one which is why it's called best of n. So you have n completions you rate all of them and you just return the best one.
So you have your prompts. So, uh, our favorite example suggests a new activity I could do with my teddy bear. You put it into your SFT model. So, you're not training your SFT model. It's asis and you're just generating, let's suppose, three completions. Uh, the first one is a great answer. The second one is no, don't spend time with the teddy bear, which is not a good answer.
And then the third one, I think it's also a good one, take your teddy bear to a picnic. You put all of these into your reward model and this is where I wanted uh to respond to you like you know what is an output of the reward model. So it would be a score like this. So for instance the first one would be you know pretty good uh completion. So it's like 0.8 the second one is not good at all.
Let's say minus two and the third one is like let's say so so in between.
So the best of end method would select the top rated one which is the first one and this is the one that you would return.
Okay. So what is one problem with this approach?
So the the answer is uh if the model is bad then it's going to be bad. That's fair. But I guess if you uh make enough completions with let's say a high enough temperature, you will have some diversity.
Um so it's it's a fair point uh which can be a problem. But let's suppose it's it's indeed a problem. Let's suppose it's not a problem. What is another problem that you may have?
Yep.
Yes. Yes. So the answer is uh you query the model many times which is indeed the problem. So you write inference what you want is to not spend a bunch of money and compute but here what you're doing is uh complete your prompt multiple times and this is indeed the main challenge with this. So you indeed skip the RL training but you're basically pushing all the work to your inference to the inference stage and uh this is just making everything so super costly. So when can it be bad? So um if for instance you're in a in a case where you need to serve your model to let's say a bunch of traffic it may not make sense for you just um just from a cost standpoint.
So um I would say before taking that route it may be good to just assess what will be your inference traffic what would be your training cost and kind of uh determine that from there.
And um yeah, I think that's pretty much it. Do we have any questions on that part?
Oh yeah. Yeah.
Um that sigmoid would be softmax. Are you talking about this one?
Um so in that algorithm you're assuming that your reward model is is trained.
So uh the best of end approach is you do this this training. So I guess the sigma here is uh pairwise because you trained your model in a pair wise fashion.
Um but in the best of end um case what you do is you just use your reward model to score end times and you just pick the the the highest one. Yeah.
Yep.
>> Yeah. The question is uh but what if all the responses are bad? This is definitely a fair a fair concern. Uh and this is indeed a concern. So your model needs to be at least a little bit good.
Yeah. Yeah. Yeah. Yeah. Exactly. Yeah.
Yeah.
Yeah.
It's a great question. And do you have any constraint on the scale? Well, depending on how your scaling is done, you actually don't care about the scale.
As long as you take the best one, the highest one, if you think about it, let's suppose if you scale it with, let's say, with a normal score, the highest one will still stay the highest one. So in this case, you actually do not care about the scale.
Yes, you also don't care about that.
Yep. But typically the scaling matters when it is coming into play in a loss function which is the case for RL and so that's where people use some like scaling mechanism. Um you know that's that's the part that matters.
Cool. And with that I'm going to give it to Shervin.
Thank you.
So, for this last part, we're going to talk about DPO.
Um, but I'm not going to tell you what DPO means just yet. We're just going to first um listen to the complaints we had in the RL world and then see um what we could do about it and then how have folks um like remedied it.
So the first thing that I'm going to repeat from one what Afinen covered is that in the RL formulation of PO you have to carry during your loss optimization a bunch of model weights.
So when you look at this um loss function you have uh so first the policy of the current model you have to optimize for the one of the old model or the reference model uh and then you have your advantage which hides the reward model as well as the value function that we discussed. So like all these four which is a lot and uh the second thing is that let's say you go with the approach of best of end you don't do RL you just generate multiple times uh and then like pick the best completion as Ein mentioned and um so you have this uh latency and cost issues and just uh for the sake of you know thoughts uh for a thought experiment. Let's suppose we had infinite money. Would it be still fine?
Um so let's suppose you generate all of these in parallel.
So you would still need to um wait for the maximum time of like the generation of an answer to give your resulting answer. And then if you look at the distribution of latency, usually it's some you know shape that is not centered on a single point. And when you look at the probability distribution of the max of let's say n answers, it will tend to be uh shifted to the right. So let's suppose you have infinite money, infinite compute, um you still have to wait more than you would have on a single pass.
Do these concerns make sense?
And okay, so there comes the question that someone was asking. Why don't we do supervision all this time? And this is the route that uh DPO folks have explored. So DPO stands for direct preference optimization. So instead of doing this uh expensive two-step process of first finding some reward model and then uh iterating on the model's weights, it optimizes a single loss function that directly optimizes for the model's weight. And as you can see, uh there is no reward involved anymore. So no R. and you have a way to formulate what you care about as a function of your preference pairs. So in that loss function all you have is a bunch of variables and then the sigmoid functions but within uh the parenthesis you have the probability of a given completion like to happen and uh you have this uh subtraction operation that compares the probability of the winning uh completion to happen versus the losing one. So as you can see it's a direct expression with respect to these preference pairs which is pretty nice and if you look at it a bit closer you will recognize what I've just mentioned regarding Bradlary formulation because you have this uh expected value of log of a sigmoid and then within that sigmoid you have a difference of terms where you can recognize the equivalent of what you saw as a reward term.
So I know it's a lot to see at once. The formula is not super pretty. Um yeah, do you have any questions?
And then we're going to discuss in a bit what that beta is and also how we obtained this formula in the first place.
Uh okay. And I one thing I will mention before going to that is that the paper is titled your language model is secretly a reward model. And the reason for that is that the reward um like placeholder that you could identify from this loss function is expressed as a function of your policy. So you don't have any Rs in there. But when you express the equivalent of a reward, you find the expression of your model um of your model directly. So which is um kind of pretty interesting insight.
Okay. So but but now let's think together about how we got there in the first place.
So you recall from a few few slides ago the PO objective that we had. So we wanted to maximize uh rewards and minimize the distance of the obtained policy with respect to the reference one which is this uh trade-off between maximizing rewards and scale diver divergence term. And then you have this beta coefficient that appears there that uh controls how how much you want to penalize going further from the base model. And what this paper did is that it wrote down this formula and then solved what would be the optimal solution. Uh so it expressed the optimal policy pi star and when you derive this expression so you get it as a function of r. So that's other term that was in the objective function and uh so like all these steps have no extra assumptions. It's just a bunch of uh derivations that was done. And uh what you can see here is like this Z term that stands for a partition function and that just normalizes things uh but it's not something that is new.
It's like a function of the other uh of the of the terms that were already in there.
So when you rearrange these terms you can uh equivalently express R as a function of this pi star.
Uh so it's just like a rearranging terms between these two expressions. Nothing um nothing too fancy.
And then uh the key part of of what this paper did is recognize this term as some reward and then express it as part of a Bradley formulation. So you remember you had this probability of um of a completion of a winning completion being bigger than a losing completion and you had this sigmoid of uh like the difference between the two Rs. So basically they took this formulation again and they plugged in the expression they had obtained for the optimal reward which we had seen was a function of the policy.
So at the end of it you plug in things you have no more rewards just a function of the other parameters. So you have this beta and uh only the the policy in there.
And the last step is the same one as Afin was mentioning. Once you have this probability that you want to maximize you can turn it into a loss function.
And this loss function is what this paper proposes to optimize in order to tune policy weights in order to align it with this preference data.
So yeah, I simplified a lot the quite heavy steps of this paper uh into like this these few steps. Do they make sense?
And uh the beta that we see here is the same beta as we had seen in the PO formulation. So typically taken um around 0.1 if you want some order of magnitude. So beta is a hyperparameter and all you want to optimize for is the policy and then the preference pairs you have them as input.
Make sense?
Okay, great. So now let's compare uh what this method gives with respect to RHF.
Um so RHF is this two-stage training process where you first fit a reward function and then use it uh in order to do the policy update steps. So as Afin was mentioning it, it's quite heavy. Um we were mentioning it just before you need a lot of models in memory and then training stability is is a concern. You don't have direct supervision. You rely on the policy of the current model to generate the next updates. So there is a lot of complexity in it and as opposed to that this DPU formulation gives a supervised way to think about preference tuning where you just take these preference preference pairs and you directly fit a loss function on them and um yeah and then instead of four models you just need two because you had this pi of theta and then a pi ref which was your base SFT model and you don't need any other model copi copies.
So now you might want to ask um if it's um if it's much easier, why um why is not everyone just simply using DPO? So it's not as easy. Uh in practice um there are some pros and cons to each method. Um and the paper that is listed below uh gives a very interesting study of um you know like the benchmarking differences and the steps to get there.
Uh so just to get the main lines overall PO performs better than DPO. Um and um one great thing about DPO which is like this supervision this direct supervision comes at the cost of sometimes fitting something that is not the exact distribution that the model has seen as training time. So this paper talks about it a bit. um if you train via SFT on the preference related data, you could get better performance. But this shift in distribution is uh you know a challenge intrinsic to to DPO.
uh because one thing that the DPO method enables you uh to have is just to not worry about that reward modeling at all and just get some preference data set out there. But the issue with with doing so is exactly that distribution shift.
Uh so you could you know either train uh like SFT on it or generate it yourself and then rate it but that is u like yet another cost that you would need to uh to pay.
Okay, great.
Um and yep, any questions?
Uh so the question is uh you are optimizing the parameters with respect to the base model. So yes but it's not the base model it's the current one. You have a base model that will stay fixed in the iterations and that occurs in the loss function and you're indeed uh updating parameters and that will be like the preference tuned one. Yeah. So you have like two weights two set of weights one frozen one one that you update. Yep. Okay, great.
Any other questions?
Yes. So the question is uh instead of the reference model, could you use another model uh another large model? So I think that is a that's an interesting point.
I guess that would be a different algorithm.
Um, and you might lose a bit the meaning of what you're trying to do because this preference tuning stage is trying to align your responses to what you prefer and you want them to align with respect to what you had learned so far. So if you like start from a different model and then you preference tune like a different probability distribution, you might be doing things that might not make as much sense from a loss standpoint. But I guess you know the direction I mean I I cannot tell like a blanket answer there. You know, maybe it could give like interesting behavior, but my intuition would tell you that the point of all of this is to start from the last train checkpoint that you had and then further like preference tune the distribution of your completions, >> right? So like the the suggestion was you know if you change the reference model that could be seen as another SFT uh version of your model. Yeah I guess so.
Okay great. And now the favorite my favorite part of of today. So do you all remember the completions we had obtained on our favorite prompts? So can I put my teddy bear in the washer? So if you remember last episode we had said uh after instruction tuning the response was no it might get damaged try hand washing it instead.
So does anyone have any complaints regarding this answer?
So I can give a hint teddy bears they indeed should be handwashed. So, it's factually correct.
But let's suppose it's factually correct. What else do you not like? I mean, could you not like about this answer?
Yeah. So, the suggestion was uh is asking to try it instead. And I guess like a broader point is it's a bit too rough, doesn't sound super friendly. And then you know someone who asked this question might love teddy bear. So did you want you might want to tell them about it in a gentle way. And uh this is exactly what a preference tuning stage is about. We don't want to learn new facts. We want to focus on the existing distribution of completions and align what the model should return with respect to what the human prefers.
So on this intuition, the preference tunes answer could be um like a gentler one. It's better not to your teddy bear could get hurt. Uh gentle hand wash is safer. So it's conveying the same points but in a tone that's uh that sounds like much better to the ears of of the person who asks.
Okay, great. Any questions on the DPO part?
Yep.
Yeah. The question is um like what are the trade-offs and then which one is used in practice. So I think it all depends on your compute budget on how much you want to worry about performance versus babysitting your training process. PO like the RL bar is very hard to tune right. It will give the highest uh results. Uh but you might get almost at that level with much less efforts. So I guess if you want to get like a quick preference tuning that is shown to like show good results but maybe not not the best results and DPO is your is your friend. And if you're an expert at uh reinforcement learning, you you know what you're doing and you want the every last bit of performance, then PO might be like a better way.
Okay. Amazing. And with that, thank you very much.
Up Next

Stanford CS229: Building Large Language Models (LLMs) | ML Lecture
@stanfordonline
1.8M views•2024-08-27

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

Multi-Task Learning Basics | Stanford CS330 Lecture 2
@stanfordonline
28.8K views•2023-03-31

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence






































