Online DPO (Direct Preference Optimization) is an alignment method that fine-tunes large language models by generating preference data on-the-fly during training, using a reward model to select preferred completions for each prompt, which eliminates the need for pre-collected preference datasets and enables continuous model improvement while achieving better results than traditional DPO.
Online DPO Fine-Tuning for LLMs: Hands-On Implementation Guide
Added:hello everyone and welcome to the channel I'm very excited to report that now hugging face celebrated TRL Library supports yet another fine-tuning method called as online DPO or online direct preference optimization before I show you the Hands-On demo as how to get this DRL installed and then how to do the online dpu on your own data set let me first set the stage by explaining what exactly is meant by fine-tuning what is this online DPO and why this is such a important Milestone first and foremost TRL is a library from hugging phase which is a full stack library that provides a set of tools to train transformer language models with reinforcement learning from supervised fine tuning to reward modeling to various other things now let's try to simplify these Concepts what exactly is fine-tuning in simple words what happens is that all of these large language models no matter what the model is they are trained on a huge generic set of data but they that data and that model doesn't really reflect your own data your own requirements if you want to align the model as per your own preferences for example you just want the model to always align with your values your goals your preferences then you can create a data set and then F tune or retrain that model as per your own data set so that for example you don't like Winters so you can train the model on that item so that whenever someone asks the model about winter model would say that model doesn't like winter because you have given that preference to the model and you have train the model on that data set that is a very high level example of course but you get the point and that is what called as direct preference optimization where you give it a preference now for the fine tuning there are lot of methods for example if you look at this diagram and by the way this and this diagram I have taken from Maxim labone who a really um good researcher out there and I have interviewed him for my channel if you're interested just search for his name on the channel Okay so now what exactly here is happening is that this is a diagram for the PO algorithm which is the proximal policy optimization now don't get too much worried about the naming here so all what it is happening here is that we are fine-tuning the models from Human preferences that is what is generically at high level called as uh reinforcement learning with human feedback where they have presented this framework where a reward model is trained to approximate human feedback and then this reward model is used to optimize the fine-tuned models policy using the proximal policy algorithm now the core Concept in this diagram is revolving around making smaller incremental updates to the policy as larger updates can lead to instability the problem where is that there is lot of loss time convergence it is difficult to reproduce and it is quite expensive to be honest that is where a new technique emerged a few months back called as direct preference optimization or DPO DPO simplifies this by treating the task as a classification problem it only uses two models train model or a copy of it called the reference model during the training the goal is to make sure that the train model outputs higher probabilities for preferred answers than the reference model and we also wanted to Output lower probabilities for rejected answer so what it what happens is that we give it a prompt we give it a preferred answer and a rejected answer and then we find you the model so that model always know what answer is the acceptable one as per her own alignments and preferences so that is what uh direct preference optimization is now recently and very very recently this new technique has emerged called as online DPO or online direct preference optimization it's a new alignment method from Deep Mind to boost the performance of llms with online DPO data is generated on the Fly by the train model instead of pre-collected for each prompt two completions are generated with a reward model selecting the preferred one what this approach does is it eliminates the need for a pre-collected preference data set it is ated online and that is really amazing it also enables continuous model Improvement and it yields better results from a traditional DPO so that is what has been now merged in this hugging faces TRL library and that is what we are going to do in this video we will install TRL library from The Source from the recent branch and then we will see how to fine tune a model on online uh DPO data set before I do that let me give a huge shout out to M compute who are sponsoring the VM and GPU for this video If you're looking to rent a GPU on affordable prices I will drop the link to their website in video's description plus I'm also going to give you a coupon code of 50% discount on range of gpus so this is my one2 terminal with 22.4 and this is my GPU card nvd RTX a6000 okay so let me clear the screen first up let me create a virtual environment with Konda let's wait for it to get activated the virtual environment is created let's install all the prerequisites here I'm installing T torch Transformers and lot of other stuff so let's wait for it to finish and don't worry about the code I'm also going to provide you the link to the code which I'm using from start to end all the prerequisites are done next up after clearing the screen let's get clone the repo of TRL and the reason why I am doing it is that this feature is very very new so we have to check out the recent Branch where they have committed this if you're using it like after maybe couple of weeks then you might not have to do it you can install it from The Source but at uh it I don't think so it has been merged into the main branch so that is why I have switched to that uh commit HSA so let me install it from that Branch let's wait for it shouldn't take too long and these are the joys of working at the bleeding edge anyway so it is installed now let's Now launch our Jupiter notebook because that is where I'm going to install it and then see how it works so let's wait for it to launch it in the browser and then I will show you how you can use this Dr library to do online DPO on your own data set so what jupyter notebook is launched let me paste the code here now what this code is doing it is simply importing the libraries which we have installed here you can see and then it is just defining some of the dummy samples grabbing our tokenizer grabbing our model I'm just going to go with this small LM with 135 million parameter just to make things faster and easier and then we are defining so first up this is the model which we are use looking to optimize or fine tune and then this is the reference model same model and this is going to calculate the K Divergence again K Divergence measures how different two probability distributions are and U this mainly helps a model understand uncertainity and similarity okay so once that's done we are defining this reward model to score completions with of course uh you will definitely need a reward model for online dpu as I explained earlier and this is whereever training and eval data set is I'm just giving it this prompt and simple but of course you can just just replace it with your own data set here and then we are just calling our online DPO trainer with the model reference reort R tokenizer and all the usual stuff and then we are running the training so once I start it you will see it is going to prepare our data set and then from there it is going to map it and then get our models there you go so it has started the training it's a very small data set so it should be fairly quick I believe and this is how you do the online DPO but the main thing here is that to understand the reward model and this stuff and then getting your data set and this is how easy it becomes once you understand that uh how you should prepare your data set how should you just put a reward model and of course you can select different models for this one maybe you could have Lama 3.1 70 billion1 as a reward model so that your performance of the F your model would be more Superior and then there you go the fine tuning is done it is giving you the statistics around it that what was a training loss which has come down how much was a train runtime aox and all that stuff now if you want to see your new model which is fine tuned just go to your um terminal maybe I'll just open a new one and then I just deactivate my original base model base gond and then go to TRL if you do LS here you will see that there should be directory called as online DPO model if you do LS there should be a check point yep and this is your model file with all the tokenizer config do GSN and your models save ters here and then there's also some other Json files with vocabulary and stuff so this is how you create a new findu model from your own data set by using this online deep which seems quite promising new fine tuning technique so that's it I hope that you enjoyed it let me know what do you think if you like the content please consider subscribing to the channel thanks for watching
Up Next

Introduction to Secure Multiparty Computation with Yehuda Lindell
@fhe_org
7.7K views•2021-02-04

Supervised vs Unsupervised Learning: ML Foundations
@digiLab_ai
108 views•2023-06-22

HTTP Requests Explained: GET, POST, PUT, DELETE
@codecademy
103.1K views•2021-10-07

Enigma Machine Mechanics: WWII Encryption Explained
@JaredOwen
13.2M views•2021-12-11
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Computer Science










![LLM Crash Course Part 1 - Finetune Any LLM for your Custom Usecase End to End in under[1 hour]!!](https://i.ytimg.com/vi/whbuNo6APVs/maxresdefault.jpg)





















![[Lab Seminar] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena](https://i.ytimg.com/vi/OywgftYSMmc/maxresdefault.jpg)






