Reinforcement Learning Through Human Feedback (RLHF) is a framework that integrates human feedback into the training of reinforcement learning algorithms to accelerate learning and improve decision-making quality; this approach involves training a reward model that assesses answer quality based on human rankings, then using this reward model with algorithms like Proximal Policy Optimization to fine-tune AI systems such as chatbots, enabling them to generate responses that better align with human preferences and expectations.
Reinforcement Learning with Human Feedback (RLHF) Explained
Added:greetings fellow Learners now before we embark on this journey into reinforcement learning with human feedback I've got a thought-provoking question for you when learning something new when has feedback from others made a noticeable impact on your decision-making or learning this could be any experience that you had in your life so please share your thoughts down in the comments below and let's have a discussion we will divide this video into three passes where we start introducing the concept of reinforcement learning through human feedback and then provide some engaging examples along the way also pay attention because I'm going to quiz you along the way now let's get to it for this first pass let's have Frank help us explain Frank say hi hello what a cutie now this year is a grid world where there's nine squares and each square has a reward inside of it the goal for Frank is to get to this plus 10 reward spot and to do so Frank makes decisions that is to go either left right up or down but Frank doesn't know how to make any decisions to begin with and so Frank learns Maybe by interacting with the environment and he does so by using a reinforcement learning algorithm we've discussed details about a few of them in previous videos so you can check them out for specifics but effectively once Frank learns how to make decisions with any of these algorithms he will effectively be able to get to that plus 10 rewards spot efficiently but wait can we help Frank out even more and it turns out that we can so while Frank is learning with a reinforcement learning algorithm us humans can also provide our feedback to Frank as a mentor this allows Frank to learn faster and it also allows Frank to give responses that are more human favored quiz time have you been paying attention let's quiz you to find out which algorithms can be used along with human feedback a q learning B DQ learning C proximal policy optimization or D all the above comment your answer down below and let's have a discussion and if you think at this point that I deserve it please consider hitting that like button because it will help me a lot now that's going to do it for quiz time and for pass one but continue paying attention because I will be [Music] back all right Frank get over here yay in this past we're going to show you how Frank learns without human feedback and then add in human feedback in see how that affects things so want to get it started Frank sure H let me go down let me go down let me go right oh bad spot learning now starting over let me go down let me go right let me go up let me go right let me go down let me go down good spot learning now great so Frank will keep doing this as he learns but now let's see how a human can help starting over let me go right all right Frank so that's fine let me go down okay let me go left hm I would prefer you actually go right here okay going right let me go down good spot learning now so in this situation Frank was following an algorithm but I was still nudging him in the direction that I thought was the correct direction quiz time it's that time of video again have you been paying attention let's quiz you to find out how does human feedback contribute to reinforcement learning as Illustrated with Frank's grid World adventure a it acts as a randomizing factor B it accelerates the learning process C it slows down the learning process or D it has no impact on decision making comment your answer down below and let's have a discussion now that'll do it for quiz time for now but I still will be back so pay attention in this pass let's talk about how chat GPT makes use of reinforcement learning through human feedback for a practical application it's split into two parts so first is train a reward model to be a human advisor to chat GPT and then the second is use this rewards model along with an algorithm called proximal policy optimization to fine-tune chat GPT let's talk about each part now and starting with the rewards model this model is a GPT architecture that takes in a question and answer as input and the output of this GPT network is a number it's a score that says how good was this answer to this input question now higher the score better the response and our goal is is to first train this model we can do this by putting a question to a pre-trained chat GPT multiple times and each time we'll get a unique answer we as humans then take these responses and we rank them based on which was the best response versus the worst response and we use this then to train the rewards model and once the rewards model is trained it should be able to assess how good a given answer is to a given question so that was the first part which dealt with training the rewards model now on to the second part where we use this rewards model along with proximal policy optimization in order to fine-tune chat GPT so chat GPT is given a question it generates a response now this response is generated using the reinforcement learning algorithm called proximal policy optimization now for more information on how po works I recommend you check this video out it's a good one you won't regret it this response along with the question is passed to the rewards model to generate a number that's a score that tells us how good was this response to this question now we use this reward in the loss function for chat gpts Network we now perform back propagation so that chat GPT learns now this is just one iteration but we keep doing this for multiple iterations and once trained chat GPT becomes this public facing app that it is today for a more Deep dive on the entire process for chat GPT you can check out my playlist of videos right here but overall I hope you understand the real world use case of reinforcement learning through human feedback quiz time oh this is going to be a good one have you been paying attention let's quiz you to find out in shat gbt what is the primary purpose of the rewards Model A to generate unique answers to questions B to serve as a pre-trained chat GPT C to assess and score the quality of answers generated by chat gbt or D to perform back propagation in chat gpt's Network comment your answer down below and let's have a discussion and as I mentioned before if you do think I deserve it please do give this video a like that'll mean a lot to me that's going to do for quiz time for now but before we go let's write out a summary re reinforcement learning through human feedback is a framework that integrates human feedback into the training process of a reinforcement learning algorithm now the reinforcement learning algorithm in question could be DQ learning proximal policy optimization or any other algorithm human feedback is used to guide and accelerate the learning process allowing the algorithm to make more informed decisions and in chat gbt human feedback is given via the rewards model the iterative training process with reinforcement learning through human feedback enhances chat gpt's capabilities making it a powerful tool for generating high quality responses and that's going to do it for today so I hope this video helped you get a good sense of what is reinforcement learning through human feedback and where it is also practically used in the guise of chat GPT now here there was a mention of an algorithm called proximal policy optimization so to understand more details about it do check out that video right over here thank you all so much for watching if you think I deserve it please do give this video a like once again and I will see you in the next one bye-bye
Up Next

Writing Basic Foundry Tests: Setup, Assertions, and Error Handling
@smartcontractprogrammer
14.3K views•2023-03-20

Building Real-Time ML Pipelines with Feature Stores and MLOps Frameworks
@ODSCAI
5.1K views•2022-02-20

Proximal Policy Optimization: RL Algorithm Explained | PPO Tutorial
@CodeEmporium
41.2K views•2023-12-04

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence






















![[인공지능,머신러닝,딥러닝] (심화) Direct preference optimization (DPO)](https://i.ytimg.com/vi/A80ue5nS_A4/maxresdefault.jpg)
![[ESC 2025-2 세션] 20250729 Large Language Models](https://i.ytimg.com/vi/hB_ykf5BGaY/maxresdefault.jpg)








![Cassidy Laidlaw - A New Definition & Improved Mitigation for Reward Hacking [Alignment Workshop]](https://i.ytimg.com/vi_webp/s_I-6AJfz58/maxresdefault.webp)





