LLM Preference Tuning Explained: RLHF, PPO, DPO | Stanford CME295

Added:

Preference Tuning
Data Collection
RL Basics
Reward Modeling
PPO Alignment
Best-of-N Sampling
Direct Preference Opt
Method Comparison

Preference Tuning

4:11
Playing Section
  • 1

    Explore the necessity of aligning models with human preferences.

  • 2

    Identify the limitations of supervised fine-tuning in this context.

  • 3

    Explain the role of preference pairs in this process.

Foundational understanding of Transformer architectures and standard autoregressive language model training (pre-training vs. fine-tuning).
Basic principles of Reinforcement Learning (RL), including concepts like policies, states, actions, rewards, and value functions.
Familiarity with Supervised Fine-Tuning (SFT) and how base models are initially adapted to follow instructions.
Understanding of optimization and gradient-based training, particularly policy gradient methods in machine learning.
Advanced alternative alignment methodologies, such as Kahneman-Tversky Optimization (KTO) and Odds Ratio Preference Optimization (ORPO).
Strategies to mitigate 'reward hacking', policy drift, and the alignment tax (the loss of general capabilities during tuning).
Reinforcement Learning from AI Feedback (RLAIF) and Constitutional AI, exploring how to scale alignment using synthetic feedback.
Hands-on implementation of preference tuning using frameworks like Hugging Face's TRL (Transformer Reinforcement Learning) or DeepSpeed-Chat.
22.2K views385likes1:47:41@stanfordonlineOriginal Release: 2025-11-14

This lecture covers preference tuning and reinforcement learning from human feedback (RLHF) for aligning large language models with human preferences. The process involves collecting pairwise preference data where humans rate which model response is better, then training a reward model using the Bradley-Terry formulation to score responses. The second stage uses reinforcement learning algorithms like Proximal Policy Optimization (PPO) with clipping and KL-divergence penalties to align the model with these preferences without deviating too far from the base model. An alternative approach called Direct Preference Optimization (DPO) provides a simpler supervised method that directly optimizes model weights using preference pairs, avoiding the complexity of maintaining separate reward and value functions.