Proximal Policy Optimization Algorithms
John Schulman et al. · arXiv preprint
arXiv:1707.06347
In short
PPO is a policy-gradient method that limits how far each update can move the policy by clipping its objective. That lets it reuse each batch of experience for several update passes, stay stable, and remain simple to implement.
Why it matters
PPO became the default RL algorithm, including for RLHF on the first ChatGPT-era models.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 13 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Reinforcement Learning · read first ✓ understood
Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.
- Agent ✓ understood
In RL, the learner or decision-maker that takes actions in an environment to maximize cumulative reward.
- Policy ✓ understood
A strategy or mapping from states to actions that defines the agent's behavior in reinforcement learning.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Policy Gradient · read first ✓ understood
RL methods that directly optimize the policy by computing gradients of expected reward with respect to policy parameters.
- Reward ✓ understood
A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.
- Value Function ✓ understood
A function estimating expected cumulative reward from a state (state-value) or state-action pair (action-value/Q-value).
- Actor-Critic ✓ understood
RL architecture with two components: an actor (policy) that selects actions and a critic (value function) that evaluates them.
- PPO · read first ✓ understood
Proximal Policy Optimization - a stable and efficient policy gradient algorithm widely used in RLHF for training LLMs.