Agents & RL Jul 2017

Proximal Policy Optimization Algorithms

John Schulman et al. · arXiv preprint

arXiv:1707.06347

In short

PPO is a policy-gradient method that limits how far each update can move the policy by clipping its objective. That lets it reuse each batch of experience for several update passes, stay stable, and remain simple to implement.

Why it matters

PPO became the default RL algorithm, including for RLHF on the first ChatGPT-era models.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 13 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Reinforcement Learning · read first ✓ understood

    Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.

  3. Agent ✓ understood

    In RL, the learner or decision-maker that takes actions in an environment to maximize cumulative reward.

  4. Policy ✓ understood

    A strategy or mapping from states to actions that defines the agent's behavior in reinforcement learning.

  5. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  6. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  7. Loss Function ✓ understood

    A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

  8. Gradient Descent ✓ understood

    An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.

  9. Policy Gradient · read first ✓ understood

    RL methods that directly optimize the policy by computing gradients of expected reward with respect to policy parameters.

  10. Reward ✓ understood

    A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.

  11. Value Function ✓ understood

    A function estimating expected cumulative reward from a state (state-value) or state-action pair (action-value/Q-value).

  12. Actor-Critic ✓ understood

    RL architecture with two components: an actor (policy) that selects actions and a critic (value function) that evaluates them.

  13. PPO · read first ✓ understood

    Proximal Policy Optimization - a stable and efficient policy gradient algorithm widely used in RLHF for training LLMs.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers