Training Feb 2026 · #63 most cited · 87 citations

Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Wenkai Yang et al.

arXiv:2602.12125

In short

The authors show on-policy distillation is a special case of KL-regularised RL, then generalise it with a tunable reward weight. Weighting the reward above 1 (“reward extrapolation”) lets students beat standard distillation and, when merging domain experts, even surpass their teachers.

Why it matters

It turns a popular heuristic into a framework with a knob that measurably helps.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 10 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  5. Activation Function ✓ understood

    A non-linear function applied to neuron outputs that introduces non-linearity, enabling networks to learn complex patterns.

  6. Softmax ✓ understood

    A function that turns a list of scores (logits) into probabilities that are all positive and sum to 1; the standard output of classifiers and language models.

  7. Knowledge Distillation · read first ✓ understood

    Training a smaller 'student' model to mimic a larger 'teacher' model, transferring knowledge while reducing size.

  8. Reinforcement Learning · read first ✓ understood

    Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.

  9. Entropy ✓ understood

    A measure of uncertainty or randomness in a random variable from information theory.

  10. KL Divergence · read first ✓ understood

    Kullback-Leibler divergence - a measure of how one probability distribution differs from another.

In the frontier

Rank
#63 of 100
Citations
87
as of Aug 9, 2026
Published
Feb 2026

Topics: RL for reasoning , Efficiency and serving , Reasoning methods

Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026

Nearby papers

Summary in our own words; read the paper for the details. ← All papers