← All topics

Tag

On-policy Distillation (OPD)

Every On-policy Distillation (OPD) story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

4 items · 2 issues

Week 29 · 2026

Read the issue →

Week 28 · 2026

Read the issue →
  • Weak-to-Strong Generalization via Direct On-Policy Distillation

    This technical paper proposes a new distillation paradigm called Direct-OPD to improve weak-to-strong generalization in reasoning models. By distilling only the policy shift induced by reinforcement learning, the method avoids the capacity limitations inherent in traditional teacher-student imitation.

    AI & ML arXiv Source ↗
  • Weak-to-Strong Generalization via Direct On-Policy Distillation

    This technical paper proposes a new distillation paradigm called Direct-OPD to improve weak-to-strong generalization in reasoning models. By distilling only the policy shift induced by reinforcement learning, the method avoids the capacity limitations inherent in traditional teacher-student imitation.

    AI & ML arXiv Source ↗