← All topics

Tag

On-Policy Distillation (OPD)

Every On-Policy Distillation (OPD) story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

3 items · 3 issues

Beats AI & ML

Week 31 · 2026

Read the issue →
  • Pass the Baton: Trajectory-Relayed On-Policy Distillation

    The researchers present Relay-OPD, a method designed to fix the issue of 'prefix failure' in large language model distillation. By allowing a teacher model to briefly intervene when it detects a reasoning deviation, the system improves student accuracy on mathematical benchmarks with minimal computational overhead.

    AI & ML HuggingFace Papers Source ↗

Week 29 · 2026

Read the issue →

Week 28 · 2026

Read the issue →
  • Weak-to-Strong Generalization via Direct On-Policy Distillation

    This technical paper proposes a new distillation paradigm called Direct-OPD to improve weak-to-strong generalization in reasoning models. By distilling only the policy shift induced by reinforcement learning, the method avoids the capacity limitations inherent in traditional teacher-student imitation.

    AI & ML arXiv Source ↗