← All topics

Tag

Reinforcement Learning

Every Reinforcement Learning story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

3 items · 2 issues

Week 34 · 2026

Read the issue →

Week 28 · 2026

Read the issue →
  • Weak-to-Strong Generalization via Direct On-Policy Distillation

    This technical paper proposes a new distillation paradigm called Direct-OPD to improve weak-to-strong generalization in reasoning models. By distilling only the policy shift induced by reinforcement learning, the method avoids the capacity limitations inherent in traditional teacher-student imitation.

    AI & ML arXiv Source ↗