← All topics

Tag

Reinforcement Learning (RL)

Every Reinforcement Learning (RL) story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

4 items · 2 issues

Week 30 · 2026

Read the issue →
  • On-Policy Delta Distillation

    This research presents OPD2, a novel distillation technique that focuses on the learning trajectory of reasoning models by using a delta signal between teacher and base models. Experimental results prove it outperforms standard on-policy distillation across multiple complex reasoning domains.

    AI & ML HuggingFace Papers Source ↗
  • On-Policy Delta Distillation

    This research presents OPD2, a novel distillation technique that focuses on the learning trajectory of reasoning models by using a delta signal between teacher and base models. Experimental results prove it outperforms standard on-policy distillation across multiple complex reasoning domains.

    AI & ML HuggingFace Papers Source ↗

Week 29 · 2026

Read the issue →