← All topics

Tag

GRPO

Every GRPO story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

6 items · 6 issues

Beats AI & ML

Week 36 · 2026

Read the issue →

Week 35 · 2026

Read the issue →

Week 33 · 2026

Read the issue →
  • On-Policy Self-Distillation without Any Supervision

    Researchers have developed u-OPSD, a technique that allows large language models to perform self-distillation using only their own generated outputs. By leveraging internal consistency through majority voting, the model can correct its own errors without requiring external ground-truth data.

    AI & ML HuggingFace Papers Source ↗

Week 32 · 2026

Read the issue →

Week 29 · 2026

Read the issue →
  • Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

    This paper presents the development of Ring-2.5-1T-Zero, a trillion-parameter model trained via zero-shot reinforcement learning to elicit emergent reasoning. The study validates that massive scaling enables models to spontaneously develop complex problem-solving strategies like self-verification without human-annotated data.

    AI & ML HuggingFace Papers Source ↗

Week 27 · 2026

Read the issue →