Tag
GRPO
Every GRPO story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 36 · 2026
Read the issue →-
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
Researchers have developed SecOPD, a fine-tuning technique that uses token-level feedback to protect AI agents from prompt injection attacks. This method significantly outperforms previous state-of-the-art defenses by precisely identifying and penalizing malicious tokens during training.
Week 35 · 2026
Read the issue →-
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
Researchers have developed SecOPD, a fine-tuning technique that uses token-level feedback to protect AI agents from prompt injection attacks. This method significantly outperforms previous state-of-the-art defenses by precisely identifying and penalizing malicious tokens during training.
Week 33 · 2026
Read the issue →-
On-Policy Self-Distillation without Any Supervision
Researchers have developed u-OPSD, a technique that allows large language models to perform self-distillation using only their own generated outputs. By leveraging internal consistency through majority voting, the model can correct its own errors without requiring external ground-truth data.
Week 32 · 2026
Read the issue →-
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Researchers have developed ABSeeker, a search agent trained using a new fine-grained credit assignment method called ABC. This approach allows models to learn from specific useful steps within a trajectory rather than just the final outcome, enabling small models to rival much larger counterparts.
Week 29 · 2026
Read the issue →-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
This paper presents the development of Ring-2.5-1T-Zero, a trillion-parameter model trained via zero-shot reinforcement learning to elicit emergent reasoning. The study validates that massive scaling enables models to spontaneously develop complex problem-solving strategies like self-verification without human-annotated data.
Week 27 · 2026
Read the issue →-
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
This study demonstrates that reinforcement learning post-training for LLMs is not a uniform process across all parameters but is concentrated in specific middle layers. By leveraging this discovery, researchers developed layer-aware training methods that outperform traditional full-parameter optimization.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free