Tag
Qwen3-8B
Every Qwen3-8B story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 37 · 2026
Read the issue →-
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
FlowBalance is a novel reinforcement learning method designed to stabilize the self-improvement process in large reasoning models. It effectively integrates sparse verifier feedback with dense, privileged-hindsight self-guidance to prevent model overconfidence and improve mathematical reasoning accuracy.
-
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
AgentGrad is a novel prompt optimization framework for multi-agent systems that improves both the accuracy and efficiency of textual gradient methods. By using sequential intervention to target specific agents and semantic abstraction to group similar errors, it outperforms current state-of-the-art approaches.
Week 36 · 2026
Read the issue →-
TTPO: Test-Time Policy Optimization
The paper introduces TTPO, a method designed to optimize LLM reasoning during test-time training by utilizing asymmetric learning signals from pseudo-labels. It effectively mitigates the risks of noisy majority-vote labels by distilling agreeing rollouts and penalizing disagreeing ones.
Week 35 · 2026
Read the issue →-
TTPO: Test-Time Policy Optimization
The paper introduces TTPO, a method designed to optimize LLM reasoning during test-time training by utilizing asymmetric learning signals from pseudo-labels. It effectively mitigates the risks of noisy majority-vote labels by distilling agreeing rollouts and penalizing disagreeing ones.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free