Tag
Qwen3-1.7B
Every Qwen3-1.7B story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 36 · 2026
Read the issue →-
TTPO: Test-Time Policy Optimization
The paper introduces TTPO, a method designed to optimize LLM reasoning during test-time training by utilizing asymmetric learning signals from pseudo-labels. It effectively mitigates the risks of noisy majority-vote labels by distilling agreeing rollouts and penalizing disagreeing ones.
Week 35 · 2026
Read the issue →-
TTPO: Test-Time Policy Optimization
The paper introduces TTPO, a method designed to optimize LLM reasoning during test-time training by utilizing asymmetric learning signals from pseudo-labels. It effectively mitigates the risks of noisy majority-vote labels by distilling agreeing rollouts and penalizing disagreeing ones.
Week 29 · 2026
Read the issue →-
J-space comparisons across open models
This research expands on Anthropic's 'Verbalizable-Workspace' paper by applying J-space measurements to open-source models. It characterizes the internal structure of LLMs as having distinct functional zones and demonstrates that steering influence decays via a power law.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free