← All topics

Tag

Qwen3-1.7B

Every Qwen3-1.7B story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

3 items · 3 issues

Beats AI & ML

Week 36 · 2026

Read the issue →
  • TTPO: Test-Time Policy Optimization

    The paper introduces TTPO, a method designed to optimize LLM reasoning during test-time training by utilizing asymmetric learning signals from pseudo-labels. It effectively mitigates the risks of noisy majority-vote labels by distilling agreeing rollouts and penalizing disagreeing ones.

    AI & ML arXiv Source ↗

Week 35 · 2026

Read the issue →
  • TTPO: Test-Time Policy Optimization

    The paper introduces TTPO, a method designed to optimize LLM reasoning during test-time training by utilizing asymmetric learning signals from pseudo-labels. It effectively mitigates the risks of noisy majority-vote labels by distilling agreeing rollouts and penalizing disagreeing ones.

    AI & ML arXiv Source ↗

Week 29 · 2026

Read the issue →
  • J-space comparisons across open models

    This research expands on Anthropic's 'Verbalizable-Workspace' paper by applying J-space measurements to open-source models. It characterizes the internal structure of LLMs as having distinct functional zones and demonstrates that steering influence decays via a power law.

    AI & ML HackerNews Source ↗