Tag
Large Language Models (LLMs)
Every Large Language Models (LLMs) story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 30 · 2026
Read the issue →-
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
PagedWeight optimizes the serving of MoE LLMs by implementing a dynamic quantization strategy that adapts to runtime memory pressure. It balances hardware efficiency with model accuracy by monitoring expert routing statistics and prompt-specific sensitivities.
-
On-Policy Delta Distillation
This research presents OPD2, a novel distillation technique that focuses on the learning trajectory of reasoning models by using a delta signal between teacher and base models. Experimental results prove it outperforms standard on-policy distillation across multiple complex reasoning domains.
-
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
This research presents a novel approach to document retrieval that prioritizes the holistic quality of document sets over individual document relevance. By introducing the SetwiseEvalKit benchmark and the Rubric4Setwise optimization method, the authors demonstrate how addressing redundancy and factual conflicts can significantly improve LLM generation performance.
-
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
PagedWeight optimizes the serving of MoE LLMs by implementing a dynamic quantization strategy that adapts to runtime memory pressure. It balances hardware efficiency with model accuracy by monitoring expert routing statistics and prompt-specific sensitivities.
-
On-Policy Delta Distillation
This research presents OPD2, a novel distillation technique that focuses on the learning trajectory of reasoning models by using a delta signal between teacher and base models. Experimental results prove it outperforms standard on-policy distillation across multiple complex reasoning domains.
-
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
This research presents a novel approach to document retrieval that prioritizes the holistic quality of document sets over individual document relevance. By introducing the SetwiseEvalKit benchmark and the Rubric4Setwise optimization method, the authors demonstrate how addressing redundancy and factual conflicts can significantly improve LLM generation performance.
Week 29 · 2026
Read the issue →-
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
The Seed framework addresses the limitations of sparse rewards in agentic reinforcement learning by extracting reusable skills from completed interaction trajectories. It utilizes on-policy distillation to provide dense, token-level supervision that evolves alongside the policy's capabilities.
-
Prompt-engineering paper accepted to ICML [R]
This research identifies that mode collapse in LLMs is driven by a cognitive typicality bias within preference datasets. To counter this, the authors present Verbalized Sampling, an inference-time method that unlocks model diversity without retraining.
-
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
The Seed framework addresses the limitations of sparse rewards in agentic reinforcement learning by extracting reusable skills from completed interaction trajectories. It utilizes on-policy distillation to provide dense, token-level supervision that evolves alongside the policy's capabilities.
Week 28 · 2026
Read the issue →-
Weak-to-Strong Generalization via Direct On-Policy Distillation
This technical paper proposes a new distillation paradigm called Direct-OPD to improve weak-to-strong generalization in reasoning models. By distilling only the policy shift induced by reinforcement learning, the method avoids the capacity limitations inherent in traditional teacher-student imitation.
-
Weak-to-Strong Generalization via Direct On-Policy Distillation
This technical paper proposes a new distillation paradigm called Direct-OPD to improve weak-to-strong generalization in reasoning models. By distilling only the policy shift induced by reinforcement learning, the method avoids the capacity limitations inherent in traditional teacher-student imitation.
Week 27 · 2026
Read the issue →-
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
This study demonstrates that reinforcement learning post-training for LLMs is not a uniform process across all parameters but is concentrated in specific middle layers. By leveraging this discovery, researchers developed layer-aware training methods that outperform traditional full-parameter optimization.
-
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
This paper presents Distill to Detect (D2D), a technique for identifying stealthy, topic-specific biases in language models that evade standard detection. By distilling the distributional shift of a suspected model into a small prefix adapter, the method amplifies hidden signals until they become visible in generated text.
-
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
This study demonstrates that reinforcement learning post-training for LLMs is not a uniform process across all parameters but is concentrated in specific middle layers. By leveraging this discovery, researchers developed layer-aware training methods that outperform traditional full-parameter optimization.
-
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
This paper presents Distill to Detect (D2D), a technique for identifying stealthy, topic-specific biases in language models that evade standard detection. By distilling the distributional shift of a suspected model into a small prefix adapter, the method amplifies hidden signals until they become visible in generated text.
Week 25 · 2026
Read the issue →-
APPO: Agentic Procedural Policy Optimization
This paper proposes APPO, a method to enhance the training of autonomous LLM agents through more granular credit assignment. It moves beyond trajectory-level rewards by identifying and branching at specific high-impact decision points within the reasoning process.
Week 22 · 2026
Read the issue →-
Scientists trained an AI model using an IBM quantum computer — and it answered questions correctly that the base model couldn't
Scientists successfully demonstrated quantum enhancement in large language models by creating a hybrid system that integrates quantum circuit blocks. This novel approach significantly improved the LLM's perplexity and factual accuracy, paving the way for more powerful, resource-efficient AI.
Week 21 · 2026
Read the issue →-
Barnes & Noble CEO backs selling AI-written books in stores
Barnes & Noble CEO James Daunt announced that the company is willing to sell AI-written books in its stores, provided that the books are transparently labeled as synthetic content. He stressed that the key criterion is maintaining clarity for the customer, ensuring the book does not falsely imitate human authorship.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free