Tag
Transformer
Every Transformer story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 28 · 2026
Read the issue →-
The Key to Going Linear: Analysis-Driven Transformer Linearization
This paper presents a method for converting pretrained transformers into linear-time architectures by focusing on the efficiency of state update designs. By analyzing softmax attention through a first-order approximation, the authors prove that delta-style updates are superior for post hoc linearization.
-
The Key to Going Linear: Analysis-Driven Transformer Linearization
This paper presents a method for converting pretrained transformers into linear-time architectures by focusing on the efficiency of state update designs. By analyzing softmax attention through a first-order approximation, the authors prove that delta-style updates are superior for post hoc linearization.
Week 27 · 2026
Read the issue →-
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
This study demonstrates that reinforcement learning post-training for LLMs is not a uniform process across all parameters but is concentrated in specific middle layers. By leveraging this discovery, researchers developed layer-aware training methods that outperform traditional full-parameter optimization.
-
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
This study demonstrates that reinforcement learning post-training for LLMs is not a uniform process across all parameters but is concentrated in specific middle layers. By leveraging this discovery, researchers developed layer-aware training methods that outperform traditional full-parameter optimization.
Week 26 · 2026
Read the issue →-
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Wan-Streamer is a new foundation model that enables seamless, low-latency audio-visual interaction by processing interleaved text, audio, and video tokens within a single causal Transformer. By moving away from cascaded modular pipelines, it achieves sub-second end-to-end latency suitable for real-time digital humans and embodied assistants.
-
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Wan-Streamer is a new foundation model that enables seamless, low-latency audio-visual interaction by processing interleaved text, audio, and video tokens within a single causal Transformer. By moving away from cascaded modular pipelines, it achieves sub-second end-to-end latency suitable for real-time digital humans and embodied assistants.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free