Tag
KV-cache
Every KV-cache story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 30 · 2026
Read the issue →-
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
PagedWeight optimizes the serving of MoE LLMs by implementing a dynamic quantization strategy that adapts to runtime memory pressure. It balances hardware efficiency with model accuracy by monitoring expert routing statistics and prompt-specific sensitivities.
-
Self Gradient Forcing: Native Long Video Extrapolation
The researchers present Self Gradient Forcing (SGF) to solve the lack of gradient flow in historical KV caches during autoregressive video generation. This method enables much more stable and consistent long-form video extrapolation without the massive memory overhead of full backpropagation.
-
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
PagedWeight optimizes the serving of MoE LLMs by implementing a dynamic quantization strategy that adapts to runtime memory pressure. It balances hardware efficiency with model accuracy by monitoring expert routing statistics and prompt-specific sensitivities.
-
Self Gradient Forcing: Native Long Video Extrapolation
The researchers present Self Gradient Forcing (SGF) to solve the lack of gradient flow in historical KV caches during autoregressive video generation. This method enables much more stable and consistent long-form video extrapolation without the massive memory overhead of full backpropagation.
Week 27 · 2026
Read the issue →-
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving
ELDR optimizes MoE model serving in disaggregated environments by routing requests with similar expert activation patterns to the same decode workers. This approach reduces memory bandwidth bottlenecks and significantly improves decoding latency.
-
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving
ELDR optimizes MoE model serving in disaggregated environments by routing requests with similar expert activation patterns to the same decode workers. This approach reduces memory bandwidth bottlenecks and significantly improves decoding latency.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free