Week 36 · 2026
Read the issue →-
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
This research highlights the critical failure of using black-box LLMs as evaluators due to inherent non-determinism in shared computing environments. It advocates for a rigorous, preregistered auditing framework to ensure measurement stability and reproducibility in AI evaluation.
-
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
This technical report examines the efficacy of using Random Attention for KV cache eviction to optimize LLM inference throughput. The study shows that this method significantly boosts tokens per second in high-concurrency scenarios without sacrificing model accuracy.
Week 33 · 2026
Read the issue →-
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
OasisKV addresses the memory wall in LLM inference by implementing a sparse prefetching mechanism for the KV cache. By leveraging speculative decoding to predict future token importance, it moves less critical KV data to cheaper memory tiers without stalling the decode process.
Week 32 · 2026
Read the issue →-
TokTier: Exact Stateful Tokenization for Agentic LLM Serving
TokTier is a novel stateful tokenization service that optimizes LLM inference for agentic workloads by avoiding redundant full-text re-tokenization. It employs incremental repair for session updates and GPU-accelerated processing for new contexts, significantly reducing time to first token.
Week 27 · 2026
Read the issue →-
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving
ELDR optimizes MoE model serving in disaggregated environments by routing requests with similar expert activation patterns to the same decode workers. This approach reduces memory bandwidth bottlenecks and significantly improves decoding latency.
Week 22 · 2026
Read the issue →-
What scanners are actually trying against AI infrastructure
This report details the rising trend of opportunistic scanning targeting AI-related services and infrastructure. It highlights specific threats to unauthenticated Ollama instances and the use of coordinated sweeps to harvest AI API keys from configuration files.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free