Tag
CUDA
Every CUDA story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 39 · 2026
Read the issue →-
NVIDIA links TensorRT and Dynamo-Triton for faster AI on multiple GPUs
NVIDIA's new integration between TensorRT and Dynamo-Triton allows a single neural network to execute across multiple GPUs seamlessly. This advancement significantly reduces latency for complex generative AI tasks like video synthesis by automating GPU coordination.
-
CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
CliffCompaction is a new autocompaction strategy for long-horizon coding agents that manages context windows by truncating token-intensive content without rephrasing. This approach significantly reduces inference costs and enables efficient test-time scaling while preventing the accumulation of context drift.
-
SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
This technical update outlines a series of low-level optimizations for production-grade AI inference serving. It covers significant advancements in speculative decoding, MoE model support, and GPU kernel efficiency.
Week 28 · 2026
Read the issue →-
Why a five-minute sniff test is your secret supply chain defense
The article advocates for a proactive 'sniff test' methodology to validate the integrity of SBOMs in containerized environments. It highlights how identifying omissions like unpinned packages or missing dependencies is crucial for preventing supply chain attacks.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free