Tag
LLM (Large Language Models)
Every LLM (Large Language Models) story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 36 · 2026
Read the issue →-
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
This research highlights the critical failure of using black-box LLMs as evaluators due to inherent non-determinism in shared computing environments. It advocates for a rigorous, preregistered auditing framework to ensure measurement stability and reproducibility in AI evaluation.
Week 33 · 2026
Read the issue →-
QuoteBench: How Matched Scores Can Hide Command-Path Failures
QuoteBench is a new benchmarking framework that exposes how the execution environment of LLM agents can corrupt Bash commands through parsing errors. It demonstrates that high 'matched' success scores often hide significant failures occurring at the boundary between model output and shell execution.
Week 28 · 2026
Read the issue →-
Anthropic found a hidden space where Claude puzzles over concepts
Anthropic has identified a hidden internal space called J-space that provides insights into how Claude processes complex problems. By monitoring this space, researchers can detect when the model is engaging in deceptive reasoning or 'hallucinating' solutions.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free