Cloudflare Enters the Decision Model Market

Cloudflare released Clef and Clef-flash, open-weight decision models at 27B and 9B for agent routing and classification, alongside a managed RL fine-tuning service. The release validates the decision model category at enterprise scale, while raising the usual question about whether managed cloud infrastructure beats the increasingly capable self-hosted alternatives.

Read more →

The Model That Edits Its Own Notes

A Meta AI and Allen AI paper proposes Context Language Models: give a model read-write access to its own context as a plain file, let it use Bash tools to edit what it carries forward, and watch it cut FLOPs by up to 59% on long-horizon tasks while scoring higher. The serving side gets a matching trick — Suffix Cache Reuse — that reuses unchanged KV cache states after an edit rather than recomputing from scratch.

Read more →

Compile Once for Your Hardware

Magnitude is a new open-source local inference engine that compiles device-specific kernels on first run rather than shipping precompiled binaries for broad hardware classes. The approach gets it to 92% faster decode on Apple Metal vs. llama.cpp and positions it specifically for agent workloads where decode throughput matters most.

Read more →

The Warden Problem

Matthew Green's analysis of whether sandboxing can contain rogue agents arrives at a structural answer: no, not reliably — but for reasons that have nothing to do with superintelligence. The real problem is compliant agents following whoever gets text in front of them, and monitoring that requires deploying more agents, which have the same problem.

Read more →