A Thousand Tokens per Second

InceptionLabs released Mercury 2.5, a diffusion language model doing 1,107 tokens per second on standard NVIDIA GPUs at $0.20/$0.75 per million tokens. For voice agents and real-time coding assistants, the latency difference between autoregressive and diffusion approaches has become concrete and measurable.

Read more →

Five Knobs, Sub-50ms

Nari Labs walked through five coordinated optimizations that bring Qwen3-TTS 1.7B to sub-50 ms p95 time-to-first-audio on a single H100, at $2 per million characters — against ElevenLabs at $100/M. None of the five changes require a new model architecture. Each targets a specific latency source, and the gains compound.

Read more →

Speculative Decoding Has an Acceptance Problem You Can Exploit

Mistletoe (arXiv 2605.14005) demonstrates a stealthy adversarial attack on speculative decoding systems: craft inputs that look normal to the target model but cause the draft model to disagree, collapsing acceptance length and throughput while leaving output quality and perplexity unchanged. The attack exploits the fundamental gap between draft and target distributions that all speculative systems rely on bridging.

Read more →