A Ghost in the Reasoning

A researcher prefilled Qwen3.8 A95B with the first 1% of a GPT-5.5 Pro chain-of-thought and measured how much the model continued following it. The resulting +18 pp jump in overlap — much larger than seen in other models — is a practical signal that Qwen's post-training data may include GPT-5.5 Pro reasoning traces, and points to a general method for probing where a model learned to think.

Read more →

What $998 Buys You

Hugo Vergnes trained a 3.8B-parameter LLM for $998 in compute costs, scoring 0.384 on the CORE benchmark. The interesting result isn't the price — it's that doubling context length from 1024 to 2048 tokens drove 83% of the score improvement on one key task, which says something about what benchmark numbers are actually measuring.

Read more →

Five Hundred Dollars of RL

Fermi Sense and Ramp fine-tuned a 9B open-source model with GRPO for $500 and outperformed every frontier configuration on a catalog review task — 87.3% vs 76.9%, at 40x lower inference cost. The benchmark has a Goodhart's Law concern, but the underlying economics of task-specific RL fine-tuning are real and worth taking seriously.

Read more →