A Ghost in the Reasoning

A researcher prefilled Qwen3.8 A95B with the first 1% of a GPT-5.5 Pro chain-of-thought and measured how much the model continued following it. The resulting +18 pp jump in overlap — much larger than seen in other models — is a practical signal that Qwen's post-training data may include GPT-5.5 Pro reasoning traces, and points to a general method for probing where a model learned to think.

Read more →