Astra, Looped, and Partially Blind

GPT-6 Astra's headline numbers are real — a jump from 7.8% to 62.7% on ARC-AGI-3, the first model at OpenAI's Critical cybersecurity tier — but the architecture that gets it there may be the most consequential detail: looped transformers that keep reasoning opaque by design, arriving the same week OpenAI's chief scientist said chain-of-thought monitoring is already becoming unreliable.

Read more →

Same Weights, Different Agent

DeepSeek V4-Flash-0731 went from a 7.3 to a 54.4 on DeepSWE with identical pretrained weights — a 645% jump achieved purely through post-training. It's evidence that the gap between "can write code" and "can act as an agent" is largely a training-envelope problem, not a capacity problem, and that post-training is now a first-class axis of model improvement alongside pretraining scale and architecture.

Read more →