When the Proof Arrives Alone

Over 600 mathematicians responded to a survey on how AI labs should release AI-generated mathematical results. The resulting statement from AGMAI distinguishes proofs humans understand from those they don't — and places responsibility on labs to fund the gap.

Read more →

miniF2F Hits the Ceiling

Mistral's Leanstral 1.5 scores 100% on miniF2F and solves 587 of 672 Putnam Competition problems using a 6B-active-parameter MoE. The model saturates the main formal-proof benchmark and finds real bugs in production code — at roughly $4 per Putnam problem versus competitors charging $300.

Read more →