The Millennium Problem and the Credit Dispute

OpenAI's agents proved that 3D Navier-Stokes equations can develop singularities — a genuine mathematical achievement. But the credit dispute that followed, involving an NYU mathematician and an Anthropic employee who'd been working the same problem for a year, raised a question that has no clean answer: what happens to independent research when an AI lab can decide to race at it the moment they hear a rumor?

Read more →

Science Is Harder Than Coding

Terminal-Bench-Science v0.1 evaluates AI agents on 70 real computational workflows from active research labs across five scientific domains. Claude Opus 5 tops the leaderboard at 30%, well below the 50–80% range frontier models achieve on general coding benchmarks — a gap that says something useful about where the actual difficulty lies in scientific work.

Read more →

After AlphaFold, Jumper Places a New Bet

John Jumper, who led AlphaFold and won the 2024 Nobel Prize in Chemistry, is leaving Google DeepMind for Anthropic. The interesting question isn't who won the talent war — it's what his choice says about where the hard problems in biology AI go next, and why a safety-focused lab might actually be the right place to work on them.

Read more →

Claude Passes an NMR Exam

Anthropic published a study showing Opus 4.7 matching or beating ChemDraw and MestReNova on 1D NMR spectroscopy tasks. The 80% J-coupling spacing accuracy — versus 26–35% for dedicated software — is the surprising number. The bidirectional structure elucidation capability has no direct equivalent in existing tools.

Read more →