Keeping the Problem Hard

A new post-training method called Frontier Learning dynamically generates reasoning problems calibrated to a model's current capability, using a regret signal to stay at the edge of what the model can reliably solve. On a probability reasoning task it achieves a 115% relative gain over fixed-pool baselines.

Read more →