Keeping the Problem Hard
A new post-training method called Frontier Learning dynamically generates reasoning problems calibrated to a model's current capability, using a regret signal to stay at the edge of what the model can reliably solve. On a probability reasoning task it achieves a 115% relative gain over fixed-pool baselines.
Read more →
