Training the Thinking Short
Fireworks Research released Ember-1, a post-trained variant of Kimi K3 that produces roughly 40% fewer reasoning tokens while holding quality steady across benchmarks and live A/B tests. The work points at a new cost axis in the reasoning model era: not smaller models, not quantization, but training the internal monologue to be shorter.
Read more →
