Key Info
Fireworks Research released Ember-1, a reasoning model built on Kimi K3 that uses ~40% fewer tokens while maintaining benchmark performance, achieved through post-training to reduce repetitive thinking.
Highlights
- Ember-1 uses ~40% fewer tokens than K3 with the same benchmark performance.
- Trained via RL on real agentic coding task loops to distinguish useful reasoning from looping.
- In a live A/B test on coding traffic, Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens at the same success rate.
- Available in Cline now, including the new Desktop app.