Key Info

Fireworks Research released Ember-1, a reasoning model built on Kimi K3 that uses ~40% fewer tokens while maintaining benchmark performance, achieved through post-training to reduce repetitive thinking.

Highlights

  • Ember-1 uses ~40% fewer tokens than K3 with the same benchmark performance.
  • Trained via RL on real agentic coding task loops to distinguish useful reasoning from looping.
  • In a live A/B test on coding traffic, Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens at the same success rate.
  • Available in Cline now, including the new Desktop app.