Fireworks AI launches Ember-1: 40% fewer tokens, same quality

Original: Ember-1

Why This Matters

Token efficiency is now a primary cost lever for AI at scale — especially in agentic, multi-turn deployments.

Fireworks AI released Ember-1 on September 23, 2026 — a specialized model built on Kimi K3 that cuts token usage by 40% while maintaining output quality. The model required 50+ training runs and 200+ evaluations, and is available today via Fireworks Serverless Training.

Fireworks Research's Ember-1 tackles a real cost problem: reasoning models like Kimi K3 burn the vast majority of their tokens — sometimes over 90% — on internal chain-of-thought rather than actual answers. In multi-turn agentic workflows, this compounds fast, since prior reasoning gets re-read on every call, making context grow roughly quadratically with conversation length.

Rather than simply dialing down K3's reasoning effort (which the team found hurt quality too much), Fireworks trained Ember-1 to reason more efficiently from the ground up. The process involved 50+ training experiments and 200+ evaluations, with new training algorithms developed specifically to shorten reasoning traces without degrading accuracy. The team ran everything on Fireworks Serverless Training, skipping GPU provisioning overhead.

Ember-1 was validated on the company's Specialized Intelligence Index, public benchmarks, live customer A/B tests, and internal coding and agent workloads. According to Fireworks, quality held across all settings. No customer data was used in training.

The model is positioned as the first in a planned series from Fireworks Research. Fireworks also announced its inaugural conference, Forge 2026, alongside the launch.

Source

fireworks.ai — Read original →