Fireworks AI Cuts the Token Meter With Ember-1
Fireworks AI released Ember-1, cutting Kimi K3 token volume by up to 50% at identical per-token prices, squeezing foundation labs while restoring enterprise software margins.
Vincent Jiang · 3 min read
The landlord gets bypassed
When Moonshot AI launched its 2.8 trillion parameter open-weights model Kimi K3 in July 2026, it shipped one of the sharpest coding engines on earth 4. It also shipped an operational trap. Reasoning models spend up to 90% of their output tokens on internal chain-of-thought scratchpads 1,2. In multi-turn software workflows, every prior turn's internal monologue gets replayed into the prompt, expanding context quadratically and rebilling the enterprise on every subsequent step 1,3.
Enterprises running agentic workflows watched their unit economics buckle. On 26 September 2026, reporting showed legal platform Harvey rebuilt its infrastructure on Kimi K3 precisely to fix gross margins that had swung from positive 50% to negative 50% under proprietary labs 5. But running native K3 remained prohibitively verbose: tasks routinely chewed through tens of thousands of internal tokens, taking nearly an hour per complex run 2.
A retrained K3, not a smaller model
Fireworks AI answered that bottleneck on 28 September 2026 by releasing Ember-1 1,2. Ember-1 is not a pruned small model or an inference parameter tweak; Fireworks took Moonshot's open model and ran it through over 50 post-training experiments to excise redundant loops while preserving self-reflection 1,3.
Compressing the meter
Fireworks did not drop its price per token. Instead, it compressed the number of tokens the customer buys 1,3. Ember-1 matches Kimi K3 rate cards dollar for dollar: $3.00 per million input tokens, $0.30 per million cached tokens, and $15.00 per million output tokens 1,2,4.
Output tokens cost five times input on Ember-1's unchanged rate card, so volume cuts hit the expensive half
Data
| Value | |
|---|---|
| Input | $3 per 1M tokens |
| Cached input | $0.30 |
| Output | $15 per 1M tokens |
Ember-1 beats Kimi K3 Max on DeepSWE and matches it elsewhere while slashing token counts
- K3 Low
- K3 Max
- Ember-1
Data
| K3 Low | K3 Max | Ember-1 | |
|---|---|---|---|
| Terminal Bench 2.1 | 76.4% | 80.9% | 82% |
| SWE-bench Verified | 80.4% | 93.2% | 92.2% |
| DeepSWE 1.1 | 55.8% | 66.4% | 75.2% |
The savings occur entirely by reducing token volume by 35% to 50% 1,3. In live production A/B tests with two enterprise coding customers, output tokens plummeted from 49,300 to 29,900 per task, reasoning tokens dropped 71.3%, and overall token consumption fell 39% 1,3. Accuracy held steady at 0.753 versus 0.751, while task steps dropped from 23.8 to 21.4 1,3. On deep engineering benchmarks, Ember-1 actually outperformed Moonshot's flagship K3 Max run 1,3.
Who pays for the missing tokens
The tension in AI economics has migrated from raw price wars to billing-unit shrinkage. Model creators like Moonshot pour millions of dollars into foundation training capex, hoping to recoup investments via high-volume token consumption 4,6. Cloud giants have rushed to capture those flows: Amazon Web Services recently brought Kimi K3 to Bedrock while Moonshot negotiated revenue-sharing terms of up to 30% with hyperscalers 6.
Fireworks breaks that monetization model by sitting at the serving layer, keeping the weights proprietary in serverless preview, and cutting the billable token stream in half 1,3. When an inference broker can post-train an open-source model to think in fewer words, the foundation lab absorbs the capex while the host captures the enterprise margin. Enterprise buyers recover sustainable margins without touching rate cards, while labs discover that when tokens become concise, their revenue burns at both ends.
How this brief was made
01Gathered & sourced383 channels · 2,396 articles▾
Agents swept 383 channels and ingested 2,396 articles, then de-duplicated and ranked them for signal.
02Verified & cross-validated6 claims · 19 data feeds▾
Every one of 6 load-bearing claims was checked against primary sources, with 19 live data feeds reconciling the figures and charts.
- 1MarkTechPost, Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens, 28 September 2026
- 2GIGAZINE, To address the problem of overthinking and increased costs in cutting-edge AI models, the AI model Ember-1, developed based on Kimi K3, has been introduced, 28 September 2026
- 3Fireworks AI Blog, Introducing Ember-1, 23 September 2026
- 4Crypto Briefing, Modal, Fireworks, and Baseten gain cost advantage with Nvidia and AMD chips for Kimi K3, 20 September 2026
- 5The Inference, Harvey Rebuilds on Kimi K3 as Legal AI Margins Swing, 26 September 2026
- 6Anue Cnyes, Kimi K3 Officially Enters AWS Bedrock, Reuters Reported Talks for Up to 30% Revenue Share, 21 September 2026
03Reviewed & edited2 human editors▾
2 editors read the draft against the evidence, tuned the framing, and signed off before it shipped.
Become a contributor
Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.
Deepdive
AI-generated from this story and its cited sources. Not investment advice.


