Fireworks AI Cuts the Token Meter With Ember-1

Fireworks AI released Ember-1, cutting Kimi K3 token volume by up to 50% at identical per-token prices, squeezing foundation labs while restoring enterprise software margins.

Vincent JiangVincent Jiang · 3 min read
Share
Four Nvidia H100 data center GPUs standing in a row against a black background
Nvidia H100 accelerators, the class of hardware serving Kimi K3 and now Ember-1 token streams for inference providers like Fireworks AI.

The landlord gets bypassed

When Moonshot AI launched its 2.8 trillion parameter open-weights model Kimi K3 in July 2026, it shipped one of the sharpest coding engines on earth 4. It also shipped an operational trap. Reasoning models spend up to 90% of their output tokens on internal chain-of-thought scratchpads 1,2. In multi-turn software workflows, every prior turn's internal monologue gets replayed into the prompt, expanding context quadratically and rebilling the enterprise on every subsequent step 1,3.

Enterprises running agentic workflows watched their unit economics buckle. On 26 September 2026, reporting showed legal platform Harvey rebuilt its infrastructure on Kimi K3 precisely to fix gross margins that had swung from positive 50% to negative 50% under proprietary labs 5. But running native K3 remained prohibitively verbose: tasks routinely chewed through tens of thousands of internal tokens, taking nearly an hour per complex run 2.

A retrained K3, not a smaller model

Fireworks AI answered that bottleneck on 28 September 2026 by releasing Ember-1 1,2. Ember-1 is not a pruned small model or an inference parameter tweak; Fireworks took Moonshot's open model and ran it through over 50 post-training experiments to excise redundant loops while preserving self-reflection 1,3.

Compressing the meter

Fireworks did not drop its price per token. Instead, it compressed the number of tokens the customer buys 1,3. Ember-1 matches Kimi K3 rate cards dollar for dollar: $3.00 per million input tokens, $0.30 per million cached tokens, and $15.00 per million output tokens 1,2,4.

Output tokens cost five times input on Ember-1's unchanged rate card, so volume cuts hit the expensive half

$0 per 1M tokens$5 per 1M tokens$10 per 1M tokens$15 per 1M tokensInput$3 per 1M tokensCached input$0.30Output$15 per 1M tokens
Data
Value
Input$3 per 1M tokens
Cached input$0.30
Output$15 per 1M tokens
Ember-1 pricing per million tokens on Fireworks AI, unchanged from Kimi K3's rate card; the savings come from token volume, not price. September 2026.1,2,4

Ember-1 beats Kimi K3 Max on DeepSWE and matches it elsewhere while slashing token counts

  • K3 Low
  • K3 Max
  • Ember-1
0%50%100%Terminal Bench 2.1SWE-bench VerifiedDeepSWE 1.175.2%
Data
K3 LowK3 MaxEmber-1
Terminal Bench 2.176.4%80.9%82%
SWE-bench Verified80.4%93.2%92.2%
DeepSWE 1.155.8%66.4%75.2%
Benchmark accuracy percentages reported by Fireworks AI across Terminal Bench 2.1, SWE-bench Verified, and DeepSWE 1.1; September 2026.1,3

The savings occur entirely by reducing token volume by 35% to 50% 1,3. In live production A/B tests with two enterprise coding customers, output tokens plummeted from 49,300 to 29,900 per task, reasoning tokens dropped 71.3%, and overall token consumption fell 39% 1,3. Accuracy held steady at 0.753 versus 0.751, while task steps dropped from 23.8 to 21.4 1,3. On deep engineering benchmarks, Ember-1 actually outperformed Moonshot's flagship K3 Max run 1,3.

Who pays for the missing tokens

The tension in AI economics has migrated from raw price wars to billing-unit shrinkage. Model creators like Moonshot pour millions of dollars into foundation training capex, hoping to recoup investments via high-volume token consumption 4,6. Cloud giants have rushed to capture those flows: Amazon Web Services recently brought Kimi K3 to Bedrock while Moonshot negotiated revenue-sharing terms of up to 30% with hyperscalers 6.

Fireworks breaks that monetization model by sitting at the serving layer, keeping the weights proprietary in serverless preview, and cutting the billable token stream in half 1,3. When an inference broker can post-train an open-source model to think in fewer words, the foundation lab absorbs the capex while the host captures the enterprise margin. Enterprise buyers recover sustainable margins without touching rate cards, while labs discover that when tokens become concise, their revenue burns at both ends.

How this brief was made

01Gathered & sourced383 channels · 2,396 articles▾

Agents swept 383 channels and ingested 2,396 articles, then de-duplicated and ranked them for signal.

02Verified & cross-validated6 claims · 19 data feeds▾
03Reviewed & edited2 human editors▾

2 editors read the draft against the evidence, tuned the framing, and signed off before it shipped.

Become a contributor

Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.

Share

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio