Two brass hourglasses in warm sidelighting — left filled with fine golden sand, right with sparse silver granules, both at identical fill level

Fireworks Trained a Model to Stop Overthinking. The Quality Improved.

/ Maxim Starkweather / 7 min read

Last week, Fireworks AI released Ember-1, a model trained to reason less. Not on simpler problems — the same benchmarks, the same engineering tasks, the same production workloads. In A/B tests with enterprise customers, it delivered 35% fewer total tokens per completed task. Reasoning tokens specifically fell 71.3%. The developers switching from Kimi K3 to Ember-1 didn’t notice the change. On Terminal Bench 2.1, Ember-1 scored 82.0% against Kimi K3’s 80.9%.

Tighter reasoning produced a better answer. Two independent data points this week — one from Fireworks, one from Anthropic’s own cost-efficiency documentation — argue that frontier models have been generating significantly more reasoning than their answers require, and that the correction has started. The number Fireworks measured in production: 71% of reasoning tokens were optional.

What Fireworks Actually Found

Fireworks built Ember-1 from the publicly available Kimi K3 weights through what their announcement describes as more than 50 training experiments and over 200 evaluations. The stated goal was to teach the model to “cut unnecessary reasoning while keeping the thinking that matters.” The specific algorithm isn’t disclosed — their blog uses phrases like “on-policy planning and learning” without mathematical specifics, a limitation that matters for anyone trying to evaluate whether the approach generalizes to other base models.

The benchmark results split into two clusters. On Terminal Bench 2.1, Ember-1 scores 82.0% against K3-max’s 80.9%, with 51.9% fewer tokens. On DeepSWE 1.1, it scores 75.2% against K3-max’s 66.4%, with 23.7% fewer tokens. In both cases, the shorter reasoning chain outperforms the longer one. The exception is SWE-bench Verified: Ember-1 at 92.2% versus K3-max at 93.2% — a 1-point regression, with 15.5% fewer tokens. Two of three coding benchmarks improve when you cut half the reasoning. One regresses slightly. This is the honest read on the data, and it matters for calibrating how aggressively to apply efficiency training.

Benchmark results showing Ember-1 outperforming K3-max with fewer reasoning tokens on Terminal Bench and DeepSWE

The production evidence is harder to dismiss than any benchmark. Two enterprise customers ran A/B tests in which users interacted with the same product while the underlying model switched from K3-max to Ember-1 without notice. The result: 35% fewer total tokens per completed task, 71.3% fewer reasoning tokens specifically, and no detectable quality change as measured by the product teams running the test. Fireworks describes the customer response as developers who “didn’t notice the switch.” Output tokens — the actual text users receive — were largely unchanged. The gap between what the model was generating internally and what the answer needed was 71% of the reasoning trace.

There is a meaningful question about which kinds of tasks the gap is largest on. Fireworks’ benchmarks are weighted toward structured engineering work — coding, clinical decision support, software engineering evaluations. Reasoning traces in those domains tend to follow predictable patterns: check the inputs, plan the approach, execute, verify. A significant portion of that pattern may be recoverable as formula rather than genuine inference. Open-ended generation tasks — writing, analysis, novel problem types — may have less compressible reasoning. The blog doesn’t address this directly, and the HN community’s response was to raise it prominently.

The Same Pattern From a Different Direction

Fireworks isn’t the only lab publishing efficiency data this week. Anthropic’s cost-efficiency documentation for Opus 5.5 includes a comparison that’s more useful than per-token pricing: on SWE-bench Pro, a benchmark of 478 hard engineering problems, Claude Opus 5.5 at its default medium effort scores 92.8% for $0.22 per solved task. Claude Fable 5.1, Anthropic’s previous top-tier model, scores 92.3% for $1.19 per solved task. Slightly lower accuracy, one-fifth the cost per correct answer.

The per-token pricing difference doesn’t fully explain this. Opus 5.5 is priced at $4 per million input tokens, down from $5 for Claude Opus 5, and it generates output more than 30% faster. But the cost-per-solved-task gap against Fable 5.1 is 5.4x — larger than a price cut and a speed improvement together can account for. The difference comes from task completion efficiency: Opus 5.5 “tends to finish the same task with fewer tokens,” as Anthropic’s documentation states, not just at a lower rate per token. The prompting guide for Opus 5.5 describes why: the model generates more thinking per individual turn but needs fewer total turns to close a task. Efficiency through better single-pass reasoning, not through shorter reasoning chains on each pass.

That distinction is worth holding onto. Ember-1 achieves efficiency by pruning the reasoning chain within each step — identifying intermediate conclusions that don’t contribute to the final answer and removing them. Opus 5.5’s efficiency operates at the task level — completing the full job in fewer model calls, with each call producing a more complete and self-correcting result. Two different mechanisms, both arriving at the same economic outcome: less compute per correct answer than the prior generation.

Two production pipelines producing identical outputs: one using 35 percent fewer intermediate steps

Anthropic’s default effort level for Opus 5.5 is medium. Every prior-generation Claude model defaulted to high. The documentation is explicit that medium-effort Opus 5.5 matches or beats high-effort Opus 5 on coding evaluations. This is a behavioral statement about architecture, not just about pricing: the new model doesn’t need the same depth of effort to reach the same result. That’s a different claim than “we made the model cheaper.” It’s a claim that the model reasons more efficiently by design.

The Skeptical Case and Why It Doesn’t Settle the Question

The HN discussion on Ember-1 — 468 points and 217 comments, and clearly the community taking the announcement seriously — ran several critiques worth engaging. The most substantive: Fireworks didn’t disclose the training algorithm. “On-policy planning and learning” is a phrase, not a method. A lab that had produced a generalizable efficiency technique would have reason to publish it; Fireworks is treating this as proprietary, which suggests they view it as a competitive moat rather than a research contribution. Whether the approach transfers to other base models, or only to K3’s specific architecture, is unanswerable from the announcement.

The second critique concerns Bedside Bench, the clinical benchmark Fireworks uses for their headline “Pareto frontier” claim. Bedside Bench is operated by Doximity, a healthcare professional network. Fireworks does not declare any commercial relationship with Doximity in the announcement, and the specific benchmark was chosen to lead the results section. A benchmark that outperforms GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost-per-task in a domain that is neither Fireworks’ stated focus area nor an independently audited evaluation deserves scrutiny. It may be accurate. It was not independently verified.

Ember-1’s research-preview status also places a ceiling on how much weight to give the production results. Two customers, internal A/B tests, no disclosed methodology for the quality assessment beyond “developers didn’t notice.” The number is directionally correct if you believe the customers were genuinely measuring output quality rather than just uptime. It doesn’t constitute a large-scale validation.

What the critiques leave intact: the benchmark pattern where Ember-1 outperforms K3-max on two of three coding evals while using half the tokens, and the independent cost efficiency data from Anthropic on Opus 5.5. Those two data points don’t depend on each other or on Fireworks’ claimed methodology.

What the Math Looks Like at Production Scale

The economic argument for efficiency training runs on the numbers Anthropic published. The $0.22 and $1.19 figures are task-level totals — the cumulative cost of every LLM call an agent makes to close one SWE-bench Pro problem. At 10,000 engineering tasks per month, Opus 5.5 costs $2,200 against Fable 5.1’s $11,900. Per year, $26,400 versus $142,800. Scale this to the kind of automated codebase operations a large engineering org runs — PR reviews, issue triage, test generation — and the difference is material at a budget level, not just a line item. The caveat: SWE-bench Pro is a coding benchmark, and these cost ratios apply to workloads similar to it. The only way to know if your workload maps to this ratio is to measure it.

Fireworks’ per-token pricing is worth noting because the HN thread flagged it: Ember-1 charges the same rate per token as the base K3 model. The efficiency gain flows to the customer through reduced token count, not through a reduced rate. Fireworks’ margin per token on Ember-1 is therefore higher than on K3-max, since the model costs less to run inference on. This is rational business behavior and doesn’t undermine Ember-1’s value to buyers — 35% fewer tokens at the same price-per-token is still 35% less spend. But it signals that efficiency-tuned variants may carry a premium in future pricing, once the market establishes willingness to pay for them separately from capability.

The correct model for thinking about what this week’s releases mean is not “cheaper AI.” It’s “the over-computation in current generation models is now measurable and correctable, and the correction is starting.” Fireworks measured 71% of reasoning tokens as optional in production. Anthropic built a model that reaches the same benchmark scores at less-than-default effort. Both findings point at the same architectural fact: the training process that produced the current frontier generation optimized for final-answer quality, not for reasoning economy. Long chains of thought scored well on benchmarks during training. No signal said “that step was unnecessary.” Efficiency training applies that missing signal retroactively, and it works.

The labs training the next generation will have to decide whether to bake that signal in from the start or apply it as a post-training optimization. Fireworks has demonstrated the post-training case. Anthropic’s architecture decisions on Opus 5.5 suggest the in-training case is being built. The models that come out the other side of this — trained with token economy as an explicit objective alongside accuracy — will be faster, cheaper, and capable of the same work. For production agent builders evaluating their inference spend right now: this week’s releases are not incremental updates. They are the first confirmed measurements of how much has been wasted.

Two brass hourglasses in warm sidelighting — left filled with fine golden sand, right with sparse silver granules, both at identical fill level

AI-generated editorial illustration · TemperatureZero · September 28, 2026

Keep reading the signal

Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.

Subscribe Free

Continue the archive

Latest BriefingsArticlesAbout Temperature Zero