AI NewsModels & agentsReported

Fireworks Research promised Ember-1 uses 40 percent fewer tokens than Kimi K3, and a 15-run test measured 23 percent

The New Stack ran Ember-1 and Kimi K3 five times each on three tasks. Ember-1 was 3.4 times faster and used 23 percent fewer reasoning tokens, below the 40 percent Fireworks Research had advertised.

AI News

Editorial3 min read

LinkedInX

Why it mattersA vendor tuning claim of 40 percent fewer tokens shrank to 23 percent under a 15-run test, and Ember-1 becomes more expensive than the cheapest Kimi K3 provider, so a team should measure both on its own workload before switching.

Buying a model on the number in its launch post is how a team ends up paying for the wrong thing. Fireworks Research launched Ember-1 on 23 September as a research preview built on Moonshot's open-weight Kimi K3, and said in the launch post that it matches Kimi K3's quality on 40 percent fewer reasoning tokens. The New Stack ran both models five times each on three tasks and got a smaller gap, a faster model and one arithmetic slip. The reporter is Jessica Wachtel and every figure below is from her write-up.

Both models were called through OpenRouter with identical prompts and default reasoning settings, and both were routed to Fireworks so the speed comparison was fair. Fireworks and Kimi K3 on Fireworks were billed the same, at $3 per million input tokens and $15 per million output tokens. Ember-1 has no cheaper provider yet, and Kimi K3 is sold by other providers for as little as $1 in and $9 out.

What the tests measured

Three tasks, each run five times per model. Logic puzzles were three problems with 4, 5 and 7 engineers, each with one correct solution. Deploy scheduling was 12 services with dependencies, one team deploying at a time, and two blackout windows to route around, with a known best of 17 hours. Probability was five questions about a retry system with a healthy and a degraded server plus a circuit breaker, and every answer had been checked against a simulation of two million requests before either model saw the prompt.

Kimi K3 got 15 of 15 runs right. Ember-1 got 14 of 15. The single miss was on the fifth probability run, where Ember-1 answered 0.94619 instead of 0.94629, and the arithmetic slip carried into two of the other answers.

The token gap by test

Averaged across the three tests, Ember-1 used 23 percent fewer reasoning tokens than Kimi K3. Only the probability test came close to the 40 percent figure in the launch post. On the logic puzzles the gap was 18 percent, and on deploy scheduling it was 16 percent.

The finding the launch post did not carry was speed. Ember-1 finished each test 3.4 times faster than Kimi K3. Kimi K3's slowest logic-puzzle run took nearly 20 minutes, against 3 minutes 46 seconds on Ember-1. Deploy scheduling was 1 minute 29 seconds against 4 minutes 46. Probability was 1 minute 47 seconds against 6 minutes 48.

What the cost picture looks like

Both models on Fireworks, the 15 runs cost $2.48 for Ember-1 and $3.26 for Kimi K3, a 24 percent saving. Change the provider and the ranking flips: the same Kimi K3 runs at the cheapest listed price of $1 in and $9 out would have cost $1.96, less than Ember-1. Wachtel did not test the cheaper Kimi K3 providers for speed, so a reader who wants Kimi K3 for less has to accept that the speed numbers above may not apply.

The article closes the way a real test should: if you want speed at almost the same accuracy, use Ember-1; if you have time and want a lower bill, route Kimi K3 to a cheaper provider.

A team weighing the two on its own workload can lift the prompts from the end of the piece and run the same five-times-per-test check.

Source

Reported byThe New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX