Replit says giving the model the choice of effort and subagents beats a router, at 72 percent on DeepSWE for $2.11 per task
Replit's engineering post argues that a router that picks a model for each turn is capped by its own picking rules, and reports that letting the model decide its own effort level and when to hand work to a subagent scores 72 percent on DeepSWE v1.1 at $2.11 per task, against 74 percent at $4.43 for the same frontier model on its own.

Image: Replit
Why it mattersA team building coding agents can stop tuning a router that is by design less capable than the model behind it, and let the model itself decide how hard to think and whom to hand a job to.
Model routers pick which large language model handles each turn. The picker is a smaller heuristic or a small model, and it never becomes smarter than the ones it routes to. Replit's engineering team published a post today that starts from this limit and describes a different design: leave those decisions to the model, and give it the levers to control its own effort and to hire a subagent when it wants one.
The post, titled "Free the models: Harness design at the frontier," is signed by Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi and Michele Catasta at Replit, and dated 29 September 2026. It reports two benchmark results and a pattern seen in production.
Two benchmarks, one design
On DeepSWE v1.1, Replit says its agent scores 72 percent per task at a cost of $2.11. Running GPT-6 Astra on its own inside a light harness scores 67 percent at $1.60 on low effort, and 74 percent at $4.43 on the highest effort. A separate sidekick architecture scores 61 percent at $1.34.
On Terminal-Bench 4.0, Replit's numbers are 49 percent at $2.53 per task for Replit Agent, 42 percent at $2.25 for Astra on low effort, 60 percent at $5.86 for Astra on the highest effort, and 33 percent at $1.84 for the sidekick setup. Replit says its agent beats the sidekick score by 11 points on DeepSWE and by 16 points on Terminal-Bench. These are Replit's own benchmark runs and Replit's own reported costs, and it does not compare against a router built by anyone else.
The three decisions the model gets
Replit describes three choices its harness passes to the model on every step. It picks its own effort level (low, medium, high, xhigh, max) and can change it mid turn without losing the prompt cache. It decides whether the next job belongs to a subagent or to itself, which controls both parallelism and how much context stays in the main loop. And it picks which specialist subagent takes the job: coding, user interface design, slides, writing, or a general worker.
In production traces Replit shares, Fable 5.1 sends a job to a subagent on 21 percent of its turns and GPT-6 Astra does so on 36 percent. When Astra delegates, it returns to a subagent it has already used 42 percent of the time, rather than creating a new one.
What Replit does not say
The post does not report end user latency for its own agent, only cost and task score. It does not describe a cache hit rate for the mid-turn effort switch it says preserves the cache. It does not compare its harness against a specific named router product; the point of comparison is a single model at various effort levels and Replit's own sidekick architecture. And the benchmark harness Replit calls "mini-swe-agent" is the one it ran Astra on for the single-model rows, not a fair copy of anyone else's production stack.
For a team building coding agents, the useful line is that a router adds a decision layer between the request and the model that is by design less capable than the model itself, and Replit's post is one team's argument, with numbers, that the layer should go. Adopting that costs some determinism, because the model now picks its own effort. Replit's answer is that the choice is worth it, and its Terminal-Bench line at 49 percent for $2.53 is the number a reader would check against their own workload.
Source
- Replit engineering blog, primary source: https://replit.com/blog/free-the-models
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


