AI NewsProductivityReported
Wagtail spent September coding on GLM 5.3 Flash, used 2 billion tokens, and only half landed on the target model
Wagtail's Thibaud Colas tried to spend September coding only on GLM 5.3 Flash, and reports that 1 of 2 billion tokens went to other models, including a $150 overrun on a vibe-coded MCP prototype.

Image: Wagtail CMS
Why it mattersA team that commits to a cheap open model needs a per-project cost cap in dollars and energy, because one wrong model pick on an overnight prototype can spend more than the planned monthly budget.
Picking a cheap open model for a team's day-to-day coding takes one afternoon, and a single forgotten prototype can spend the whole month's budget in a weekend.
Thibaud Colas, a core team member at the open-source CMS Wagtail, published his report on 2 October after spending September trying to code entirely on GLM 5.3 Flash, an open-weights model the team had set as its target at the start of the month. The piece is the latest in the Wagtail project's Agentic engineering recommendations series.
Half the tokens, not the target
Of 2 billion tokens used across September, GLM 5.3 Flash took about 1 billion, Colas writes. The rest went to DeepSeek V4.1 Flash, Qwen 3.8 Flash and other models. The GLM share stayed within budget at $68, which Colas estimates at about 4 kilowatt-hours of energy use and 365 grams of carbon.
The first half of September held to the plan. The second half did not, for three measured reasons.
The $150 prototype
Wagtail is building an experimental MCP server for the CMS, and Colas calls it a "vibe-coded prototype" by design. He picked the wrong model for the prototype's agent, and spent 450 million tokens, $150 and about 5 kilowatt-hours of energy on it near-overnight. The server itself ships as a working demo, so the money was not wasted. Colas writes that the same output was reachable at about a fifth of the cost with a few more choices up front.
The infrastructure the giant labs do not share
The second block was capacity. Wagtail runs its open-model inference through third-party providers rather than the big labs, and the providers that serve a popular small open model do not have the same spare GPUs as the frontier labs. Colas notes performance degradation on GLM 5.3 Flash in particular, which forced the team onto DeepSeek V4.1 Flash and Qwen 3.8 Flash when the primary route slowed down. The switch itself is one configuration change, but the need to switch was unplanned.
Benchmark work has its own tokens
The third block was deliberate. The team is building a Wagtail-specific benchmark across many models, and that comparison needs real traffic on each candidate. Colas calls this cost essential, because without it the only guidance a maintainer can give to another team is a story rather than a number.
Taking the four lessons together, Colas treats the month as useful more than successful. His October rules: measure usage locally in energy and dollars as well as tokens; budget separately for prototypes; use orchestrator, scout, implementer and reviewer roles with bounded goals; and watch the decision-diffusion models in the Jev family for tasks where their speed makes the switch worthwhile. His target for the next month is still that the majority of day-to-day inference should sit on cheap flash-tier models, measured in dollars and energy rather than token counts.
For a team planning its own open-model switch, the useful fact is the shape of the overrun. The budgeted work stayed inside its planned spend of $68. One prototype with the wrong model picked once, running agentically overnight, cost more than twice that in a day. A cost cap set per project, not per month, is the control that would have caught it.
Source
- Thibaud Colas, One month on GLM 5.3 Flash, Wagtail CMS, 2 October 2026.
- Hacker News discussion: One month coding with GLM 5.3 Flash, 2 October 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


