Wagtail team fails month-long GLM 5.3 Flash challenge, 50% of 2B tokens

Wagtail team fails month-long GLM 5.3 Flash challenge, 50% of 2B tokens

The Wagtail team set itself a challenge: spend the whole of September on a single efficient open model, GLM 5.3 Flash. In a blog post on wagtail.org, the team calls the result "task failed successfully". Across the month the team used about 2B tokens, and only 1B of them, 50%, went to the target model.\n\nThe first half of the month went as planned. GLM 5.3 Flash usage stayed well within budget: $68, about 4 kWh of energy and 365 grams of carbon emissions. The team tracked the token split with AgentsView, one of its own recommended tools for keeping tabs on AI usage. The second half went badly, with 1B tokens going to other models. The post names three causes.\n\nThe first was the cost of vibe coding. The team's experimental Wagtail MCP server is openly a vibe-coded prototype, which the author says is fine for a prototype. But the author chose the "wrong" model for it, and the team burned 450M tokens, $150 and 5 kWh almost overnight. The post does not name that model. The MCP server works well and the team now has a good demo of its capabilities, so the spend was not for nothing. Still, the team says similar results were most likely possible at about 5x less cost with not much more effort.\n\nThe second was infrastructure. The team's inference providers work most of the time, but they are popular and do not have the capacity of the big labs that, in the post's words, hoard all the GPUs. The team saw performance degradation on GLM 5.3 Flash in particular, most likely because it sits so high on the Pareto frontier of models relevant to its work. It switched to similar models, DeepSeek V4.1 Flash and Qwen 3.8 Flash, which the post says is easy to do but was unexpected.\n\nThe third was experimentation and R&D. Beyond using one model for day-to-day engineering, the team felt it had to keep trying a wide range of models, especially as it starts to benchmark model performance on Wagtail tasks, which needs data across many models. The post shows a sneak peek of the benchmark in an image. The team argues that concrete data makes it easier to steer people toward leaner options, and it wants to make those options more viable through agent skills and a new CLI prototype intended to work well with agents.\n\nThe bottom line: only 50% of tokens went to the target model, and total energy use was about 35 kWh instead of 10. The team still says it learned a lot. For October it lists four fixes: constant local measurement of tokens, energy use and spend, ideally tied to concrete outcomes; a separate budget for experimentation; better prompt selection and multi-agent techniques (orchestrator, scout, implementer and reviewer agents, with bounded goals); and continued pushing for more efficient techniques and models. The post says Jev-style decision diffusion models look very promising if they can run so efficiently, and that the latest flagship models look like a step in the right direction.\n\nThe team's conclusion is that for day-to-day developer work it is totally viable to focus on one or two flash-tier cheap models. A viable target, it says, is probably that the majority of AI inference work be done with such efficient models, measured in cost or energy use rather than tokens. That is the goal for October, and the team plans to report how it pans out at Wagtail Space 2026 in November.

Key facts

  • The Wagtail team tried to use only GLM 5.3 Flash for September and calls the challenge a technical failure: 1B of about 2B tokens (50%) went to the target model.
  • GLM 5.3 Flash usage cost $68, about 4 kWh and 365 grams of carbon emissions; total energy use for the month was about 35 kWh instead of 10.
  • A vibe-coded Wagtail MCP server prototype, built on the "wrong" model, used 450M tokens, $150 and 5 kWh almost overnight; the team says it could most likely have cost about 5x less.
  • Provider capacity limits and degraded GLM 5.3 Flash performance pushed the team to DeepSeek V4.1 Flash and Qwen 3.8 Flash.
  • The team still calls it viable to rely on one or two flash-tier cheap models for day-to-day developer work, and sets a similar goal for October.

Why it matters

Most teams talk about cheap models in the abstract. This is a small team's own month of numbers: what a flash-tier model cost ($68 for roughly 1B tokens by the source's split), and where a plan to use one model broke down. The main lesson is that the bill came from decisions around the model, a prototype built on the wrong model and a need to experiment, more than from day-to-day coding on the cheap one. The team also argues for measuring AI use in cost or energy rather than tokens.

Who it affects

Developers and engineering teams who use AI coding agents and want to control spend and energy use. It is most relevant to teams that build prototypes with agents, rely on third-party inference providers, or plan to benchmark several models.

How to use it

The post's own list for a repeat attempt: track tokens, energy use and spend locally and continuously, ideally alongside how well the work leads to concrete outcomes; budget for experimentation, not only day-to-day tasks; decide more deliberately which prototypes are worth building and how; use bounded goals and split agent roles into orchestrator, scout, implementer and reviewer. The team uses AgentsView to see the token distribution. For daily developer work it suggests focusing on one or two flash-tier cheap models and switching to similar ones, such as DeepSeek V4.1 Flash or Qwen 3.8 Flash, when a provider degrades.

How solid is it

This is a first-person account from the Wagtail team on its own blog, and the source names no individual author. The headline figures (about 2B tokens, 1B on the target model, $68, about 4 kWh, 450M tokens, $150, 5 kWh, about 35 kWh) come from the team's own tracking. The 5x saving is an estimate that the post itself hedges with "most likely". The cause of the GLM 5.3 Flash degradation is likewise given only as "most likely". The benchmark is shown only as an image preview, and no figures from it are given in the text.

Risks and caveats

The experiment came from one team on its own workloads and providers, so the numbers may not transfer. The post does not name the providers or the model that caused the overnight spend. It does not break the month's 35 kWh total down by model beyond the GLM 5.3 Flash figure. The claim that flash-tier models are viable for day-to-day work is the team's judgement. The team also notes that heavy use of popular cheap models can run into provider capacity limits.

“So technically this challenge was a failure.”

— Wagtail team, "One month on GLM 5.3 Flash"