Detrended

Subscribe.

No spam. Unsubscribe anytime.

The race to zero.

Justin Pyvis4 September 2026Artificial Intelligence · Prices

Most of my own AI use is through a self-hosted Open WebUI front-end plugged into a few cheap, zero-retention open weight models. For real work, I use the OpenCode harness to give agents access to a local folder and prompt them to build.

For most tasks (i.e., all but the most complicated), my model of choice has been DeepSeek V4 Flash, which since it was released at the end of July has been leading the open-weight field on cost per task and usage metrics. But then last week, something happened. A stealth model dubbed "Ox Alpha" arrived at a price per token of $0, and people jumped all over it (orange bars in the chart):

OpenCode Market Share by model.

Notice what happened. When Ox Alpha dropped, people quickly substituted away from other models, including the already-cheap DeepSeek V4 Flash. But the second it was revealed that Ox Alpha was in fact GLM-5.3-Flash in disguise, and that users would have to start paying for it, its market share plunged. That's despite GLM-5.3-Flash officially launching with a discounted price that knocked DeepSeek V4 Flash off the AI Pareto Frontier:

AI Pareto frontier.

There are a few reasons for the rapid reversal. The first is GLM-5.3-Flash's slow speed, which people tolerated while it was free but became costly once they had to pay with real dollars. People care about price, and they care about intelligence, but ultimately they want to get their work done as quickly as possible. On that metric, GLM-5.3-Flash is an outlier at its intelligence level; it generates about 40-50 tokens a second, making it roughly half the speed of the flagship it undercuts. It's also verbose, burning far more tokens to finish the same task. Once users were asked to pay a dollar price, the calculus changed and they quickly abandoned it.

Intelligence vs time per task.

Another reason for GLM-5.3-Flash's rapid loss of market share was the arrival of a new competitor in the form of Meta's open weight Muse Spark model (purple bars in the first chart). Meta recently made it "free" in exchange for input data and coding interactions to train their AI models. Given what we know about people's willingness to trade their data for other "free" services like social media, it's going to be tough to compete with such a model while it remains in data harvesting mode.

OpenCode Go indicative usage allowances.

So, what does all of this mean? It's tough to say definitively, because there's clearly selection bias at play: OpenCode customers aren't your typical AI users (they're perhaps the most price-sensitive segment). The arrival of Meta's Muse Spark and its "free" contributory model simply reinforced that: OpenCode's users will happily give up their input data to save a few dollars.

Can we generalise that behaviour across the entire market for AI? Probably not. But I think it's directionally correct, and raises the question of whether enough people will be willing to pay for frontier models to justify the huge training costs involved as the marginal cost of inference falls closer and closer to zero. I'm just not so sure they will; even the less discerning corporate customers of Anthropic aren't using its top model, Fable, all that much, primarily because of its "high price and the fact that older models are capable of handling the bulk of business demands".

The obvious objection is that a model smart enough to do things current ones can't would command a premium. Maybe so, but the market evidence so far suggests plenty of users are perfectly happy with cheaper, less intelligent models that handle most tasks just fine.

If that's the case and people aren't willing to continue paying for even more intelligence, then it may not be worth making them much smarter than today's models. In economic terms, the marginal value of additional intelligence is falling relative to its marginal cost, because current models are already good enough for most tasks.

That's the point at which AI moves close to being commoditised at the inference layer, and whatever product advantage leaders like Anthropic and OpenAI have erodes (barring a regulatory moat). The market is then dominated by interchangeable low- and mid-tier open weight models hosted wherever energy is cheapest and data centre construction costs are lowest. Total spending might still increase as usage expands to fill cheap capacity, as happened with cloud compute, but who captures it changes: monetisation migrates from point of use to other avenues, such as data harvesting, product bundling, or Silicon Valley's favourite, advertising.

I'll leave it to the reader to decide what that might mean for the respective valuations of the leading AI companies.


ShareXLinkedIn

About the author

Justin Pyvis

Independent economist and casual techie based in Perth, Western Australia.More →

Subscribe to the newsletter.

Get new essays delivered to your inbox. No spam. Unsubscribe anytime.