The 3 Trillion Dollar GLM 5.5 Problem

Share

Big Labs have a big problem, and that problem is GLM 5.5. When I refer to GLM 5.5, I don't mean that exact model (wxhich doesn't exist yet, GLM 5.2 is the latest in that series and came out two months ago). GLM is a stand in for the frontier open weight LLMs (it could be Qwen or Kimi or any other model series). What GLM 5.5 represents is when open models are good enough.

Note that in this article, I will assume there is demand for these models in the future, obviously if society pushes back on their adoption then the value will also go down.

What do I mean by "good enough"? I mean that the use cases that generate the vast majority of revenue for OpenAI and Anthropic do not have infinite scaling requirements. These models are used predominantly for coding (along side other white collar tasks). Most of these tasks have a threshold of capability requirements. For example, if you imagine the use case of some small business who wants to create a website, and uses an LLM to code it up, that doesn't require Fable 5 much less Opus 5. Fable-5 goes for $10/$50 (per million input/output tokens) on openrouter currently, but I just tested with a model 2 orders of magnitude cheaper ($0.15/$0.15) and was able to get something quite good. Likely the requirement thresholds for email related tasks were met by open models years ago.

Why is this a >3 trillion dollar problem? Anthropic and OpenAI's multi-trillion dollar valuations depend on every month having higher revenue than all previous months combined, and if something blows a hole in coding revenue that would be nothing short of disastrous for their valuations (which in turn would be disastrous for the companies whose value exists due to the >1T of backlog that the big labs need to pay for with future revenue). Once an open weight model reaches a performance that is sufficient for the majority of coding use cases (which is probably mostly just SaaS drivel), revenue will evaporate. Why would businesses pay 1-2 orders of magnitude more for something that achieves the same thing? LLM models have as close to 0 switching cost as could possible exist (literally just type "/model" in OpenCode, you can do it even in the middle of a session). Clearly this only applies to token costs, as long as Anthropic and OpenAI are selling $5,000 for $200, then of course people will take the free money. So if this point is reached while heavy subsidies are still in place, then the effect will be muted in the near term.

One could argue that this is already taking place (and it's just in the early adopters phase, since many users are subsidized through monthly plans). For example, Coinbase recently put out a graph saying that effectively 100% of their code was LLM generated (fittingly, with an AI generated graph). If we take this as true (maybe it is, maybe it's made up, maybe some code is written by hand and people just run git push in Claude Code so it shows up as AI authored, who knows), this seems to indicate that the above has already come to pass in terms of capability saturation. New models that improve 2% on whatever benchmark won't matter to Coinbase, only faster and cheaper models will. One could argue that new models will unlock capabilities required for new/larger/more complex projects that they are not undertaking because these model capabilities don't exist yet. But the employees still exist, if there was some project the business would benefit from that the models they are using couldn't handle, then they could just build it themselves. Since that isn't happening (since none of the code is hand generated), those projects must not exist. Now the skeptic might say, "hey actually with layoffs that render employee morale so low they have no energy or interest in advancing new projects, combined with their extensive LLM usage that has already turned their brains into mush, rendering their tolerance for difficult thought and real work so low that anything beyond drooling into a prompt window eludes them, and thus people will just not work on anything that they can't easily do with LLMs". And while that is possible, such pessimistic speculation is can neither be confirmed nor denied.

Why do I use GLM 5.5 to represent this? GLM is the series that is the closest to frontier performance of the open weight models (at the time of writing this was true, now it is Kimi K3) and is substantially cheaper ($0.4/$1.25 currently). GLM 5.2 reviews indicate that it made substantial gains, but wasn't quite sufficient yet, and thus that is why I use GLM 5.5 to represent the model that will be sufficient (as a placeholder, not as a prediction).

Even if costs of open weight models increase to be merely equal, and not massively cheaper, open weight models have a multitude of benefits over closed weight models: you have control of your data, you can modify/improve the model to your specific use case, you can control the deployment (allowing for better speed), etc.

Additionally, even if GPT-7 or Fable-7 is perfectly trained and achieves the irreducible validation loss on all of the massive piles of data they have stolen, that only matters if it unlocks substantially new classes of revenue. Although the new models get a lot of a hype for different areas, the simple fact is that the market already converged on the economic value of solving Erdos problems and it's a few billion dollars short of CRUD app coding revenue.

This problem isn't a problem until it's a problem, at which time it becomes a big problem for big labs. Some investors are starting to realize this and the reactions when it comes to whether it is a problem or how to fix it generally fall into a few categories. The first is something along the lines of 'western businesses won't use Chinese models because China is bad'. This relies on the idea that the open weight models are made in China which isn't necessary for the economic issues, but does seem to be the trend. This also often conflates the idea of "Chinese AI" with "sending your data to China". Just because a model is trained in China, doesn't mean you have to host it there (after all, torch.matmul is the same in the US or China). An endless list of US inference companies exist that will run this model on their hardware in the US so the data never goes to China (besides giving your data to a US company just adds a middle man to your data ending up in China). How can Anthropic and OpenAI solve this problem of theirs? Provide cheaper more competitive models? Release some of their own open weight models? Rather it seems like they want to skip ahead to the 3rd phase of how Big Tech handles open technologies and get the government to restrict their competitors.

Since writing this blog, Kimi K3 was released, which I won't expound upon here, but represents a serious threat to the big labs, albeit it a different one than I described. K3 isn't cheaper than other frontier models by the same amount (roughly 1/2 the cost of GPT 5.6 and 1/3 the cost of Fable 5), but is (on several supposedly important benchmarks) better than their models. This shatters the illusion that China is "behind" or limited to distilling frontier models and means that not only can the big labs lose on cost, but they can lose on raw capabilities as well. This also indicates that Chinese labs continue to advance and don't seem likely to plateau before they are able to reach this capability threshold.

These ideas were covered in a video here: https://www.youtube.com/watch?v=1DeoJ3DLTv0