On July 16, Moonshot AI's Kimi K3 climbed to number one on the global coding benchmark. 1,679 points on the Frontend Code Arena. Claude Fable at 1,631. GPT-5.6 at 1,618. A Chinese open-source model at the frontier. Within 48 hours, demand overwhelmed their own infrastructure and they had to pause new subscriptions.

The reaction: panic. China is catching up. American AI is under threat. The hyperscalers are spending trillions on a bubble. Anthropic and OpenAI are in trouble.

All of it is the wrong read. Kimi K3 is not a sign AI is slowing down or cracking. It is the clearest signal yet that AI is entering its second era. And the second era scales faster than the first.

Era 1: The Model Monopoly

Two or three frontier labs controlled the model layer at extraordinary margins. All eyes were on who had the best model. Infrastructure was an afterthought, too slow, too constrained, perpetually behind. The labs were the story. Everything underneath them was just plumbing.

What that structure actually meant: the labs became monopsonies. The only serious buyers of compute, memory, power, data center capacity. Suppliers had no leverage. Infrastructure investment moved at the pace the labs tolerated. Bottlenecks persisted because there was no competitive pressure to solve them urgently. And at the model layer itself, with only two players at the frontier, competition was minimal. No one was forced to compete on efficiency or cost or consumer product quality. The incentive was to maintain the lead, not to accelerate the field.

The monopoly looked like dominance. The more accurate word was ceiling. The first era of AI had a structural cap on how fast it could develop, and that cap was built into the market structure at the very top.

The Transition

Kimi K3 is the signal that Era 1 is ending.

When a Chinese open-source model reaches the frontier, the two-lab monopoly cracks. When that model's weights go fully public on July 27, any company in the world can run a frontier model on their own infrastructure without paying inference margins to Anthropic or OpenAI. The extractive power of the model layer breaks. And when that happens, something important shifts: attention and capital flow back down to infrastructure.

That is where Era 2 begins.

Era 2: Scaling in Both Directions

We saw this movie once before. When DeepSeek dropped in January 2025, Nvidia lost hundreds of billions in market cap in a single session. Same panic: China built a better, cheaper model, the infrastructure buildout is a bubble, sell everything. The recovery came when the market figured out the right answer: it does not matter who wins the model race. You still need the same compute to run it. Infrastructure came back stronger. Kimi K3 is DeepSeek at a higher level, except now it is not one model, it is systematic fragmentation of the entire model layer. The point stands with more force.

The second era has a structure that is genuinely new. Two things scale simultaneously, in opposite directions, amplifying each other.

Infrastructure scales up. As model margins compress and the monopoly breaks, capital floods into every layer of the AI supply chain that was previously underfunded. The memory shortages, the NAND bottlenecks, the one-year backlogs on chip substrates, the photonics capacity constraints, all of these get funded aggressively because now it is profitable to solve them. Hyperscalers commit hundreds of billions. The buildout accelerates precisely because the demand signal is no longer controlled by two labs extracting maximum margin at the top.

Models scale down. With real competition at the frontier, labs can no longer afford to compete only on benchmark scores. They compete on what actually matters: token efficiency, intelligence per dollar, consumer product quality, distribution. Models get better and cheaper simultaneously. Access expands. The cost of deploying frontier AI at the application layer falls.

These two vectors do not cancel each other. They compound. Cheaper models create more users. More users demand more compute. More compute investment solves bottlenecks. Solved bottlenecks enable even more capable and efficient models. The loop runs faster with every turn.

This is why the second era of AI scales faster than the first. Not because the technology suddenly improved. Because the structure that was limiting it got disrupted.

The Bubble Answer

Here is where the skeptics push back. The hyperscalers are spending trillions. The use cases are not there yet. This is the dot-com bubble with better PR.

The argument is wrong, and Jevons Paradox explains why.

In 1865, economist William Stanley Jevons observed that as steam engines became more efficient, total coal consumption went up, not down. More efficient engines made coal-powered production economically viable in more applications, which expanded demand faster than efficiency reduced it per unit. The efficiency gain multiplied use cases. Use cases multiplied demand.

This is the iron law of every general-purpose technology. Transistors got cheaper and more efficient. We used more of them, not fewer. The internet got faster and cheaper. Traffic exploded. Cloud computing got more cost-efficient. Cloud spend tripled. In every case, efficiency gains unlocked applications that were not viable before, which created demand that dwarfed whatever compute the efficiency saved.

AI is the most general-purpose technology ever built. The question is not whether a 100x more efficient model requires less compute per task. It does. The question is what happens when that model reaches a billion users across ten thousand new applications that were not viable at the old price point. Total compute demand goes up by an order of magnitude.

But there is something else. The bubble accusation assumes you are building infrastructure for a specific use case that may or may not arrive. That is what the dot-com bubble actually was: fiber laid for broadband video before broadband existed. The use case was specific and premature.

AI infrastructure is different. The two-direction scaling means you are simultaneously building infrastructure and reducing the cost of discovering use cases. Every time models get cheaper, more developers, more companies, more people in more niches start using AI for things nobody planned for. Healthcare, science, manufacturing, logistics, education, research, creative work. The infrastructure does not need to know which use case wins. It just needs to exist when the discovery happens. And discovery is happening faster with every price drop.

The Flick

At some point, the two vectors meet. Infrastructure built out enough. Models cheap enough. Use cases discovered and compounding. And it becomes undeniable, clear as day, that AI is the most cost-efficient intelligence ever created and it improves every part of human endeavor.

That moment is what this whole scaling dynamic is building toward. Not a specific application. Not a single breakthrough. A critical mass where AI utility is so obvious, so embedded, so cheap that the debate about whether this was real stops entirely.

You only reach that moment by scaling in both directions at once. Infrastructure without cheaper models means use cases stay locked behind high access costs. Cheaper models without infrastructure means you hit the ceiling immediately. Both together is how you get to mass adoption. Both together is exactly what Kimi K3 signals is now happening.

Why Compute Is the Trade

This entire thesis points to one conclusion: compute is the best-positioned trade in the market right now. Not a specific model. Not a specific lab. The infrastructure underneath every model that runs, trains, or gets discovered.

The common objection: it has already run so much. Three years of extraordinary performance. Feels risky to be buying here.

This is where the Lindy Effect is worth understanding. Named after a deli in New York where comedians observed that the longer a comedian had been popular, the longer they were likely to remain popular, the Lindy Effect states that the future life expectancy of a non-perishable thing grows with its current age. Applied to technology trends: a structural buildout that has run hard for three years is more likely, not less likely, to continue for another three. The trend has already survived every reason it was supposed to stop. DeepSeek Friday. Valuation concerns. Rate hikes. Regulatory threats. Each time it came back stronger. That survival record is itself a signal.

And then there is the Red Queen Effect, borrowed from evolutionary biology. In a competitive ecosystem, you have to keep running just to stay in the same place. Every hyperscaler, every neocloud, every data center operator is in a race where stopping means falling permanently behind. The competitive pressure creates a structural floor on investment that only moves in one direction. These companies are not spending because they want to. They are spending because the alternative is failure.

The risk is not buying compute at a high. The risk is underestimating what comes next and missing it entirely.

Given the pace of progress in 2025 and 2026, betting against this in the next twelve months is a bet against a trend that is playing out in front of your eyes exactly as it should for it to go even more parabolic.

The monopoly was the bottleneck. It is breaking. The second era just started.

Read our previous piece on why memory is the bottleneck of all bottlenecks.

This is not financial advice. neym is an independent research newsletter. The author may hold positions in assets discussed. Do your own research.