Why smarter AI models could drive up compute prices 10x
11m 18s
The rapid 10x growth in Anthropics revenue—despite only 3x annual compute scaling—highlights a fundamental shift in AI economics. To close this gap, labs must either increase margins, raise compute prices, or shift compute use toward inference. All three factors are already occurring: margins have surged to over 80%, compute spot prices are 40% higher, and inference compute share has risen to 50%+. This suggests AI is increasingly efficient, with models now monetizing compute far more effectively than before—potentially creating 15x more value per unit. The rising cost of compute stems from scarcity, not supply expansion, as physical constraints (Moore’s Law, fab capacity, wafer allocation) limit growth. Top labs, like Google and Anthropic, are paying premium prices for secure, large-scale compute (e.g., $900M/month for 110,000 GPUs), reinforcing the value of efficient, high-performing models. This reflects a core insight: AI’s economies of scale—where training costs are one-time and skills are shared—outperform human labor. While this may lead to greater market concentration and power imbalances, it also signals a future where compute becomes a scarce, high-value asset. The current regime may peak by 2025–2030, after which compute could become affordable again through automation. However, in the near term, the value of AI compute is rising rapidly due to efficiency gains, challenging traditional supply-demand models and echoing historical errors in scarcity predictions.
Today, I want to talk about what the compute situation for the labs will look like over the next few years. For the last three consecutive years, Anthropics revenue has 10x to year over year. And it's likely to do so again this year. So the end of last year with 9 billion revenue, I think they'll probably end this year with somewhere between $100 billion to $150 billion in revenue. Now, for this trend to continue, Anthropics would need to make $1 trillion in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of the capabilities. Does AI get that useful by the end of next year? But suppose the trend does continue. Well, I want to think through what happens in that world. Now, the other big trend in AI is that lab compute only three x's year over year. For a lab to keep 10x in revenue year over year while compute only three x's, one of the following three things needs to happen, or some combination of the three needs to happen. One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that labs spend on inference rather than training has to increase. My understanding is that basically all three of these things are already happening. With regards to the margins and Anthropics, inference margins reportedly went from 40% to the middle of last year to upwards of 80% now available. With regards to compute, the spot prices for compute are more than 40% higher than they were in the February trough that we had earlier this year. And with regards to the share of compute that goes to trading versus inference, in 2024, according to epoch, opening AI was spending just a quarter of its compute on inference. And that number is likely closer to 50%, if not higher now. Now, labs have deferred not to do this final thing. Of increasing the share of compute they spend on inference. The way that labs see the world, the whole point of inference revenue is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, and you're basically declaring that AI progress has stalled, and you're just now in the business of being a cloud provider. Now, this is a less compelling business than building AGI. And so the labs do not want to be in this business, nor do they think they are in this world. They think that within a year they'll have built models that make the current ones look extremely shitty. But they need to invest a lot of their compute, the majority of their compute, into doing the trading and experiments that are necessary to build the next model. So that leaves only two options, or how you can get out of this gap, between the fact that lab compute only increases 3x year over year, but revenue increases 10x. Either the labs margins have to increase, so that they get this surplus, or the price of compute has to increase, so that everybody in the stack below the lab gets this surplus. It's not clear to me which world we end up in. Do we end up in a world where we go from 80% margins for some of the top models to greater than 90% margins if the lab margin effect dominates? Well, that would require the leading model to be so far ahead of the competition, because the nature of margins, why they exist in a market economy, is that the thing you are serving is so much better than what somebody else could go get and replace you on the market. But it just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competed away at that level. So that leaves only one other possibility of this escape valve between these two trends, which is that the price of compute has to increase. As I mentioned, this is already starting to happen. And the effect is even stronger, when you look at the tranche of compute that the frontier labs actually need to accumulate, because they can't just go out and buy a spot instance. They need to make sure that they get enough scale to get really good efficiency and flexibility, and also that they have the kind of compute that lends itself to the security they need for their own weights and for their customers' information. So I think a relevant case study here is to look at the compute that Google and Anthropic are renting from SpaceX. Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB 200s and GB 300s. The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than 40% higher than it would have been in February. I want to emphasize the key conclusion here, that as AI models gets smarter, they'll be better able to monetize the same amount of compute. If a true human level software engineer could run on a natural 100 equivalent, then at today's prices for software engineers, that H-100 should rent for over 250K a year. That's over 15x, the current spot price for a natural 100. And this is not even accounting for the fact that your AI can work nights and weekends. Of course, you might expect that if we had 10 million extra software engineers suddenly appear in the economy, the marginal value of a software engineer would decrease. And thus, the revenue that that H-100 would be able to generate would not be 15x higher than news right now. But actually, don't know if this is true. If we apply this argument to people instead of AI's, then this would be the classic lump of labor fallacy. For example, economists generally believe that high school immigration does not decrease wages in the long run because of how innovation and specialization increase the value of labor. Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore. But if you believe what standard economic says, then the marginal value of labor and thus the marginal value of compute should stay astonishingly high. So let's think about what changes in such a world. One of the things that would happen is that as the top labs get better and better at monetizing compute and the cost of compute increases, it becomes harder for anybody else to compete against them because they have to bid for this resource against somebody who is basically able to make better use of it. Another thing that will happen, and I think this is actually the most interesting implication of this whole thought exercise, is that if you can train the best, most efficient model, then you'll be able to charge much higher margins than you can today. This is the Altkin Allen Effective Economics, and what it's basically saying is that it costs $20 an hour to rent an H-100, and that would be extremely stupid to use a weaker or less efficient model because it's going to burn more tokens on your expensive compute to get the exact same result. So labs would be able to charge a much larger premium if they can train a model that better economizes this scarce input. Basically, if you have a model that can get the same result by using less compute, then you've in some sense created more compute, and the value of compute is going to increase. Another thing that will happen is that a lot of current popular applications of AI will probably get priced out. The reason AI is relatively cheap right now is that AI just can't do a lot of things that top humans can do, but this at some point will no longer be the case. And at that point, Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI stop talk. I'm a bit worried that this kind of analysis honestly pattern matches a lot onto the ways that people in the past have been wrong about scarcity. I'm, for example, thinking of the famous Simon Erlich bet. Paul Erlich was this famous doomer about population growth, and he made this bet that a basket of commodities would increase in price rather than decrease in the decade proceeding 1990. And this is a very famous bet because it's supposed to illustrate how Erlich's Malthusian worldview was wrong and how he did not anticipate the way in which market signals and human ingenuity can find better ways to economize scarce inputs. I'm guessing that the analogy to this bet is probably wrong. Other analysis has shown that if that bet had been made in a different decade, Erlich might well have won. But more generally, I think the supply of compute is much less elastic and much less capable of absorbing large demand shocks and much less capable of being accommodated by using different substitutes than the extraction of different metals. To illustrate why I think this 3x and compute capacity year over year is hard to budge or potentially even sustain is that I don't see how any of the three elements that constitute that 3x can be much accelerated. So 1.4x of that is coming from Moore's Law. Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years. 1.2x is coming from building new fabs. This process is ultimately going to be bottlenecked up to 2030 and potentially even beyond by just building new ASML UV machines. Dylan, when he was on the podcast a few months ago, talked about this in great detail. At 1.8x comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs. This is probably going to hit a wall by the end of next year when at the leading edge N3 nodes at TF7C AI will have gone from 60% to 86%. At some point you have just absorbed all leading edge wafer capacity for AI and you can't keep increasing the number. So I don't know how we get even to continue to do 3x compute scaling year over year for the next few years. Let's go beyond that. At the end of the month, I go through the time on a tradition of closing my books. I started by opening Mercury, which is my baking platform, to make sure that all my transactions are properly categorized. Auto-categorization rules handle the predictable stuff pretty well, but I'm constantly working with new contractors, you know, tutors and researchers and videographers, and I'm also trying new tools. Manually categorizing all of these transactions would add a couple of hours of overhead every single month. Instead of going through them one by one, I have Command, which is Mercury's built in AI, to get staff at all of them at once. Command proposes a category for each transaction and provides its rationale. I just review, I fix anything that's off, and I approve. And once all this work is done in Mercury, it syncs everything with QuickBooks. And Command's judgment calls are genuinely good. It does the obvious things like looking at the vendor, but it also investigates who on my team made the purchase and looks at nodes and memos to build up as much context as possible. This is just one of the ways you can use Command to automate the back end of your business. To learn more, go to mercury.com/command. Mercury is a Fintech company, not an FDIC-insured bank. - Bigging services provided through choice financial group
column and a members FDSC. AI generated responses and suggested actions may vary and are not guaranteed. Now, I want to clarify that at some point in the future, compute will get cheap again. At some point, we'll just have robots that can convert shores of silica sand and mines of copper into new computer chips, and then the price of computers basically the raw inputs and the tools required to do this processing. I'm just talking about this current precingularity regime where AI compute merely 3x is year over year, which is not enough to offset how much more valuable AI is becoming over time. By the way, the fact that Anthropical Revenue has been 10x in year over year, whereas their compute has only been 3x in year over year. I think illustrates how strong the economies of scale are in the model business. Unlogically, this makes sense. When you train a model, you just have to spend this one-time cost to learn all these different skills that then get to be shared across all your users. This is very unlike human labor, where each instance has to be retrained from scratch. I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration. But it seems we do. OK, this was an iteration of a blog post that I also released on my website at dworkesh.com. Trick it out for other posts or to be notified when I release a post in the future. Otherwise, I'll see you for the next full episode.
Podcast Summary
Key Points:
Anthropics revenue is growing 10x year over year, projecting from $9B to $100–150B this year, with a potential $1T target by next year—driving a need to reconcile this growth with only 3x compute scaling.
Three pathways exist to close the revenue-compute gap
Lab margins have surged from 40% to over 80%, spot compute prices are 40%+ above February 2024 lows, and inference compute share has risen from 25% to likely 50%+ in 2024.
Labs resist shifting to higher inference use because it signals stagnation and undermines their AGI-building mission.
A key conclusion is that rising compute value stems not from supply increases, but from AI’s superior ability to monetize compute—effectively creating more value per unit.
The most plausible outcome is rising compute prices due to scarcity and high efficiency of top models, especially as frontier labs secure dedicated, secure compute (e.g., Google’s $900M/month rent for SpaceX GPUs at 2x spot prices).
AI’s efficiency gains—like a human-level software engineer running on H-100 hardware—could make compute 15x more valuable, challenging labor supply assumptions.
This mirrors historical miscalculations of scarcity (e.g. Erlich’s bet), suggesting compute scarcity may be more rigid than assumed due to physical and technological bottlenecks (Moore’s Law, fab capacity, wafer allocation).
Compute scaling may plateau by 2025–2030 as AI absorbs all leading-edge wafer capacity, making 3x annual growth unsustainable.
Long-term, compute will eventually become cheaper via automation, but in the current era, its value is rising rapidly due to AI’s transformative efficiency.
Summary:
The rapid 10x growth in Anthropics revenue—despite only 3x annual compute scaling—highlights a fundamental shift in AI economics. To close this gap, labs must either increase margins, raise compute prices, or shift compute use toward inference. All three factors are already occurring: margins have surged to over 80%, compute spot prices are 40% higher, and inference compute share has risen to 50%+.
This suggests AI is increasingly efficient, with models now monetizing compute far more effectively than before—potentially creating 15x more value per unit. The rising cost of compute stems from scarcity, not supply expansion, as physical constraints (Moore’s Law, fab capacity, wafer allocation) limit growth. , $900M/month for 110,000 GPUs), reinforcing the value of efficient, high-performing models.
This reflects a core insight: AI’s economies of scale—where training costs are one-time and skills are shared—outperform human labor. While this may lead to greater market concentration and power imbalances, it also signals a future where compute becomes a scarce, high-value asset. The current regime may peak by 2025–2030, after which compute could become affordable again through automation.
However, in the near term, the value of AI compute is rising rapidly due to efficiency gains, challenging traditional supply-demand models and echoing historical errors in scarcity predictions.
FAQs
Lab revenue has grown 10x year over year, while compute only grows 3x. This gap suggests that labs are becoming more efficient at monetizing compute or that compute prices are rising.
Lab margins have increased significantly, rising from around 40% in the middle of last year to over 80% currently, supporting higher revenue despite limited compute growth.
Spot compute prices have risen by over 40% compared to earlier in the year, and labs are increasingly paying premium prices for secure, scalable compute resources.
Labs prioritize training to build next-generation models; spending on inference is seen as signaling stagnation, which undermines the narrative of advancing AGI.
Yes — if a model uses less compute to achieve the same results, it effectively creates more value, allowing labs to charge higher prices and increasing compute's economic value.
Yes — growth is constrained by Moore’s Law, new fabs (bottlenecked by ASML machines), and AI’s absorption of leading-edge wafer capacity, which is expected to peak by late 2024.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.