The age of abundant AI is over, marked by a shift from abundant cloud resources to a supply-constrained token economy. Memory chips are sold out through 2026, and GPU orders are projected to reach a trillion dollars by 2027, with data center vacancy at just 1%. This scarcity has ended the discount era, giving pricing power to token providers. The token economy fundamentally changes enterprise software costing, replacing traditional metrics like seats and instances with tokens as the unit of value and cost. Demand for tokens is unpredictable, and costs can vary up to 10x between models, especially in agentic workloads where input tokens vastly outnumber output tokens (e.g., 25:1 ratio). This volatility, combined with constrained supply, forces organizations to prioritize procurement leverage, capacity strategy, and relationship-based access to state-of-the-art models. The FinOps discipline is transforming accordingly, with new roles like Director of AI Ops Engineering emerging that require GPU cost optimization and strategic negotiation skills. FinOps now sits at the top of AI operations, reporting to CIOs and CTOs, and practitioners must move into this conversation to advance their careers. The key takeaway is that FinOps is no longer just about cost optimization but about managing value and strategic capacity in a resource-limited environment. The next 6–12 months will see rapid change as organizations adapt to this new reality.
The End of Abundant AI: New Era Hallmarks
The age of abundant AI is over.
JP Morgan's analyst desk reported recently that memory chips are sold out for all of 2026, and Tom Tungooz says it will stay that way for years to come.
Hello from San Diego, welcome to the Synops mostly daily Brief.
I'm Jr.
Stormont, executive director of the Finn OPS Foundation.
And this is the show where we work through a story or two a day most days on what's happening in SAS cloud and AI cost management with real practitioners facing real problems driven by the patterns we're seeing across this interesting and broad community.
I wrapped up a three day run over the weekend on what's happening at Fin OPS X with breakout talks.
Today I want to step back and look at the bigger picture, the token economy, why it is reshaping the Fen OPS discipline and, and and frankly every discipline with acronyms and every role and every job out there.
There are some new jobs that are being created right now because of this shift.
And we are just 28 days out from Fen OPS X in San Diego.
And it's important, Lynn, to start it's lens to start exploring as we're hearing more and more about this concept.
Quick orientation on the term, the token economy is what economists and tech leaders and CF OS and a bunch of thought leaders on the Internet are starting to call the new shape of enterprise software costing.
Now I've had a bunch of conversations with execs in the last few months and this term has come up in all sorts of interesting formats from people I didn't expect.
You don't really know much about this space.
And every interaction that I've been having, whether it be with them or in our community calls or or working groups, is that AI costing and allocation is consuming conversations.
And in that process to the token economy, it's also consuming a lot of tokens.
Now tokens are billed, tokens are metered.
They have become a unit of value and a unit of measurement, a unit of cost that shifted from seats and instances and VMS and databases to tokens consumed and value produced per token.
Now they are, to borrow a line from the last podcast, inelegant, but an accurate way to record consumption.
Inelegant because they're actually just a portion of the cost of the token economy, and there's a portion of the way that agents create costs.
But this pricing model change tied to the idea of tokens is one of the biggest shifts in enterprise software pricing in decades and adding a lot of variable based consumption to all sorts of spend.
For most of the cloud era, fin OPS is all about that variable based consumption and against predictable demand and predictable supply of resources.
You knew how people were largely going to use the software and you could estimate how much compute they would use, and you worked around various levers like reservations, rightsizing negotiated agreements, and managing unused capacity.
The token economy breaks both halves of the equation.
Demand for tokens is not predictable even by those writing the software, and the consumption per request varies by orders of magnitude depending on which model you call and how the agent handles its own context and context windows.
More on that in a minute.
The job is changing.
The job of fin OPS.
That is because the math and the inputs underneath it is changing.
So back to the original story from the cold open.
That thing about the scarcity and the scarcity of memory and all these things that lead up to creating tokens.
This is a piece that nobody was talking about a year ago.
All we were hearing about is all the hardware open AI was building and the agreements they were doing and everybody's talking about it.
Now there's rumors of of scarcity for Anthropics ability to service some of the upcoming models.
JP Morgan put out numbers last month and they said orders for GP US are a trillion dollars to 2027.
Memory for said AI stacks is sold out for all of this year and data center vacancy globally is at 1%.
Ninety 2% of capacity under construction data centers is already pre leased.
A modern data center takes a couple years to build and the power to run that data center requires five to seven years, 10 years to put in a nuclear in time.
So JP Morgan added a line that although we're unlikely to run out of compute, we are sitting in a market condition that will increasing the reflective supply constrained environment translation.
The discount era is ending.
Pricing power is moving to providers of scarce inputs, in this case production of tokens.
Procurement and margin management are going to become disciplines that are critical in the coming years.
Tom Canton Goose, an analyst and writer, talked at the Clean about the He gave the cleanest summary of where all this comes together.
He said there are 5 hallmarks that are defining this era.
First, state-of-the-art models are becoming relationship based rather than Open Access and in that he was talking about procurement and purchasing.
I heard a podcast over the weekend that encouraged companies to quickly set up an enterprise agreement with Anthropic to get access to more tokens.
AI is going to the highest bidder when capacity is constrained.
We're hearing about model prices increasing, not decreasing.
Available but slow is the term we're hearing for new models because even if you can pay for them, there's no guarantees the model will be fast because they are token constrained.
Inflationary commodity pricing as demand compounds against fixed supply.
IE pricing is going up with basic inflation.
What is the word?
I'm looking for basic inflation math.
We'll go with math and force diversification across different types of models and different locations to get models on Prem hyperscalers build by lease.
It's going everywhere from small models, less frontier models to on premise deployments.
The open AICFO put it on the record in their last earning calls.
They said their team is making tough trade-offs on whether they will pursue what what they will not pursue, rather, because they simply don't have enough compute to service everything.
And Anthropic has limited Mythos, their newest offering, to roughly 40 organizations, ostensibly because of security concerns.
But there are also questions out there about whether they could service service broad use of Mythos with their current compute.
So access to the bleeding edge is becoming a gated privilege.
Key Differences: Token Economy vs. the Cloud Era
So what makes all this different?
Well, there's a few things here that differ from the cloud era.
The 1st is the asymmetry between input and output needs and costs.
In an agentic session, the model sends its full context as input on every turn.
Every request sends that full context.
That means the system prompt, every file the agent has to read, every edit it makes, every error message and the full conversation history.
On the other side of that, the output is fairly small, but the input grows with every turn, every conversation, every change, every response.
The numbers tell a pretty dramatic story.
If with A50, a typical like 50 turn agentic coding session, that can quickly become a million input tokens while only 40,000 output tokens, the ratio is about 25 to 1.
And this ratio is the thing that makes a gentic sessions different than anything else your team is doing with AI, the sync.
The second thing that makes a difference is the difference in the price spread between the models.
The gap between the cheapest model and the most expensive model can be up to 10 times per token in agentic workloads.
That ten time application on every turn is an accumulating at a massive context size.
If you scale that to just a 25 person engineering team against two significant agentic sessions per developer per day, you're talking about 1000 sessions a month on a high end model.
That team's burning through 72 grand a year.
And that's for a tiny team who's really not doing that much.
I'm hearing stories out there from companies quoting companies 8090 people spending five, six, $7,000,000 a year.
I'm hearing back channel from bigger companies who are already spending hundreds of millions of dollars per year related to token costs either buying them directly or generating them themselves.
So think about how this scales.
Think about how this scales to an organization who might have 10,000 or 5000 or 100,000 people in their organization.
And now it's not just developers who are using agentic.
Another thing that makes all of this token economy different is something we talked about at the top supply scarcity.
The cloud era had kind of abundant supply.
There were very few organizations who are hitting capacity limits.
Reservations and rightsizing were the levers because waste was the enemy.
Now there was capacity reservations and you know, if you're a really big org and you need 10,000 of something, you sometimes had to wait or move regions.
But token in the token era, all the capacity is constrained.
Procurement leverage, IE negotiation, relationships, agreements and capacity strategy are becoming a new lever because access is now the enemy, not necessarily waste.
Can I get enough tokens?
And you hear about anxiety from individuals and organizations about not being able to get the tokens they need to do the work they're doing.
The New FinOps Role: Strategic Shifts and Career Path
So the discipline if an OPS has to move with this new math.
I was really excited to see last week CVS posted a job or maybe it was earlier this month for a new role and it's a director of AI OPS engineering.
It's a Greenfield role, Greenfield set up and the higher really is defining an operating model and engineering standards that will govern CV, s s, AIR infrastructure for years to come.
The role reports to the global head of infrastructure and AI operations.
Common theme we're hearing it's got 9 functional parts.
It talks about platform reliability, infrastructure, network observability, security, SRE, 24/7 operation centers, change and release management, all these things.
And it's all part of fan OPS.
It all uses Fen OPS as a substrate to connect these various business and engineering processes.
And where it starts to get interesting is we're seeing jobs like this talk about GPU cost governance, utilization optimization, tenant quota enforcement and charge back models in partnership with their IT finance friends, colleagues over in finance teams.
And this is a dedicated Fen OPS function at this particular organization.
This one was CVS at a healthcare company and we're seeing this common theme align where these roles are sitting next to platform reliability and security as a peer.
Now the requirements list for these types of roles are going to be, yes, demonstrated fan OPS experience over years.
A lot of people have that.
But now GPU cost optimization experience as a must have, not a nice to have, a must have.
Do you have listener, dear listener, GPU cost optimization experience.
That's what going to become a new must have and job requirements in this particular role.
They were also looking for years of SRE leadership experience with Kubernetes, experience with open shift GPU computing, death depth.
This is the new shape of the role.
Now this one does read as a lot more technical than some of the ones we're seeing, but I also know at this particular organization, they have a well fleshed out and very mature than OPS teams with a number of other individuals there.
So this organization is not hiring a cloud analyst role.
It is not hiring a budget owner role.
From what I can read from the job description, it's a senior engineering leader who understands GPU economics, but also might have the ability or the need to negotiate procurement with hyperscalers.
And they might also be looking at chargeback models touching on the business and finance side and they need to lead A-Team.
So this post was just an example of what we're seeing as a trend, but it was a strong signal and we're seeing similar ones come out and other fence serves in healthcare and financial institutions in any of the large orgs that are ahead of tackling this massive problem.
And the pattern is consistent.
Fenops is no longer reporting up through engineering, simply talking to directors, managers.
It's now sitting at the top of the AI operations conversation in a fairly senior role with peers and security and reliability typically rolling up into a CIO or CTO.
So couple of implications to think about.
First one is team capacity.
If your phenoms function is size for the cloud era is likely undersized for the token economy, the work is harder, the cost shape is more volatile, the procurement leverage is more contested.
You may need to reallocate capacity inside your team.
You may need to add headcount.
I said that adding headcount in the age of era of AI.
The second one is skills mix.
The skills makes it matter in a cloud era for Fen OPS are necessarily not are not necessarily sufficient.
Tagging, yes.
Allocation, yes.
Commitment management, Different shape, right sizing, yes, but different.
All still real.
But the practitioners showing up at Fen OPS X in the stories I've been hearing have real systems that are built in GPU economics and model selection, trade-offs and agentic cost patterns and the procurement conversations that are coming with supply scarcity.
That skill mix is what these new senior roles are asking for.
The third one is positioning.
The discipline of fin OPS is moving up.
We've been talking about this for a while. 24 months ago it was simply something that happened within a director level of engineering, sometimes in finance. 18 months ago it started being consistently reporting up to the CIO, and now it's sitting inside that core AI operations conversation.
If you're a fin OPS practitioner today, the make it a break question, to quote our buddy Rob Martin from Episode 7, applies here directly.
This token economy is the moment the career move is to go up into the AI operations conversation.
Not sideways, not to sit where you are, and certainly not to think that the skills of fin OPS from 2025 are going to let you grow and scale your career.
The number one thing driving all of this is the speed of change. 98% of fin OPS teams are managing AI costs now.
That happened over 2 years.
These teams need to manage inference.
They need to manage agent prototypes.
They start looking at modest token bills.
But the change is coming dramatically.
And that second wave agentic coding is what's going to drive this $1,000,000.
GPU contracts are showing up all over the place.
AI workloads are needing to be Co located or hybridized in an on premise environment to make the demands available and possible.
Every one of these shifts is happening in front of the fin OPS owner.
The years we did spend talking about scopes and technology categories and data center and SAS and data clouds are all converging at this moment.
And This is why at the conference, all of this, this year in 28 days matters more than ever.
The community is going to move now through the biggest transition it's ever experienced very quickly.
The next 6 to 12 months are going to be a real change.
SO3 Quick things to take away 1 Token economy breaks the assumptions that the cloud era was built on.
Demand is not predictable per token costs may be 10 times different between models.
Supply is constrained to the Fen OPS job is bigger than the old Fen OPS jobs.
Yes, add on new types of technology, GPU cost, governance, procurement, leverage is more important than ever.
Chargeback models are more complicated.
Tenant quota enforcement, yes, all of these things, but in addition, managing AI and applying AI to fin OPS and understanding how cool agents are going to drive costs and getting ahead of that.
And three, the career move is to go up into the AI operations conversation.
The discipline has moved.
A lot of you were doing cost optimization.
It's got to be more about strategic procurement capacity then OPS has always been about value.
If you're still thinking about cost optimization, you're behind.
We don't want any of you to get left behind, which is why we're leaning into all of this.
OK, so that's going to do it for today's daily brief, mostly daily brief, 28 days to fend up sex.
The conference is the place where we're going to lay out the new shape of this discipline and where it's going to get defined over the next year.
So come hear from the practitioners who are doing this in person.
I reached out to a practitioner yesterday at a massive well known company and I asked him if he were he was hearing about these themes and he said, quote, it is all I'm hearing about in my daily life right now.
So if today was useful for you, share this episode with somebody or put it on LinkedIn or hit that follow button on Spotify.
I'm Jr.
Stormont and I will talk to you soon.
Podcast Summary
Key Points:
The era of abundant AI is ending due to supply constraints on memory chips and GPUs, with pricing power shifting to providers.
The "token economy" replaces traditional cloud pricing models, where costs are based on tokens consumed rather than seats or instances.
Demand for tokens is unpredictable, and costs vary by up to 10x between models, especially in agentic workloads with high input-to-output token ratios.
Supply scarcity is driving model access to become relationship-based and gated, with increasing prices and forced diversification across models and locations.
The FinOps discipline is evolving to focus on GPU cost governance, procurement leverage, and strategic capacity management, with new senior roles like "Director of AI Ops Engineering" emerging.
FinOps practitioners must move into AI operations conversations to stay relevant, as the job shifts from cost optimization to value-driven strategic management.
Summary:
The age of abundant AI is over, marked by a shift from abundant cloud resources to a supply-constrained token economy. Memory chips are sold out through 2026, and GPU orders are projected to reach a trillion dollars by 2027, with data center vacancy at just 1%. This scarcity has ended the discount era, giving pricing power to token providers.
The token economy fundamentally changes enterprise software costing, replacing traditional metrics like seats and instances with tokens as the unit of value and cost. , 25:1 ratio). This volatility, combined with constrained supply, forces organizations to prioritize procurement leverage, capacity strategy, and relationship-based access to state-of-the-art models.
The FinOps discipline is transforming accordingly, with new roles like Director of AI Ops Engineering emerging that require GPU cost optimization and strategic negotiation skills. FinOps now sits at the top of AI operations, reporting to CIOs and CTOs, and practitioners must move into this conversation to advance their careers. The key takeaway is that FinOps is no longer just about cost optimization but about managing value and strategic capacity in a resource-limited environment.
The next 6–12 months will see rapid change as organizations adapt to this new reality.
FAQs
CVS posted a 'Director of AI Ops Engineering' role that requires GPU cost optimization, SRE leadership, and procurement negotiation, reporting to the global head of infrastructure and AI operations.
Because consumption per request varies by orders of magnitude depending on the model called and how an agent handles its context windows, making it hard to estimate usage.
In a 50-turn session, the model sends full context (system prompt, files, edits, history) as input on every turn, leading to 1 million input tokens versus only 40,000 output tokens.
They may fail to address GPU economics, model selection trade-offs, and agentic cost patterns, which are now essential due to supply scarcity and volatile token costs.
The gap can be up to 10x per token, and with large context sizes accumulating per turn, even a small team can burn $72,000 annually on a high-end model.
Even if you can pay for a model, there's no guarantee it will be fast because token production is constrained by limited compute resources.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.