Go back

Everything You Need to Know About AI Tokens

50m 30s

Everything You Need to Know About AI Tokens

The episode serves as a comprehensive primer on AI tokens and their economics, addressing growing corporate anxiety over AI costs in the agentic era. It begins by explaining tokens as text chunks (typically smaller than words) that models read and write, with tokenization varying by language—non-English text can cost 2-5x more—and code, where indentation and brackets add tokens. The hosts trace token consumption through four phases: token oblivious (subsidized, flat subscriptions), token maximizing (leaderboard-driven, like Meta's internal usage of 60-74 trillion tokens monthly), token anxious (employees self-censoring prompts to avoid costs, as seen with Uber's 1,500-token caps), and the desired token smart era, where spending is wise, not sparing. They stress that the most expensive token is one your best employee fears to spend. Key insights include that agentic tasks consume 5-30x more tokens than simple chat, with 60% of costs tied to refining answers; tokens are not equal across providers due to different tokenizers and pricing; and model updates can inflate token counts (e.g., Opus 4.7's 30-45% increase). They highlight three token layers—input, reasoning (invisible, 4-20x cost), and output—and warn that long sessions compound costs as history re-sends. The core recommendation is to shift focus from per-token price to cost per accepted task, auditing usage to avoid leaks, and governing AI adoption to prioritize value-generating use cases while preventing accidental overspending, like a yes/no question triggering a 4-million-token deep research.

Transcription

9553 Words, 52306 Characters

English
Today on the AI Daily Brief, an operator's cut up episode with new far, everything you need to know about AI tokens. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Oh, right friends, quick announcements before we dive in. First of all, thank you to today's sponsors, Rackspace, Blitzie, Sexy and Air Table. To get an ad free version of the show, go to patreon.com/af or you can subscribe and up a podcast and to learn more about sponsoring the show, send us a note at [email protected]. All right friends, well, new far gas bars back today and new far and I have been cooking up a lot recently. A whole slew of you have done our most recent program, have explored our most recent program, the Choose Your Own Adventure Style AI Summer Adventure, plus we've been cooking up on expanded set of educational resources that we'll be telling you about soon. But one of the realities that both do far and I have been living in is every company we interact with dealing with the same questions of AI tokens and token economics. We are now firmly in the agentic era of AI where companies have to think not only about how to get adoption and how to maximize AI's value but how to do so in a way that doesn't just totally break the bank and where the right intelligence is being used for the right problems. Anyone who's ever built an open clock and tell you that getting the right models to do what you want them to without going off into endless cycles of spin takes some real consideration. Today's episode is designed to be the ultimate primer on AI tokens. What we're talking about when we say that term, what the new challenges are and some of the key pitfalls to avoid as well as strategies to maximize the way that you and your company use AI tokens. All right, new far back with another operator's cut talking about the topic to your, the topic on everyone's minds. We are talking tokens. How are you doing? I'm good. Very psyched to talk about tokens. Yeah, I think it's a, I love this period in a discourse where we've gone from sort of pulling hair out, freaking out about the new change to actually settling into new tactics, new strategy. And I think this is a perfect fit with that. So tell us a little bit about what we're going to be talking about and let's dive in. Good. So the reason why I wanted to do this episode is because every one that I walk into these days, like literally, everyone have some version of the same token conversation. Some practitioners feel like they're being watched when they use an expensive model. The regular users wonder whether one ambitious prompt will eat their weekly allowance and the leadership teams. They see a bill growing faster than expected. And then they start asking a lot of questions on whether all of these tokens produced anything useful. So I don't know where you guys are sitting, but there is a new anxiety around using too much intelligence. And I actually want to flip the conversation. First of all, to make sure that everybody understands what tokens are and what the bill actually means. And then how to spend them wisely rather than sparingly. So that's why I'm here and what I'm planning to do today. One of the places that I found myself with this conversation is there's been such a visceral reaction now as the cost has gone up. I have a bigger concern around people retreating back to known ROI biases and not boring, but ultimately low stakes use cases, let's say, as a compared to what AI can actually do that I found myself in the position of having to defend things like token maxing and token leader boards, just relative to where the Tony shifted. I think obviously we'll get into today the smarter version of that conversation. So I'm excited for it. Exactly. All right. Because there is like a going conversation around that, I think that there is a very like a better language around the feeling of tokens shouldn't be just a financial thing. And just recently the open ICO proposed the scorecard and it was called useful intelligence per dollar. And that's built around the one question, what does each successful task actually cost? And that's the conversation that I think people should have. And I want to help you with the whole story. What are tokens, what your work costs and where usage creates value and where it quietly leaks value other than adding. So to kick us off and how we got in here, I want to walk you through four errors of token consumptions and probably we'll recognize where you are. And we all started by being basically token oblivious. That was the all inclusive era where model companies subsidized the usage and the flat subscriptions, he'd the meter and many individual users and many of them are still there. Just see the ceiling not a per token price. So that's where we started. And then like you said, we got into the era of token maximizing. That was the leaderboard era where usage became the badge of I maturity. And we all remember some of the conversations around meta who tracked employee I usage on an internal leaderboard. They used to call it tech and they used roughly between 60 to 74 trillion tokens in a single month. And according to the data that was published, the top individual user used 280 billion tokens. So to give a sense of how many this is, that is roughly 2.3 million books worth of text. So if you want to try and imagine that that's about 50 books every minute, continuously for a month. So that was the meta store and then overlaunched also an adoption leaderboard and burned through the entire 2026 AI coding budget in about four months. And there was another company unnamed, but according to TechCrunch, they ran up 500 million dollars of cloud bill with no usage limits in place. So then in that era usage became the metric and dashboard measured activity while claiming to measure value. Obviously this was unsustainable and I know you have some opinions on that, but I'll be curious to say to hear your points. But I also want to say that it actually got us to be as always the pendulum took us way too far to the era that I call token and just which is where we are. But if you have anything to say in defense of a leaderboards, I'm here to listen. Yeah, look, the defenses less leaderboards are a great concept. They come with I think a set of very predictable challenges. In fact, so predictable that I would say that the hand-ringing around the idea of people gaming them has always struck me as a little absurd. Of course, people are going to game systems if you put real stakes around them. But that's pretty predictable and also fairly from first principles, you could figure out a lot of ways to deal with that. So I think one, it's overrought all the sort of people who are freaking out about that. Secondly, my bigger point was that a company that wildly overspends right now via a token leaderboard or anything else, I will bet any amount of money that they will be farther ahead than a company that underspends because they're overly concerned with proving out ROI or whatever it is on a sort of year time scale. Now the Goldilocks scenario, which I think is what you're going to get into, is being able to experiment, being able to learn, being able to build, and being able to actually understand consumption while also not being afraid of it. But I agree. I think that the pendulum swung so aggressively far too aggressively back from token maxing an excitement to token anxious, and that's the paradigm that we've been living in in the recent few weeks, a couple months, whatever it is. Yeah, good. So I think you love my model of how to use tokens wisely. But with regards to being token anxious, what I'm seeing in many companies now is that many employees are self-sensoring themselves, basically trying to avoid costs. And even if we look at the same companies that were token maxing, so meta went from the leaderboard to sending a memo that constrained the AI usage. And now the press is calling it token minimizing instead of token maxing. And UberCaps employees at 1500. So even the same companies who are token maxing are now significantly shortening that. And when employees self-sensore, it gets them to feel like every prompt is an AI conversation. And that's not something that we want to have. And I think that this is a very bad era to stay in because I believe that the most expensive token is the one that your best person is afraid to spend. So this is where I want to direct all of us to being. And I call it the token smart era, meaning that you need to spend wisely and not sparingly and understand what creates value and where usage quietly leaks. That's entire kind of theme of this episode. And what I wanted to go by in this era, in this case, is to talk about the four elements of what a token actually is, why tokens were not born equal, how to audit your own usage and how to govern or do it better within your company. And I'm trying to make it relevant to anybody, whether you are the practitioner that needs to apply some cost engineering playbook and be smart about that or the executives and admins that need to be proactive and avoid having a difficult conversation with the CFO without a proper response to let's just cut the bill without talking about the business implications as you just said. So that's the plan for us today. So I want to start with introducing you to the token because I feel that even though it's the most used term in AI, not too many people truly understand what it is because that's in the root of every bill quota and rate limit with AI in tokens. So in a simple word, token is a chunk of a text that the model reads and writes. It's typically bigger than one character and it's usually smaller than a word. And if you have never ever seen a tokenizer or how tokens look in action, open AI has a very good page that is open to everyone that you can just take a look at how tokens actually look. So it looks something like that. You can paste the text and then you will see how words are being chunked. So you can see the some words are stayed staying like as one token while others might be separated into multiple tokens and interestingly, numbers often are being chopped in the middle and so on. So we'll put it in the show notes, but a very interesting experiment if if you have never seen how you'll text. looks and by the way if you paste a non English or non Latin language you will see that typically the amount of tokens is much larger than an English language. So that's the open-air tokenizer and a few things just to lend it home. In general, like the ratio in English is around three quarters of a word to token, meaning that if you have a page of text it's roughly one thousand tokens and some languages that are like in the Thai Greek other languages like that might get two to five X more tokens for the same content and because billing is per token then some questions if you ask them in other languages might cost you much more and that sometimes refer to as language tax with AI with code also it's different and it has its own way indentation and brackets and white spaces they all become tokens. There are some newer ways to tokenize text that are more code friendly in order to do that but still numbers is a huge problem so you've seen the one two three four five being chopped in the middle and by the way that's also why whenever everybody's doing like the strawberry test for AI and it very badly fails in trying to count how many hours are in the word strawberry. In many cases that's just a tokenization feature rather than a failure and the model just has never ever seen the individual letters it just saw the straw and the berry as separate words and that's why it's counting it off so model models have various workarounds but many of the like AI's so dumb memes are literally just tokenizer issues so that's tokens in terms of what everyday work costs I think that's a good kind of mental model to have so for example drafting an email is around 500 to 700 tokens a page of text has noted is about 1000 tokens you can see a longer text it can be more than that if you send a model or a tool to do like a AI web search often it will add a few thousand more tokens for the results sometimes much more images interestingly are in many cases not that large in terms of how many tokens they're roughly around slightly more than 1000 tokens interestingly deep research can very easily be 70,000 or hundreds of thousands of tokens but just the other day one of our learners in one of our courses I had a yes no question and accidentally instead of asking for a web search for the problem he was asking for do like the genetic tool to do a deep research the tool spawned about 100 sub agents to do the research and then his yes no question cost over four million tokens just to answer this question so it can very easily amount to much more than that specifically a few additional places where you can find very token heavy workloads will be data analysis that can easily get to one million or more tokens per task and heavy coding can also be very aggressive similarly with many agentic working flows so just to give you a sense of where it is and to be a little bit more concrete here the everyday stuff as you've seen like the emails and so on is almost free so it's around half a cent and nobody should ration emails it's not where the money goes search and research can multiply very quietly so that can be a place to look for efficiency and the top of the ladder that's a completely different spot so if you compare like email to agentic coding task it can be a factor of a thousand or even more and another thing that you need to pay attention is that every conversation compound so the model doesn't remember your previous messages and as such it sends all of the previous conversations within the same session back to the model so by let's say turn number 10 it may be processing so much of the earlier exchange alongside your new message that the total goes much faster than the number of turns suggests that even happened before the system point and we'll talk about strategies in later on but this is one of the things that can very easily just having very long sessions can very easily amount to a ton of tokens being consumed I think this is one of the reasons why this is such an important conversation is another way to put this is that the more advanced and ultimately higher value use cases consume more tokens which is intuitive that more intelligence is required for bigger challenges but the direction of use cases is proceeding this way and so the reason that the token anxiety is going to create problems if not addressed is that it will incentivize people to stay swimming around less sophisticated use cases so this is the trajectory is clear in terms of less token consumption the you want as a leader group your people to be doing more advanced more useful things with AI it's just how they do it well so you want them to do deep research where deep research is required but you don't want them to accidentally do a deep research on a yes or no question that they can google in a second good so speaking of the agentic where the more advanced capabilities those can significantly grow the amount of tokens because agents work autonomously in loops and as such they consume by very widely cited industry estimates five to thirty times the tokens of a simple chat and poorly designed a agentic loops or agentic harnesses can be even worse than that because a typical task involves between 10 to 20 model calls caring instructions and history and tool definition and previous results and I think according to McKinsey they estimate that the roughly around 60 percent of an agentic tasks cost is tied to the checking and refining and the like regeneration of the answers after the first response so the expensive part is often getting from the answer to the accepted results so that's an interesting one and now it gets even more complex because tokens were not born equal so by the way the point here is not to not use the agentic tool just to know that in as you said intelligence cost but it gets even more complex because tokens were not born equal and every model lab has its own tokenizer you've just seen the open AI but different model labs have different tokenizer so for example the open AI current tokenizer has a vocabulary of about 200,000 tokens Gemini has around 256,000 Lama Bimetta has about half of that and Claude is unpublished and the reason why we all should care is that the price per million tokens is denominated in each lab's own tokens and often we don't know them and the same document can be 10 to 20 percent more tokens on one provider than another and even more so for a call the non-English text so the model behavior widens the gap and one model may answer in a single pass while another reasons longer and writes more and takes more agentic steps or needs retries and the tool around the model adds its own system and context and the look design so the two stacks doing the same task can have different token counts and different completion rates and as a result it's completely different build so the per token price is kind of the sticker but the cost per accepted task is the operating metric because otherwise there is no way for you to compare between different providers and different tools. All right an important story that also illustrates that what happened when opus 4.7 came on board the tokenizer basically under the hood changed and it was this April and the price sheet was identical to the previous model the same dollar per million token but the model was shipped with a new tokenizer that produced by interpix on the commentation they didn't hide it roughly 30 percent more tokens for the same text so there were quite a few independent analysis of over a million requests that found native tokens they grew and the count grew by about 32 all the way to 45 percent and the real world builds grew by 12 to 27 percent because some of the difference was absorbed by caching even Simon Wilson he measured one of his own points at around almost one and a half x more tokens so even though it was documented the fact that we paid more for the same intelligence and this is like a shrink flation right the same sticker price but a smaller candy bar so nobody prints now 30 percent for your words per dollar which is the case that happened there so that's something that is constantly changing every lab turns the tokenizer and often for good reasons but the operator lessons here is that we have to talk about dollars per task and not dollar per token because the budget is like a moving denominator and it's not the way for you to try and understand how much is gonna cost let's talk about what tokens are used for by the AI tools and you have to understand that every AI request has three token layers and they are priced very differently we have their input tokens those will be their prompts and the conversation is three hundred files and the tools definition and everything that is part of the input this is what the model reads and this is the cheapest per token but can accumulate fast because if the history is being recent or if a lot of context is being read that can cost quite a lot then we have the reasoning tokens that's the second layer these are the tokens being used for the model internal thinking before answering for the most part it's gonna be invisible to you but it's built at the output rates meaning at the high rate of per token cost and those can add between 4 to 20 x cost per request and finally we have the output that's the answer that you actually see and this is typically 3 to 5 x more expensive than the input price per token and I think the reasoning layer is the one layer that catches everybody by surprise because you might have a 400 token answer but under the hood it carried like the I don't know 4,000 thinking tokens underneath because the model was having an internal monologue and doing a lot of thinking in order to give you the answer and if you want the analogy it's like thinking about the part of the restaurant build that is labeled the kitchen time so you don't get to see it it's not part of the dish but you still have to pay a lot of it for that and the models with the high reasoning effort are often the one with the 20 x amount of tokens being consumed versus the long run. reasoning efforts, it can be the same question with a significantly different price tag. And sometimes spending a higher reasoning does not get you better results. So some metrics even showed it for simple questions. It's better to use lower reasoning because the overall cost per task will be significantly lower and the quality will be improved without having the model overthink everything. So it's not always that smart or spending more time thinking gets you better results. One of the more interesting shifts in enterprise AI right now is how quickly the conversation is moving towards infrastructure and operations. As AI moves into core workflows, regulated data environments and agentic systems, enterprises need governed infrastructure and inference that can operate reliably day-to-day with clear operational accountability built in from the start. As those systems scale, the operating model increasingly becomes part of the AI strategy itself. Rackspace technology is the operator of the full enterprise AI stack from agents to infrastructure across private cloud, hybrid cloud and edge environments. Rackspace builds and operates governed AI infrastructure, inference and production AI systems for organizations where sovereignty, compliance and uptime are non-negotiable. Therefore, deployed engineers stay embedded beyond deployment to help operationalize and run AI in live environments. To learn more about where enterprise AI runs and outcome scale, go to Rackspace.com. Every AI coding tool on the market does the same thing first. It starts writing code. Let's eat. Does the opposite? Before writing a single line, Blitzie spends days reverse engineering your entire code base. Thousands of agents ingest millions of lines, mapping every dependency, every undocumented constraint, every architectural decision made over the last decade. The result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building. Other tools guess it context with grip searches and markdown files, Blitzie never guesses. It builds true understanding first, then delivers over 80% of entire software epics autonomously, validated and to end tested production grade pull requests. That's why Fortune 500 engineering teams trust Blitzie with the code bases that matter most. See for yourself at Blitzie.com. That's B-L-I-T-Z-Y.com. Here's a harsh truth. Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized. Half of companies have AI tools, but only 12% use them for business value. Most employees are still using AI to summarize meeting notes. If you're the one responsible for AI adoption at your company, you need section. Section is a platform that helps you manage AI transformation across your entire organization. It coaches employees on real use cases, tracks who's using AI for business impact, and shows you exactly where AI is and isn't creating value. The result? You go from rolling out tools to driving measurable AI value. Your employees move from meeting summaries to solving actual business problems, and you can prove the ROI. Stop guessing if your AI investment is working. Check out section at sectionai.com. That's SEC-T-I-O-N-A-I-D-I-D-C-C-A-O-N. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. New users get $1,000 in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing's agent turns competitor moves into landing pages. Sales is agent and reaches leads, drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at HyperAgent built by the team at AirTable. Claim your $1,000 in inference at hyperagent.com/aiDailyBreath. All right, and I think that the gap between input and output pricing keeps widening at the frontier. If you look at the Fable 5, it's about 10, not about its 10 dollar per million input tokens and $50 per million with the GPT 5.6 solids 6x ratio. So we're seeing the gap even widening. And those effort levels, that's also something that highly adds the complexity because these frontier models increasingly letting you dial the reasoning effort. With higher efforts, the reasoning tokens are significantly higher. And that's probably the dial that you should even be more mindful of even beyond the models because those can easily cost you 10 to 12 x token increase between high or extra high effort to the lower medium. All right, one last thing here, price tag that experimented that came from Databricks because a smarter model might not always be more expensive than a less expensive model. What Databricks did is they tested coding agents on relangenering tasks from its own codebase and they were using Sonnet 5. It was 1.7 times cheaper per token than Opus 4.8. However, Sonnet cost $2 per task or $2.09 per task versus $1.944 Opus. So because Sonnet needed more iterations and more reasoning had to spend way more tokens to get to the same results. Overall, Opus, which is significantly on paper, more expensive model, it was cheaper to operate, which means that we shouldn't just reach to the cheapest model possible. We need to reach to the right model for the task and that's not easy to get, but something to be mindful. The other thing that matters to the bill is the tool itself. So Databricks in their same experiment ran the same model at the same thinking effort to different agent harnesses and they saw that more than a 2x difference in cost per task with the same quality using different harnesses. Just because primarily one tool was feeding the model roughly three times less context than the others and thereby the overall cost was lower. So very difficult bill to read and very difficult bill to navigate and I'll try to help you as best I can. So bottom line we're dealing with cost per task and not tokens because otherwise we will not be able to actually compare apples to apples and the cost will include the retry, the review, the every like a collection that needs an every additional iteration and then you need to divide by the number of accepted results. That's your cost per accepted task. That's the metric that you should aim for and optimize for and this also brings the conversation much more into return on investment and business value rather than just having that conversation around tokens that is very hard as hopefully by now you understand to meter. Good. The practical test that you can do you need to take between 5 to 10 representative tasks of what you do. Run them to two model or two options in order to have a good understanding of in your option space what you should do. Hold the input and the quality bar very constant and compare first pass success attempts human correction and elapsed time and the total cost. The winner is the stack that gets your actual work done reliably. I know it sounds like a lot but if you have a good autonomy that was optimized for yourself and you know for the tasks that you do which models overall get you better results potentially with your tokens or if you can do that for your team or your company and you will need to do that recurrently because things change quickly then at least you can teach folks that if you're doing that type of research the recommended to model to get you to the overall best quality and the value per task is the following and so on that's the current reality that we live in. Okay now I want to give you a language on how to look at your own tokens and hopefully using that you will be able to distinguish between the tokens that add value to the ones that not so much and every token that you all organization spends in my opinion is one of three kinds. There are tokens that I call tokens that teach and this is running in both directions meaning that you're teaching yourself as the Daniel said before we don't want to stop the experimentation so the tokens that include the experimentation the failed workflows and let me try these three different ways so I will learn and so these are tokens that are worth spending because you get learning out of them and you can look at them as tuition and by the way those also include what you're teaching AI about yourself so those will be the identifiers and the curated context and the knowledge packs the memory can be counted as those because teaching your AI who you are what is your context and learning what works for you and AI is very critical for you to continue moving forward they look a little bit like a waste on a dashboard or a lot like a waste on a dashboard because no deliverable ship but I claim that these are tokens that you need to defend fearlessly because if you want defend those you will very quickly go back to just help getting AI's help to draft emails and translate between languages rather than moving towards the workflows that matter and especially if your company and yourself has a lot of ketchup to do on where AI is currently at so I want the tokens that teach to be defended because those are the the things that will move the needle beyond the next category which I call them the tokens that produce so obviously those are the most defensible ones because those are the tokens that use to create work the chips it can be the like the final proposal or the research or the code obviously that's the thing that is much easier to show the ROI but lastly we also have the tokens that should be eliminated and those are tokens that I call tokens that spin those can be machines talking to themselves or automations that nobody is looking at their output or automations that are running too infrequently idle agents bloated context misused tools and context using fabled to item email so using the wrong model optimized flows and so on those will be activity without sufficient output so the tokens smart move if I need to summarize is to kill the tokens that spin to tune the production to make sure that it is cost effective and protect the teaching and that's the order like first go and do the audit or on your spin tokens and then do the rest. I have a very embarrassing tokens that spin story, which I will share in a minute. But I do want to also note with regards to tokens the teach that a failed experiment is as important as a successful experiment. So you should definitely encourage your employees to fail to try because otherwise, the tokens that produce will not yield as much value as possible. So it's embarrassing as an AI expert to talk about it, but my open claw was a chief of staff was because it's currently disabled, a chief of staff that I called Chloe. And it was using the Anthropic API. And because it was using an API, it was like auto renewing all the time. And the bills were sent to a secondary inbox. I wasn't really monitoring them. And I was seeing that the charges seemed quite high, but because I was getting a ton of value and because I was not paying attention to how frequently I'm getting a new bill, I wasn't noticing. And then early June, I was traveling. So I was not using my open claw at all. And still I exceeded the bill kept coming. So I was saying, why am I still getting some bills? So I opened the dashboard on it to realize that I spent in two weeks, $1,500 on an agent that I was not using. So I opened the dashboard like double clicked and I realized that I have almost 400 million tokens in and almost zero tokens out. So it was a ratio of almost 3,000 to one from input to output. And that's literally the definition of a machine talking to itself and billing me for like an internal monologue that it was running with itself. Looking further, there was a bunch of core and jobs that the open claw created for itself. And it was like a compaction job that ran every 30 minutes on empty sessions. And even worse, like the trend was going up. So I first of all closed my open claw and only to optimize it differently. But if it happens to me in this setup, it can happen to literally everybody. And especially when the credit card is owned by your company and not by yourself, often you will not pay attention but you're not sitting on the billing. - Yeah. And also the fact that you have a bunch of other things that are working well that you might assume. It's those things that are amounting for the cost. So one thing that I wanted to mention with spin is that I think a lot of the framing of spin, if people were to pick this up, they might assume that it's only mistakes or errors that produce that spin. But that's not always going to be the case. Like sure, this is sort of a in-between example where it wasn't exactly an era because it was doing something that it was meant to but you weren't really paying attention so it was doing more of it than it needed to. But I think a lot of times spin will also be just ill defining the parameters for a job that you actually do want. An example of this that I had is I had an open claw going for a while that was perpetually researching new data sources in AI that could help us figure out where the state of certain adoption metrics was. Every day there's new studies that come out that measure this or measure that and that tell you about data readiness or systems integration or use cases or whatever. And it's too much to monitor for humans but agents are really good at it. And so this open claw agent was a researcher that its only job was to set schedule based on a heartbeat. Go out and check for new things. And it was never meant to stop. It was always, it was on a specific schedule but it basically was this continuous research process that was crawling to the ends of the internet every day and it ended up just not being valuable enough for the cost but it was doing what it was supposed to. And so I think part of the auditing spin is also just figuring out what things have accidentally become spin even if they started in the right area. And I think that's why this idea of auditing I think is a good framework because sometimes it's going to be about just updating or changing a process that was valuable as well as catching mistakes. I agree. And even more, I see many automations that people created because they think they will be useful. Like, oh, I can't read. I have too many Slack messages. Let me just create like a Slack miner that runs every hour and reads my entire set of channels. And that can easily become $1,000 in tokens that literally do a job that moves the needle for nobody or the morning brief that you created wholeheartedly with the intention to read it every morning but for some reason you don't find value and you don't read it. So these are the things that you should definitely audit and kill. And my whole of thumb is if you created an automation and for one or two weeks you have never ever used the output, you should definitely kill 'cause it's a definition of a spin. Or if you are using this automation but there is a very bad proportion between the value of summarizing all of your slack channels to the bill at the end of the month, that's also something that I consider to be a spin. And I think that now that everybody gets co-work or GPT work that even more and more within companies 'cause it's so easy to build these automations and without sufficient literacy about how to effectively use the tokens people create a ton of these automations that look good on paper but don't look so great on the paper of the bill at the end of the month. All right, so let me give you a list for the suspects for silent token spenders. First of all, it's gonna be your idle agents and the over-frequent jobs. These are gonna be the things that run without any meaningful output or way, way too frequently. We also, in many cases, see automations that nobody uses so it can be like the weekly report or the dashboard that nobody ever goes to read. Additional thing can be what we refer to often as the pre-prompt tax. So anything the model runs and reads before the very first prompt. So those include the always on rules or instructions, the skill definition, the tool definition and so on. And those can very easily, if not properly organized, amount to many thousands of tokens each run without you typing a single word. So those amount significantly. Many folks also hold the immortal conversation, meaning that they will continue an endless session that keeps carrying old history and old context often if even creating a poor quality. We also, in many cases, see users never filtered data retrieval. So instead of just getting 20 rows from a database, they will pull 500 rows or they will process the entire inbox to look for a specific mail that they know what was the subject line and so on. Many other folks will have the context all over the place. So the agent will have to read through a ton of documentation just to understand what are they talking about and what's the truth here, as well as rework loops. So anytime that you're agent or your skill or your just day-to-day usage gets you to do more iterations just to get the same result, this is just more tokens being spent on nothing. So these are the immediate suspects. And I wanna show you how you can try to potentially identify whether your system is in a spin situation or that the spin-to-production ratio of tokens is not well-fomulated. And the thing here is that not everybody can detect in the same way. Some folks have concrete meta. Those will be people who are using the API version of the models they have the API console that they can use or if they are Claude code or cursor users, they have usage view. And of course people with admin privileges they have an admin dashboard. So if you are one of those, you can do the following things. One thing that you should definitely do is the weekend test, meaning that if you didn't do anything with the I but you look at your bill and you see that your bill keeps compounding, you know that there are things that are adding to your value without to your bill without any value. That was what's happening to me. Also look for very extreme input to output ratio. So agentic work legitimately runs with high ratio, but if you get to a point where it's many thousands to one between input and output, in many cases that's empty loops. In my open-close case it was 2600 to one, which is ridiculous. And if you see that you're spent keep rising while the work or the value that you do stays flat, that's also potential and indication that you're in a scenario of spin and you need to go and further understand what's the case. However, there are many folks that don't have direct meter because they're not using one of these tools or they don't have the admin privileges, which is probably most of the regular users. For them, you should probably use boxes. So just go directly to list all of your automation and the scheduled jobs that you own and ask which one of them added business value last week. If you don't know, that's a suspect. Then also watch your quota. And if you're burning to your weekly quota extremely fast, especially if you compare to other people in your setup or in your in similar roles, that might be that you're doing something wrong there. And if you're in an enterprise plan, your admin do have the view, at least of how much you're consuming and also the typically they input output, you can just ask them. And there are many places that you can look. There are specific like a slash context and slash usage in Claude code. There is also in application visualization now both in Claude code and in cursor that you can just click on the usage meter and try to understand that. So regardless of watch and how visible it is for you, you should definitely put some caps on how much you spend rather than letting the bill just extend all the time and put some alerts if there is some kind of a significant jump in how much you consume. That can be an indication that something is up in your system. So that's for identifying spin. And now the habits that we should all adopt to mind our tokens. These are several things that anybody can do immediately that typically improves the token consumption without reducing the business value. New task is a new session. This one is an interesting one because we were talking a minute about also model routers, but at least for now for the most part be intentional about which model you use for what task. Sometimes it's actually going up to like an opus or even fable class models because they will get the job done in one iteration and overall reduce the spend. In some other cases, it's not doing a web search with Fable, but rather going to the high-cool or the lower cost of models. Right size your context. Tell your eye what it needs to know. This is a classical Goldilocks, not too much, not too little, but sufficient such that it will not go into endless internal reasoning token loops just to try and understand what you're talking about. Build reusable capabilities often when we're just vibing with our model and trying to use it ad hoc, rather than sitting down and creating the skills, creating the proper automation, creating the proper agents. We're just wasting a ton of tokens to re-ask the tools to do something again and again. So, seeing and building proper systems often is one of the best levers that you have to use their tokens wisely and filter everything that you can. Tell it in which rows of the table the data exist, in which parts of the project board the data accounts for, which select channels and so on, the more you point the model to the right place, the better the results that you will get. And lastly, in many cases we start doing the work, we realize that the model is completely off. Maybe it's the wrong model, maybe it's missing something, don't let it spin. Just kill the job early and start again while understanding what you do. And this is one of the cases where looking at the model reasoning will go a long way to understanding that it's completely off in the wrong direction. So, I would recommend whenever you send the model to start doing something, especially if it's a significant portion of work, open the thinking to understand what the model is understanding from the task that you gave it. And if it seems to be off, stop and improve the instructions rather than letting it. So, that's the habits for everyone. Two additional levers that you should consider and some of them are very new. So, if you are a cloud code user, you can use this /doctor command. This will basically check not only how much like a past installation stay on your machine, but also how are you token divided? Whether you have stale skills, stale tool configuration, whether you're overall instructions are overly long or overlapping. So, it's a very good command that's top created for us that you can go and execute if you are a cloud user. If you're not a cloud user, you can just have your AI tool investigate what the /doctor command does and basically recreate it for your own tool. Because it's not a very complex thing to do. It just audits all of your system for you and gives you a structured report with concrete recommendations of things that you can kill because you haven't run them for a while or things that are duplicated or stale or contradictory that you can potentially use significantly. And with regards to routing. A lot of the industry conversations sit right now around the model routing and you were just talking today or the other day around some interesting M&A around model routing. And picking the right model is still one of the highest return things that you can do even if you are able to use like the cursor automated router or some of the other solutions that are coming our way. Because it's not always going to be even if you have like a router in the background, it's not always going to be as precise as you knowing which model to use. And in many cases, you still don't have in your existing tool a good enough or even an existing router. Yeah, I think that we are very early in figuring out the right patterns around routing. Obviously there are a million solutions coming to market. They're all taking slightly different approaches. You have independent experiments from enterprises who are building their own systems that route between custom models that they've trained as well as the premiere model. Like it is there's no one clear approach yet. And even when there do start to be clear use cases and patterns, they may not fit everyone in every use case. I think it would be entirely unsurprising to me or I expect that routing norms around certain types of software engineering get solved first because it's more deterministic and clear. You can kind of actually have more sort of verified success or not. I think when it comes to knowledge work tasks more broadly, it's going to be immensely more complicated. Especially considering how much of our personal model routing that we do right now is about not what the benchmarks would say on a test, but how we like the particular nature of one type of response versus another for a particular context. So I continue to believe that understanding different model capabilities and having model preferences is still a very high leverage activity and is going to be for quite some time. I agree. And I think the ultimate test was when GPT-5 was automatically routing us and all super users or just like more than occasional users, we were all very frustrated by what we got from the auto mode. I think that's the original test that we want control and we will probably even with a great router for many things will continue to be opinionated and rightfully so. Just for people who are also building their own, obviously they have additional levers like you can create more caching and so on, but still for them, it's much the same physics. Like the more control you have, the more you are able to be smart about the way you use the models. That's the additional levers. We talked about tokens that teach and I think that up until now we were very much focused on things that we can reduce the bill, but here I want to fight a good fight and say that we want to protect those tokens because those are in many cases the tokens that you spend in order to get much better return. And it's not just about optimizing the bill to go downwards, but rather to improve also the return that we're getting and often to improve the return, we need to improve the tokens that teach and we're talking about two ways, whether it's you teaching yourself, meaning that you run the same task using three different models in order to get to this taste of which models you like for each task or you try the same task in three different ways until you learn which one works best or your experiment with a new tool or you try a new skill or a new automation. And it doesn't work and you try something else. So all of these typically gets you overall to much better results from AI. So those should be protected firstly and also the other side of you teaching AI who you are building the systems, adding more context such it will get much more personalized results or much more organizational aware results. Those are almost always with direct correlation to how much value you get from AI and so does data. I've seen a study of 20 K developers that found that the heaviest AI users were roughly twice as productive in terms of the amount of production code that was shipped. So in many cases it's actually becoming much more like a smart exploratory user will get you to better results. So we so to summarize what you need to do in two sides. So for the individual users, these are the things that you should definitely do go and see whether you have tokens that spin. I'm sure that all of us have those idle automations or maybe some of us have like an even worse scenarios of the amount of tokens being spin without any business value practice those six habits you can even put them on a post it and just get yourself to work more effectively with the tokens that you have do spend the time to invest in reusable capabilities and improved context that the model can be much more selective and discover the relevant context where it matters. I also want you to audit the things on a schedule meaning regularly go back to the system and see what is now stale or maybe something that was working well has become a stale automation. Maybe you need to improve the context the instructions maybe you can remove some of the instructions per the new advice coming from and topic that the modern models need fewer instructions not more and make sure that you protect the learning budget and as needed. Go and negotiate that with the people responsible for the budget to make sure that you are not now being reduced to the amount of tokens that leaves you with very little room for exploration for the organizational side make the usage visible and then teach the people because when managers and employees see their own day are much smarter about how they use but make sure that they are not being encouraged to spend as little as possible but to spend smartly and also make sure that the budget is by workload and by individuals. If someone is building skills and context and usable capabilities for the entire team they need to get significantly higher budget than the person that just uses the tool as an extended Google and all the time we need to make sure that it's by that you tear it up in some organization and some individuals get significantly higher while others potentially less and not just one size fits all for the entire organization and make sure that everybody listens to something like that or that you do an internal training that teaches you to do. Meaning that teaches people on how to be smart about tokens but not how to spend as little as possible but also how to be mindful about the ROI and aiming to use tokens for the things that move the needle for the company. So that's the concrete actions for you and the team and if you want to be even more token smart so beyond the audit of your own usage we created for you a token gym that you can go and learn and flex your token smart muscles. And if you want to go even further and to learn how to build and work with the I and agents properly we do have our existing trainings and the next cohort start on early September so we'd love to have you there in the executive catch up for the executive agent leadership that will bring you all the way to be very smart about the I over smart about agents depending where you are. That's it. Awesome I think that this we're always at the beginning when we're talking about things on the show but this one is I think particularly inflection point let's say to use a word that doesn't exist we are so clearly just at the beginning of figuring out how to organize the relationship between people and the computer intelligence that they're going to consume and it is going to be it or it. and messy, which is why I think so many of these ideas that you presented are shared as frameworks, you know, patterns to explore, right? It's a set of steps that you can take to try to get a handle in these problems, but every organization at the beginning is going to solve them or not in different ways. So thank you for sharing some starting points, and you know, we'll continue to evolve this conversation as the tools around us change too. [BLANK_AUDIO]

Podcast Summary

Key Points:

  1. AI tokens are text chunks used by models; costs vary by language, code, and context length, with non-English text often costing more.
  2. Token consumption has evolved through four eras
  3. Advanced use cases like agentic tasks consume 5-30x more tokens than simple chat, with most costs tied to verification and refinement loops.
  4. Tokens are not equal across providers; tokenizers differ (e.g., OpenAI 200k vocabulary, Gemini 256k), causing 10-20% cost variations for the same text.
  5. Model updates can change token counts; e.g., Opus 4.7's new tokenizer increased tokens by 30-45% without price changes, raising real-world bills.
  6. Every request has three token layers
  7. The key metric should be cost per accepted task, not per token, to compare providers and tools effectively.

Summary:

The episode serves as a comprehensive primer on AI tokens and their economics, addressing growing corporate anxiety over AI costs in the agentic era. It begins by explaining tokens as text chunks (typically smaller than words) that models read and write, with tokenization varying by language—non-English text can cost 2-5x more—and code, where indentation and brackets add tokens. The hosts trace token consumption through four phases: token oblivious (subsidized, flat subscriptions), token maximizing (leaderboard-driven, like Meta's internal usage of 60-74 trillion tokens monthly), token anxious (employees self-censoring prompts to avoid costs, as seen with Uber's 1,500-token caps), and the desired token smart era, where spending is wise, not sparing.

They stress that the most expensive token is one your best employee fears to spend. 7's 30-45% increase). They highlight three token layers—input, reasoning (invisible, 4-20x cost), and output—and warn that long sessions compound costs as history re-sends.

The core recommendation is to shift focus from per-token price to cost per accepted task, auditing usage to avoid leaks, and governing AI adoption to prioritize value-generating use cases while preventing accidental overspending, like a yes/no question triggering a 4-million-token deep research.

FAQs

AI tokens are chunks of text that a model reads and writes, typically larger than one character and smaller than a word. They form the basis of billing, quotas, and rate limits in AI usage.

Non-English or non-Latin languages often require more tokens for the same content, sometimes 2-5 times more, due to tokenization differences. This is sometimes called a 'language tax' because billing is per token.

Everyday tasks like drafting emails are cheap, but deep research can use hundreds of thousands of tokens, and agentic coding tasks can use millions. Data analysis and agentic workflows are also very token-intensive.

Each conversation turn sends all previous messages back to the model, so by turn 10, the total tokens processed can be much higher than expected. Long sessions can quickly accumulate a large token count.

Input tokens are the prompts and context, which are cheapest. Reasoning tokens are for the model's internal thinking, priced at output rates and can add 4-20x cost. Output tokens are the visible answer, typically 3-5x more expensive than input.

Token prices are not comparable across providers due to different tokenizers and model behaviors. The same task can have different token counts and completion rates, so the operating metric should be dollars per accepted task, not per token.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.