Go back

The data black hole at the center of AI

11m 57s

The data black hole at the center of AI

Sample efficiency—the amount of data needed for a model to learn a task—is a critical yet underappreciated aspect of AI intelligence. Despite significant scaling in model size and compute, progress in sample efficiency remains minimal. Most AI advancements stem from vast, domain-specific datasets generated through reinforcement learning, where models are trained on millions of human expert demonstrations. These datasets are highly task-specific, requiring hundreds of experts to generate examples, rubrics, and reasoning chains. In contrast, humans learn efficiently through natural development, with only a fraction of their lifetime experience in sensory input. Even a teenager’s 20 hours of driving practice or a human’s lifetime of sensory exposure pales compared to the trillions of tokens used to train frontier AI models—millions of times more. Scaling model parameters alone cannot close this gap, as scaling laws show only a modest reduction in required data. Humans possess a fundamentally different learning curve, suggesting AI still lacks true generalization. While AI can automate repetitive white-collar tasks due to data abundance, complex, open-ended tasks like software engineering remain challenging due to the lack of efficient learning. The ultimate path forward may lie in AI-driven AI research, where automated systems tackle the sample efficiency problem. This requires a deeper understanding of how intelligence emerges—not just through more data, but through more efficient learning structures.

Transcription

2513 Words, 14292 Characters

English
So, one definition of intelligence is sample efficiency. That is to say, how much data do you need to give in domain to operate fluently and comfortably? And it's actually not clear that we've made that much progress in training sample efficiency over the last few years. It seems like more so, we've just dramatically widened and improved the data distribution. The main way that AI has been getting better is we're adding more and better data and scaling the compute required to develop that data in the first place. Obviously, RL is the main way that this has happened. You can think of RL as basically a kind of synthetic data generation, where you dump a ton of compute against a verifier or a rubric if you have an LLM as a judge. And you do this in order to find out what the good data is in the first place. And then you train your model to predict these correct rollouts much in the same way that you might train that model to predict the next word in internet text. For this process to work, the model must have at least some prior probability to anticipate the correct solution in the first place, which is why you need mind stretching amounts of human expert trajectories in every single field and skill that you want the model to eventually be competent in. It's hard to overstate how task-specific and bespoke this human expert data is. If you want some intuition, I recommend checking out the drop descriptions on Mercor or search's websites. There are listings for word specialist who will convert legacy documents into polished word files. And legal experts will write realistic eminent diligence or securities filings, and management consultants who will write up template market research. And there's not only that the data have to be so domain-specific, but there has to be so much of it. Each skill corresponds to at least hundreds of human experts who are generating example completions, writing rubrics, and explaining their chain of thought. There's a reason that the data industry that is producing these expert labels and the oral environments in which these meticulously catalog skills can congeal is earning billions a year in revenue, soon to be deck and billions. Now imagine if it took a couple decades worth of courses with hundreds of concurrent professors and millions of practice tasks for you to learn how to polish a word file. Even the task count difference here understands that, because the models have to grind their far more numerous tasks, each far harder, whereas a human student might practice a textbook problem once or twice. With GRPO, these models are generating hundreds to thousands of rollouts per task, and they need to assault the credit assignment problem. The correct way to think about these models is not like a human who has learned all these different skills that you see these models displaying. It's more like a Frankenstein's monster, which has been built out of a billion graphs of carefully constructed examples all sewn together. Epoch recently reported that open models lag state-of-the-art frontier models by four months. I think the reason it is relatively easy for open source and previous laggards to catch up to within months of the frontier is that data is the real driver of progress. And data can be easily distilled from public APIs, whereas hyperparameters and training tricks and architectural optimizations cannot. And if the latter were driving most of the progress, then catching up would be far harder than we are observing it to be. It is easy to forget how much data these models are trained on, and how much more it is than what we human see in our lifetimes. We see these AI's as the galaxy glittering with capabilities. But at their center, invisible to the naked eye, holding all the constellations together is an unimaginably massive black hole of data. Just a couple of points of comparison to help drive home how big this difference is. Here's one. If a person sees and hears on average, let's say generously, 2,000 more to an hour, then between the time they're bored and the time they're an adult, they'll see about 200 million tokens. Now, by contrast, these frontier models are trained on somewhere between tens to hundreds of trillions of tokens. That is close to a million-fold difference. Here's another point in comparison. If you wanted to, you could learn to tell or operate any random humanoid or robot arm within hours. And if you could get AI's to learn just as fast, robotics would be a decade trillion dollar industry, and you'd have an endless army of unitary G1s doing all kinds of useful work in the world. But the reason we can't do this is that our AI's learn much less efficiently than we do. And even with the millions of hours of demonstrations that we collected, this is not enough to allow them to perform complex open-ended tasks. And a final point of comparison, a teenager can learn to drive a car with about 20 hours of practice. And even if it include their 16 years of growing up and understanding how the world works and building physical intuition, then it's still three to four orders of magnitude less data than Waymo and Tesla are using to train their self-driving car models. Now, I want to deal with a couple of common responses and objections that people have to these kinds of comparisons. One thing people will say, and I think our probably said this when it came onto my podcast, is that for humans, many billions of years of evolution had to go into basically pre-training us. And so we're being unfair when we're comparing how little data we see within our lifetimes to what these coal-started LLNs were just starting off with a totally random initialization have to learn from. I think this is not the right way to think about it. Our genome is only three gigabytes big. And only one to two percent of it is protein coding. And that is simply not enough space to store the parameters of this network that supposedly evolution has pre-trained. I think the closer analogy is more that evolution found the right hyperparameters and the right loss functions. And that within our lifetime, we are still from scratch building up the connect home in our brain. That is to say, the analogous thing to the weights and parameters of the neural network itself. And even if you granted this comparison, you said, yes, the hundreds of trillions of tokens that these models see to get pre-trained is similar to just catching up to evolution. That still doesn't explain why any new marginal capability that you want to get these models takes so much data. So once you have been educated, again, you don't need a hundred different professors to teach you how to learn a new programming language. But these AIs, even once they're pre-trained, still require enormous amounts of data to learn the next marginal skill, and the next marginal skill after that. Another objection to this kind of comparison is that we're not including multimodal data that we're seeing in our lifetimes. So you include all this sensor information that we see from birth to adulthood. That's probably tens to hundreds of billions of tokens of data. And my response to this objection is simply that blind and deaf people who have been cut off from all the sensor information still have general intelligence. And that suggests to me that all these billions of sensory tokens are not really the thing that is making humans smart. And in fact, deaf people who don't have the ability to hear any tokens, who just have to consume them via sign language and reading, are probably ingesting far less than the 200 million language tokens that we ballparked earlier, which suggests that even the million-fold difference that we calculated earlier might be an understatement. OK, the third common objection people make is that we just haven't scaled enough. We have the scaling laws. They tell us the bigger models are more sample efficient. The human brain, we know, is about 100 trillion synapses. And we have frontier models that are currently at around five trillion parameters. And so maybe we could just achieve human level sample efficiency if we made these models one to two orders of magnitude bigger. The reason the subjection is offmark is actually quite interesting. So if you look at the way the scaling laws equations work, they tell you that the parameter and data terms are added to the loss independently. So suppose you have a model and you've trained it compute optimally. And you say, I want to be sample efficient. I want to use as little data as possible. And I'll throw in as many parameters as necessary to make that happen. So take the constants from the chinchilla scaling law paper, even if you increase the number of parameters by infinity, that would only decrease by a factor of 10, the amount of data that you need in order to keep the same loss. Humans are somewhere between thousands to millions of times more sample efficient than these models. So scaling the size of current models simply can't make up for that discrepancy. And this really does suggest that humans are in a different scaling curve altogether. As soon as I earn money, I want to put it to work. But I also need to say for things like upcoming expenses and estimated taxes. So to figure out exactly how much I need to set aside, I ask command. Command is AI that is built into Mercury, which is my baking platform. And since I already use Mercury to run my entire business, command has access to all the information it needs to get worked on. I just tell command the date I'm interested in, and it does the rest. It takes my current balance. And as whatever invoices will we do by the cutoff. Then I reviews my last six months of transaction history. So I can subtract out my monthly average expenses along with any scheduled payments. And if there's anything relevant coming up that's not in Mercury yet, I can just fly it. Things like heads up, there's a $12,000 contractor payment that's slated for July. And that gets included in the final output. Because this is all happening in chat and every answer has links to the underlying data, I can easily double check commands work. And once I'm convinced, I can just tell command, all right, that looks good. Just transfer the surplus to my personal account. And you will immediately draft the transfer for me to approve. Command is live now. Visit mercury.com/command to learn more. Mercury is a Fintech company, not an FBIC in short bank. Baking services provided through choice financial group and column NA members FBIC. EI-journated responses and suggested actions made very in our knock guarantee. Okay, all these nerdy comparisons aside, you might ask, why do we even care about sample efficiency? Is this actually necessary for the labs to achieve the two overarching objectives they have, which are one, automate white color work and two, automate AI research itself. The bet that the labs are making with white color work is that the common tasks that are software engineer or analyst or accountant needs to do are common. And as a result, you can bring them into the training distribution quite easily. If you look at the revenue curves of these labs or the last few months, it does suggest that there's an enormous amount of value from bringing into distribution these kinds of common tasks, even if we can't replicate whatever is making human learning so special. And it might be more inefficient to train AI's to do these kinds of tasks than it is to train humans. But so what? Human lifespan simply does not allow for the quantity and the breadth of training that these models experience. If you, as a human, had some weird learning disability where you needed to read through every public repository on GitHub before you could be a competent software engineer, there would simply not make sense to train you up. You'd be on social security by the early stages of your education, and even once you were trained, you would only be able to work on one project at a time. But AI's can learn these skills by firehosing gigawatt, the training at a time, and what they learn can be amortized across billions of sessions at once. So we can be ludicrous in the inefficient in training them up, and still be wildly in the green. And then there's a question of, well, how much idle distribution thinking do white collar employees need to do that you simply can't train for an advance? This is more a question about the nature of different jobs than it is a question about AI research. And it also depends on which job you're talking about. Some jobs are so mechanical and predictable that we were able to automate them long before the modern era of AI. For example, bank tellers or travel agents, but there are other jobs which require dealing on a daily basis with problems that are quite distant from the data distribution. I think software engineering is probably one such. This is the job that AI is supposed to take first, but I would be willing to bet that there's overall more demand for human software engineers in 2027 than there is right now. Largely due to the complementary input of AI. The last plans for this latter category of jobs is first to automate AI research and then have the automated AI researchers solve the sample efficiency problem. So then the question is can AI's, which do not have human level sample efficiency, nonetheless solve the remaining research problems that stand on the way of human-like intelligence and learning? This is a very complicated question, and I'll have to address it in a much longer future block post. But just a tease at a bit, I think that the way that people currently think about an intelligence resolution is very clumsy because either people dismiss the possibility of AI speeding up AI progress altogether, or they assume that some kind of God pops out the other end. They don't reason carefully about what it looks like to have a period where AI progress is much faster than usual, but have that happen atop LLMS and the particular kinds of intelligence that LLMS are. But I'll save that for next time. In the meanwhile, if you want to read this block post or all the other block posts I write or be alerted when I write a future block post, go side enough for my newsletter at my website,

Podcast Summary

Key Points:

  1. Sample efficiency—how much data a model needs to learn a task—is a fundamental measure of intelligence, yet progress in this area has been limited despite advances in model size and compute.
  2. AI improvements are primarily driven by massive data scaling, especially through reinforcement learning (RL), which acts as synthetic data generation by simulating millions of expert human trajectories.
  3. The data required to train AI models is astronomically larger than human experience—frontier models see tens to hundreds of trillions of tokens, compared to humans’ roughly 200 million lifetime tokens.
  4. Human learning is highly efficient and generalizable; AI, in contrast, requires domain-specific, task-specific, and vast amounts of expert-generated data for each new skill.
  5. Scaling model size alone cannot overcome the sample efficiency gap; scaling laws show that even infinite parameters only reduce data needs by a factor of 10.
  6. Humans are far more sample-efficient than current AI models, and this efficiency is not due to sensory input volume but to biological and cognitive structure.
  7. The real bottleneck for AI is not computation or model size, but the sheer volume and specificity of human-curated training data.
  8. AI has the potential to automate white-collar work and accelerate AI research, but only if it can overcome the inefficiency in learning new skills.

Summary:

Sample efficiency—the amount of data needed for a model to learn a task—is a critical yet underappreciated aspect of AI intelligence. Despite significant scaling in model size and compute, progress in sample efficiency remains minimal. Most AI advancements stem from vast, domain-specific datasets generated through reinforcement learning, where models are trained on millions of human expert demonstrations.

These datasets are highly task-specific, requiring hundreds of experts to generate examples, rubrics, and reasoning chains. In contrast, humans learn efficiently through natural development, with only a fraction of their lifetime experience in sensory input. Even a teenager’s 20 hours of driving practice or a human’s lifetime of sensory exposure pales compared to the trillions of tokens used to train frontier AI models—millions of times more.

Scaling model parameters alone cannot close this gap, as scaling laws show only a modest reduction in required data. Humans possess a fundamentally different learning curve, suggesting AI still lacks true generalization. While AI can automate repetitive white-collar tasks due to data abundance, complex, open-ended tasks like software engineering remain challenging due to the lack of efficient learning.

The ultimate path forward may lie in AI-driven AI research, where automated systems tackle the sample efficiency problem. This requires a deeper understanding of how intelligence emerges—not just through more data, but through more efficient learning structures.

FAQs

Sample efficiency refers to how much data a model needs to learn a task effectively. More efficient models require far less data to achieve competence compared to less efficient ones.

AI models have improved primarily through increased data volume and compute power, especially via reinforcement learning, which generates vast amounts of synthetic training data.

AI models need domain-specific, high-quality human-generated examples—such as legal filings or market research—to learn complex skills, as these examples provide accurate and realistic training data.

Frontier AI models are trained on tens to hundreds of trillions of tokens, which is nearly a million times more than the 200 million tokens a human sees and hears in a lifetime.

No, scaling model size alone cannot compensate for the vast gap in sample efficiency. Theoretical scaling laws show that even with infinite parameters, data requirements only reduce by a factor of 10.

Unlike humans, AI models start from random initialization and must learn through massive amounts of data, requiring hundreds of thousands of examples per task to master a new skill.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.