Ep 860: Managing the AI Capability Gap: AI Is More than Ready. Most Companies are Not (Start Here Series Vol 19)
35m 34s
The Everyday AI Podcast's Start Here series addresses the critical AI capability gap, which host Jordan Wilson describes as the most urgent and real problem facing businesses today. Unlike theoretical concerns about artificial general intelligence, this gap is a present-day issue that will stall company growth if left unaddressed. Frontier AI models now match or exceed human professionals on most defined knowledge work tasks. OpenAI's GDPVal benchmark shows the best models winning 83% of blind head-to-head evaluations against industry experts, nearly doubling from 47% just five months earlier. Anthropic's labor data study reveals that while AI could theoretically automate 94% of computer and math tasks, actual usage sits at only 33%, with other sectors showing even lower adoption rates below 20%. The bottleneck is not the models themselves but organizational adoption, workflow design, and training. Recursive self-improvement, where AI models improve their own code, began in late 2025 and has accelerated capability growth far beyond what enterprises can absorb. Meanwhile, only 6% of organizations generate meaningful profits from AI despite widespread adoption. The top-performing companies invest over 20% of their digital budgets into reworking processes, while most invest only 1-5%. Wilson recommends tracking the percentage of AI-assisted workflow steps accepted without rework or incident, broken down by risk tier, on a monthly basis to begin closing the gap.
Welcome to the Everyday AI Podcast.
My name is Jordan Wilson, and for the past three and a half years, we've put out more
than 800 episodes.
Yet, one of the most common questions I get, I didn't really have an answer for.
Where do I start on the Everyday AI Podcast?
And that's why we started the Start Here series.
And with the fall now back in full swing, the Everyday AI Podcast is going back to school
and playing back the entire Start Here series from front to back.
We've hit pause on our normal Monday to Friday programming to run back our most popular
series ever for the next 30 days.
We made the Start Here series for beginners and AI champions alike.
So whether you're just trying to get a grasp on large language models or grappling with
the best coding harness for multi-agentic workflows, the Start Here series covers it all.
Any language, no jargon, and easy to follow along each day.
So make sure to subscribe to the podcast and check back each day for new insights day
by day.
The series is a culmination of spending more than 10,000 hours covering generative AI over
the past three and a half years.
So you don't want to miss a single episode of the Start Here series.
Let's get into it.
There's been a lot of talk lately about these AI models that are scary good and even more
chatter about how many of the brightest minds in AI and all the tech AI CEOs feel that we've
already reached artificial general intelligence and are racing toward super intelligence.
But what does that matter to your company?
Because right now those are kind of just theoretical issues.
It's kind of like when you see an email on a Friday and you look at it and you're like,
yeah, that's a Monday problem.
But do you know what's scarier than some private AI model that's scary good or artificial
super intelligence?
The AI capability gap.
That's the scariest of all.
Why?
Because it's very real.
And unlike that Friday afternoon email that you're maybe going to tackle on Monday, the
AI capability gap is a today problem.
And if you don't address it head on, it's going to slowly stall your company's growth.
And if your company isn't growing, the AI capability gap will undoubtedly stop your company
dead in its tracks.
And this is not an exaggeration and the timing here is specifically urgent.
So we're going to hit rewind today and explain the capability gap in AI and why AI is racing
ahead of your company and we're going to show you how to slow it down a bit and catch
it.
All right.
You ready?
This is part of our start here series on every day AI.
Let's get into it.
So if you are new here, well, let's just get to the big picture.
AI capability right now is outpacing business adoption.
And it has changed recently and the pace is too fast for any of us.
Right now, frontier AI and when we talk about actual capabilities, this is not exaggeration.
Frontier AI models, when used correctly and that's the big asterisk here, they now match
or exceed human professionals on most defined knowledge work tests.
There's a massive gap that persists between what AI could in theory automate today and
what our organizations are actually using it for, which is often just topical, right?
There's a McKinsey study that said only about six percent of organizations are generating
meaningful profits despite wide spread AI adoption.
And the bottleneck is not actually the AI models.
You could have made that argument maybe mid 2025, but not anymore in 2026.
It's actually about organizational adoption workflow design and training and education.
So here's what we're going to tackle on today's show.
So stick with me for the next, I'm going to try to make this one 25 minutes.
We'll see.
All right.
Stick with me for the next 22 ish minutes and you're going to learn the benchmark proving
that AI outperforms human experts and it's not even close.
You're going to know more about Anthropics new ish study revealing most knowledge workers
barely touch AI's true automation potential.
You're going to learn why the AI capability gap only started to show itself over the past
few months is actually kind of new and what the top 6% of AI performing companies actually
do that everyone else completely ignores.
All right.
Welcome to every day AI and this is our start here series.
This is the essential podcast series to both learn the AI basics and for AI experts to
double down.
All right.
I started this thing because after 750 plus episodes, everyone always said Jordan, great
podcast.
Right.
When they found it, they're like, where do I start?
I have no clue.
So that's why I started these start here series.
It's best if you are brand new here, start an order.
All right.
These are faster, you know, podcasts, usually about 25 to 35 minutes.
But if you listen on 2x, they even half that, right?
But this is a great way to listen from one to now volume 19 and then also make sure to
go to start here series.com.
That's going to give you free access to our exclusive inner circle community.
You can't find access right now any gap report car.
All right.
So a little bit more on that at the end.
And trust me, you are going to want to repose today's episode.
I'm just saying, all right.
And if you miss our last start here series in volume 18, and this was episode 750, we went
over the vibe coding boom, why vibe coding isn't going away and how it's both good and bad.
So make sure to go check that one out with now.
Let's talk about how you can actually manage the AI capability gap and why AI is more than
ready and companies are not.
All right, and I'm not the only one talking about this.
I've been talking about this now since I think late 20, 25, but a lot of super smart people
are starting to dive in deep on the AI capability gap.
I really like what Jack Clark said.
He is the anthropic co-founder of Anthropic.
And here's what he said in a post a couple of weeks ago.
He said most of AI progress has this flavor.
If you have a bit of intellectual curiosity and some time, you can very quickly shock
yourself with how amazingly capable modern AI systems are.
But you need to have that magic combination of time and curiosity.
And otherwise, you're going to consume AI like most people do as a passive viewer of
some unremarkable synthetic slop content or at best just asking your LLM of choice how
to roast turkey and keep it moist or Tony Box lights spinning, but not playing music.
What do I do?
And all the amazing advancements are mostly hidden from you.
All right, so that's what Jack Clark said in a kind of viral post that he had a couple
of weeks ago on X.
And I think this is very telling.
And the two things that he talked about, I think are very true, all right?
To start, well, first to even realize the AI capability gap, let alone close it.
You have to be extremely curious.
And the reason why I say that is because if you work, how you've been working for the
last few decades, you're not going to discover or your team, your organization is not truly
going to discover that capability gap because AI is not meant for the average knowledge
worker, right?
That's why I talk about all the time, how I hate, absolutely hate the concept of upskilling
because AI is not something that you sprinkle on top of a season knowledge worker, right?
Like myself, I've been working full time for 20-ish years, right?
I can't just sprinkle AI on the top.
I have to unlearn, right?
So you do have to have the time and you have to be curious on, hey, what does my role
look like if I completely started over?
And if I kind of forgot everything that got me to the point I am in my career, and that's
what you have to do.
And you also have to have a lot of time and you have to devote a lot of time to it.
Olivia Moore from A16z talked about, she said, open AI dropped a state of enterprise reports
across a million customers.
This was a couple of months ago, but her response, she said, the golf between AI power users
and everyone else is wide.
The 95th percentile user spends six cents, six times more messages than the median with
coding, writing in analysis, showing the biggest gaps.
So yeah, those people that are power users, well, because they actually understand the
capability gap, they're using AI all the time.
This is why myself, and I'm not saying this is like some weird flex, I'm saying this
because I am part of this, you know, power user group.
When I say that I use AI for 10 to 12 hours a day, right?
That's how much I'm working, you know, it's not like I'm working 16 hours a day.
I am only working in AI every single day, no matter where I am.
I'm using AI from beginning to end in literally every single step in between.
And it has completely reshade my workflow.
Kevin Rousse, New York Times columnist said this, and I really like how he put this.
He said, I follow AI adoption pretty closely, and I have never seen such a yawning inside
outside gap.
People in SF San Francisco are putting multi agent cloud storms in charge of their lives,
and solving chatbots before every decision,
wire heads.
to a degree only sci-fi writers dared to imagine. People everywhere are still trying to get approval
to use co-pilot in teams if they're using AI at all. It's possible that the early adopter bubble
I'm in has always been this intense, but there seems to be a cultural kick-take-off happening
in addition to the technical one, not ideal. And I'll say from my experience, I think there's
five main reasons that the AI capability gap has really been exposed in 2026. And number one,
the models of what they can actually do, right? I'm just going to roll through these things
really quickly. So number one, what AI models can actually do. Number two, what business leaders
think AI can do. Number three, AI literacy. Number four, human skill set. And number five, AI
access. Now let me break those five things down a little bit more in depth. So number one, what AI
models can actually do a year ago, right? So I don't know when you're listening to this start
here series episode, you know, could be in April 2026. You might be listening to it at the end
of 2026. I don't know, right? But a year ago in the beginning of 2025, you could for the most part
understand what AI models are capable of today. You absolutely can't. And it is literally getting
to the point of science fiction, right? All of these math problems and science problems that have
been plaguing researchers for decades are now being solved. I look at them and I have no clue,
right? I read the papers and I'm like, okay, I don't have any clue. I can follow and understand it
a year ago. What AI models can actually do today is mind boggling. I get to talk to a lot of smart people. It's really, right? I've had hundreds
of guests on the show. I'll say maybe a handful that I've talked to and I'm like, yes,
this business leader fully understands what AI can do for the most part. And that's not their fault
necessarily because most people in their role, they have to be an expert at one application,
right? You can't be in like very few business leaders know what AI can do. I feel fairly confident
that I'm in that category, but only because this is all I do. If I had another job, right? And I
was just an AI champion at a company, you can't, like you literally can't. You used to be able to,
you can't anymore because literally every single day, right? Whether you're talking about Google,
Microsoft, OpenAI and Throbbing, now meta is back in the conversation, perplexity, etc. Right? You
literally can't have an actual job where you have to produce something for a company and still
understand what AI can do. Not anymore. You used to be able to juggle that now unless you're literally
working 20 plus hours a day, you can do it. All right. And this kind of
goes, you know, hand in hand with, you know, number two, what business leaders understanding. So
there's a difference between understanding AI's capabilities and then AI literacy because you
have to actually be able to speak the language foundationally, which again, most people skip over
that, right? They go straight to the bells and the whistles. They go straight to clicking the
button, not knowing what happens, you know, from A to X. They only want to get to Y in the result
that brings on C. Number four is human skill set. All right. This is different than literacy.
You have to have the foundational understanding, but then you also have to have these skills, right?
I know a lot about basketball, but my basketball skill set, not great anymore, right? I think I
peaked in eighth grade, but you have to have that human skill set. And then last but not least is
AI access. So that's access to the tools to the technology in your role. Okay. So what can AI
models actually do? Right? And I do want to spend a little bit more time on that because that will
better explain this gap that I think started manageable. And now it's very hard to manage that
capability gap because it has turned into a friggin canyon. So the reason being is because AI
models are better than expert humans. All right. I've talked about this this evaluation a couple of
times, but the more I follow these AI benchmarks and all of these other things, right? There's there's
dozens of them. I think there's about five that matter, right? And one of them, and I think probably
the most important one is GDP value. So this is a benchmark from open AI. And the reason why I think
is the most valuable benchmark to look at is because it's about creating business value. So more or
less this benchmark evaluates AI models against real professional deliverables across 44 different
high GDP occupations. So these are things that you and I may do, you know, going in and doing research
on a fast moving market segment and then creating an artifact like a, you know, a pitch deck
or creating a spreadsheet, something like that. These are real world domain specific tasks where
both a human and an AI model from start to finish completed a task and then submitted their final
version to a panel of experts who are experts in that field. All right. And expert judges then
compared the unlabeled AI and human outputs in a blind head-to-head pairwise comparison. And what
we've seen now is as an example, the best model for actually creating front to back economically
valuable outputs like a human would is the open AI's GPT-54 model. And it matches or exceeds
industry professionals in 83% of these evaluations, right? I remember what some of these first GDP
value numbers came out and I'm like, oh, that's, that's pretty impressive, right? When we were
in the 30, 40%, but now it's, it's undeniable. And I do assume that by the end of the year,
that number is going to be like in the mid 90s, right? My like one of my predictions is it was going
to get to 80% and we're already at 80% is just a huge jump. That's why I keep talking about the AI
capabilities are just running too fast for most business leaders to understand and keep up.
And it's nearly doubled. That GP value score has nearly doubled in five months. So when I talk
about the last, you know, since essentially late 2025, I'm going to talk about why, but it has been,
we've seen more developments in the past five months than we have in the five years prior,
at least when it comes to largely which models and it's not even close. Right? So the in October,
the best AI model scored on that was a 47%. So it hasn't quite doubled, but it's nearly doubled
in about five or six months. And right now, AI handles most defined professional tasks yet,
most organizations barely use models in this way. It's like, did you know, right? As an example,
your, you know, GPT and clawed models can do this. They can create artifacts by default. You
don't have to have anything special turned on. It can just create spreadsheets, presentations,
word docs, et cetera, all personalized based on your company's info. All right. Next, come on
to talk very briefly about the anthropic labor data study that came out about two months ago.
I did go over this in more depth. So if you're interested in this, go listen to episode 730.
Here's essentially, similarly, whereas the GDP Val benchmark looks at actual economic output,
this anthropic labor data study looks a little bit more of the capability gap. So that's why I think
these two kind of this benchmark, open AI benchmark in the anthropic labor data study,
labor data study in tandem, really tell a powerful story of where we're at and why this capability
gap even exists. So in this study, anthropic mapped over 20,000 work tasks, right? So they actually
used some public, some public jobs data from the federal government. So they mapped over 20,000
work tasks across 800 occupations to millions of real AI conversations with their clawed chat bot.
These were anonymized. But essentially, they said, okay, what are people using clawed for? And they're
matching millions of these chats to 20,000 work tasks across these 800 occupations from this
federal data. And what they found in theory was that computer and math roles could
solve 94% of, sorry, today's most capable AI models could automate up to 94% of computer and math
tasks, right? But they were only seeing 33% usage, right? Because people assume, oh, if
computer and math, yeah, everyone's using AI for that because everyone knows that's the lowest
hanging fruit, right? Anything in coding software development, etc. Yet even in the area where people
assumed, oh, yeah, literally everyone, if you're working in anything computer, right, software,
engineering, anything dev, anything with math, of course, you're using AI and they said, well,
actually not, right? And even office, admin, business, financial, and legal roles all revealed the
same deep kind of adoption, shortfall, right? So for our live stream audience, you can always see
the video version of this on our website at your everydayai.com. So I have the theoretical capability
and observed usage by occupational category, kind of map that Anthropoc put together,
very fascinating. But all this shows is in certain categories, right? Such as management,
business and finding.
computer and math, you know, architecture and engineering, legal is another high one, arts and media,
right? Like AI's capabilities when mapped to literal, the actual work tasks that the federal
government uses to define these roles. I mean, most of them have coverage in the 80 to 90 percent,
yet the observed, or sorry, that's the theoretical coverage. So in theory, today's most powerful
models could do anywhere from 80 to 90 percent of the actual work, right? Computer and math was
the high one where they had an observed coverage of a little more than 30 percent, but these other
areas are not even 20 percent, right? They're for the most part, you know, in the 20s, some of them
below 20. So you have this huge gap where in most of these general cases, right? Like management,
you know, in theory, today's AI models, if you understand the capabilities, they can automate
85 plus percent, yet you don't even have a 20 percent coverage. That is a crazy gap, right?
Where essentially you have a magic wand that if you know how to use the magic wand and you know
the magic words, and you say poof work begun, poof the work is gone, but people don't know because
it's impossible to keep up with. And some other things that they found on this study, which is
actually the workers that are most impacted, they said that workers most exposed to AI right now
earn 47 percent more in hold graduate degrees at nearly a four times rate. So essentially, I think
people early on assume that AI would displace or would potentially have the capabilities to
displace workers who were maybe a little more junior or, you know, weren't as high of the kind of
quote unquote knowledge work totem pole, so to speak. And what they found was the exact opposite,
right? It is those people who have those higher degrees, and then conversely people with physical
and manual occupations, registered the lowest exposure according to inthropics study.
So how do we get here so quickly, right? How do we get as an example from GDP value, where
literally five and a half months ago, it was even, right? It was a little less. It was about
a 47 percent tie or win rate against humans, right? It was a coin flip. And if the off the shelf
model that you can pay $20 a month for was as good as a, you know, and human with a decade of
experience who specializes in that to now humans aren't going to be able to compete in a couple of
months, right? By blind benchmarks, almost every single time within probably six months,
humans are always going to prefer the AI model. How do we get here so quickly, right? Two years ago,
the AI models were not good at all, right? They're, they're actually pretty bad, right? Especially
obviously we have today's comparison to drawback on. But one of the biggest reasons I think is
recursive self-improvement. Stick with me here if you're not super technical. All right, so recursive
self-improvement or RSI, it's a concept within artificial general intelligence where essentially an
AI system improves its own source code, architecture or training data, leading to a more capable model,
which then improves itself further, creating a self-reinforcing loop. All right, so why am I talking
about this kind of strange niche concept called recursive self-improvement on something about
managing the AI capability gap? Well, because at the end of 2025, the AI companies started to either
directly admit or to kind of allude to the fact that they were all now using recursive self-improvements
on their models, right? So now, their quote-unquote big models were good enough that it could start
writing its own code. It could start improving itself, right? I think, you know, inthropic with
their cloud code, probably one of the most consequential products of the AI life cycle outside
of ChatGPT. You can make the argument that cloud code is one of the most important outside of ChatGPT,
right? The lead at inthropic's cloud code says, "Yeah, we don't write code anymore." You know,
cloud codes, cloud code, right? And we've seen the same inferred by different researchers at
sorry, at OpenAI as well, you know, Google has been a little less direct, but they've still alluded
to the fact that a lot of their models today, not just the smaller versions that are distilled from
the biggest versions, but the biggest versions are being improved by themselves, right? And these
new big scary models, right? Like inthropic's mythos or, you know, OpenAI's whatever it's called,
you know, Sput or Glacier, right? If you're listening to this in six months, these names,
these code names don't matter anymore, but one of the reasons that maybe the general public isn't
getting their hands on them is, well, they're too computeintensive, but they are using these models
internally to improve the best consumer models and to also ship new products. And that's why
inthropic, as an example, their ship rate in February and March was straight up off the charts,
right? And then we found out later, well, one of the reasons was because they were using this
mythos model to put out a lot of this, a lot of these new features that we're all using now.
And this is why, right? Because a year ago, it would have taken a team, even using AI would have
taken teams way longer. But now that you have this kind of recursive self-improvements or models
improving themselves and building new features that use those models, right? This is why it's hard
now to keep up. So this started, like I said, in 2025, and ever since large language
model capabilities have far outpaced enterprise training and learning and development, even the
companies that wanted to do it right, I think they could do it in quarter to quarter three of last
year. But now you can't. Unless you have an entire segment of people, maybe 5% of your total
workforce that literally has no deliverables. And all they do, I've said this all along, your team
needs a bunch of needs where all they do all day is they just scope different AI models. They play
with AI releases, they sandbox thing. They're not building anyone for anything. They're just
building solutions for what they think the company needs. And then they're training those people.
But companies don't have that, right? And that's why at least right now it is nearly impossible
to keep up. And most companies, though, claim AI adoption, but they can't actually prove real
results. And it's getting even harder and harder. Right? That McKinsey study that I talked about
in 2025 said that 88% of organizations are using AI, but only 6% are generating meaningful
business profits. And right now, I think leadership confidence is running far ahead of what
frontline practitioners report actually seeing on the ground. That's the thing, right? So a lot of
times the people who are in charge of AI at certain companies now, because it's getting
easier to build, I think maybe their hands are on keyboard or their eyes are on monitoring agents
a little bit more. And they are getting removed from what the frontline practitioners are actually
experiencing. So not only is it pretty hard to manage that gap just from a technical perspective,
but from a change management and a people management perspective, it's getting even harder.
And I think that's why some AI capability gaps are closing, but others are still remaining stuck.
So as an example, right? AI is helping AI enabled humans close some gaps, but many still exist
and are getting worse. So as an example, coding performance, right? This in a lot of different
benchmarks, this surge from single digits to 90% unstructured benchmarks in three years. In terms of
what the AI itself was capable to do. And that helps obviously the humans that use this as part of
their daily workflow, close a big chunk of that gap individually, right? But you now have this,
you know, PhD level models with, you know, science and standardized math capabilities that are
rapidly approaching the ceiling of the current tests. So you aren't even necessarily able to know by
the benchmarks what the capability gap is because these benchmarks are becoming saturated. So we
even the AI community doesn't even fully understand these models capabilities because the benchmarks,
right? A lot of them have been stuck in the 90 to 95% tile. You know, a lot of them are at 98, 99.
So I think the benchmarks themselves are getting saturated. So even the people building the models
aren't even fully aware or understand what these models are actually capable of.
So let's talk about how you can actually manage this gap and start to close it. All right. Right now,
the top performers, right? The top companies are investing more than 20% of their digital budgets
to fundamentally rework their processes. All right. Let me repeat that. 20% of their digital
budgets. So whatever your digital budget budget is, I'm guessing most companies are probably at 1%.
Maybe maybe 5%. Very few. Only the top performers are investing that 20% of their digital
budget to rework their day-to-day knowledge processes, right? Because you need to learn to separate
workflows into risk tiers with verification and human approval matched to each level.
And organizations that wait for AI to become ready, they're going to find that their competitors
have already captured the advantage. Because in the same way, right? Just to draw a little parallel here.
What inthropic, and I keep saying they've won 2026 so far, well, it's because they were
first, right?
They were first to at least internally close that capability gap and put it to work, right?
And that's why they've been able to outpace their competitors.
I don't know how long it'll last.
We'll see because I think the other labs have closed that gap and I think we'll start
to see that soon.
Think of that within your own organization, right?
The first to close the gap is going to be able to accelerate at a pace that we haven't
seen before, right?
That's why you see, unfortunately, a lot of these big companies, you know, block as an
example, cutting 40% of their workforce and they seem fairly confident that they're going
to be able to actually grow revenue because of the way they've completely reworked their
organization.
Obviously, Jack Dorsey had a very fascinating essay that we shared about in our newsletter
about how they're essentially flipping the work pyramid on its head.
So here's what I want to have you focus on, one metric, okay?
I want to make this very digestible for you and hopefully actionable as we close out
today's show.
I want you to track the percentage of AI-assisted workflow steps right now that are accepted
without any rework or incident.
And here's why, because you probably, most organizations, they kind of get their AI plan
or their AI training for the year, right?
Or some companies, unfortunately, it's kind of like one time.
Maybe the companies that are really investing heavily, you might get it once a quarter.
But when was the last time that you used the most capable thinking models?
I'm talking about GPT-54 Pro, I'm talking about Opus-46 with extended reasoning.
I'm talking about Gemini-31 Pro with a higher thinking budget, right?
What was the last time that you used those and tracked the percentage of your workflows
that get accepted without rework or incident?
And then I want you to break that metric down by risk tier, right?
So the low risk, you know, drafting versus the high stakes legal or financial outputs.
So organizations that measure and understand that operational reliability for the high percentage
of AI-assisted workflows that can be accepted at our lower risk.
You will be surprised that depending on what your team does, depending on what your company
does, depending on what sector you're in, you'd be surprised without too much investment
aside from just reverse engineering your current day-to-day processes.
You'd be surprised to say that about 30 to maybe 60% of a lot of the work that many
of us do, right?
I'm not talking about, you know, people in specialized role, I'm saying, if your organization
has 1,000 employees, you know, if you look at that work collectively, you'd be surprised.
I would say 30 to 60%.
If you rescope everything would fall in that can essentially be automated without rework
or incident and is in a lower risk or a medium risk tier.
And if you're not measuring that on an ongoing basis, you can't scale.
But that's step one to managing the AI capability gap and you have to be doing this, right?
At least monthly.
Again, a year ago, you could get away with quarterly, the way models, right?
If you're using a model from last quarter, good luck, right?
That's showing up to an F1 race in a bicycle.
Good luck.
You're going to get smoked.
You don't see it a chance.
So you can no longer take these year-long pilots, these quarter-long plans.
You have to be agile to actually understand number one, the capability gap, but to begin
to manage it.
All right.
I hope this was helpful in our start here series.
Like I said, make sure to repost today's episode on LinkedIn.
Here's why we put together the AI capability gap report cards to help you know.
So where your organization stands when it comes to number one, understanding this gap and
number two, how you can actually tackle it.
So in this report card, it's a great guide for you and your team to go through it together
to understand the latest capabilities of all the models and how they break down for different
types of work.
All right.
So if you didn't know, yes, if you're listening on the podcast, this is actually live streamed
on LinkedIn.
So in the podcast show notes, we always put a link to today's LinkedIn show.
So just go click repost and I will send that capability gap report card your way.
All right.
Thank you for tuning in.
If you haven't already, please go to start here series.com.
That's going to give you free access to the inner circle community.
And you can go listen to all of the start here series in order in the playlist that we
have.
And you can go read and listen to all of the shows there.
So thanks for tuning in.
We hope to see you back tomorrow and every day for more everyday AI.
Thanks, y'all.
And that's a wrap for today's edition of Everyday AI.
Thanks for joining us.
If you enjoyed this episode, please subscribe and leave us a rating.
It helps keep us going.
For a little more AI magic, visit your everyday AI.com and sign up to our daily newsletter
so you don't get left behind.
Go break some barriers and we'll see you next time.
Podcast Summary
Key Points:
AI capability is outpacing business adoption, creating a growing "AI capability gap" that is a present-day problem threatening company growth.
Frontier AI models now match or exceed human professionals on most defined knowledge work tests, with OpenAI's GDPVal benchmark showing AI winning 83% of head-to-head evaluations against experts.
Anthropic's labor study found AI could theoretically automate 94% of computer and math tasks but actual usage is only 33%, revealing a massive adoption shortfall across industries.
Only 6% of organizations generate meaningful profits from AI despite 88% claiming adoption, with the bottleneck being organizational workflow design and training rather than model capability.
Recursive self-improvement, where AI models improve their own code and training, began in late 2025 and is a major reason model capabilities now far outpace enterprise learning and development.
The workers most exposed to AI disruption earn 47% more and hold graduate degrees at nearly four times the rate, contradicting early assumptions about who AI would displace.
Top-performing companies invest more than 20% of their digital budgets to fundamentally rework knowledge processes, while most organizations invest only 1-5%.
Organizations should track the percentage of AI-assisted workflow steps accepted without rework or incident, broken down by risk tier, on at least a monthly basis.
Summary:
The Everyday AI Podcast's Start Here series addresses the critical AI capability gap, which host Jordan Wilson describes as the most urgent and real problem facing businesses today. Unlike theoretical concerns about artificial general intelligence, this gap is a present-day issue that will stall company growth if left unaddressed. Frontier AI models now match or exceed human professionals on most defined knowledge work tasks.
OpenAI's GDPVal benchmark shows the best models winning 83% of blind head-to-head evaluations against industry experts, nearly doubling from 47% just five months earlier. Anthropic's labor data study reveals that while AI could theoretically automate 94% of computer and math tasks, actual usage sits at only 33%, with other sectors showing even lower adoption rates below 20%. The bottleneck is not the models themselves but organizational adoption, workflow design, and training.
Recursive self-improvement, where AI models improve their own code, began in late 2025 and has accelerated capability growth far beyond what enterprises can absorb. Meanwhile, only 6% of organizations generate meaningful profits from AI despite widespread adoption. The top-performing companies invest over 20% of their digital budgets into reworking processes, while most invest only 1-5%.
Wilson recommends tracking the percentage of AI-assisted workflow steps accepted without rework or incident, broken down by risk tier, on a monthly basis to begin closing the gap.
FAQs
The AI capability gap refers to the difference between what AI models can do and what organizations actually use them for. It's significant because AI is outpacing business adoption, and if left unaddressed, it can stall company growth and even halt progress entirely.
Frontier AI models now match or exceed human professionals on most defined knowledge work tests. In particular, AI models have achieved 83% or higher performance in real-world professional tasks like creating pitch decks or spreadsheets, according to OpenAI's GDP value benchmark.
The study found that while AI models could automate up to 94% of computer and math tasks, only 33% of those tasks are actually being used. In other high-skill roles like management, business, or legal, adoption is even lower, with most areas showing less than 20% usage despite theoretical automation potential.
The gap exists due to poor AI literacy, lack of human skills, limited access to tools, and slow adoption of workflow redesign. Most organizations rely on basic AI use without understanding its full potential, while top performers invest heavily in reworking processes and training.
AI capabilities have advanced rapidly, with performance benchmarks nearly doubling in just five months. This includes breakthroughs in math, science, and coding, making AI models more capable than human experts in many defined tasks.
Recursive self-improvement (RSI) is when an AI system improves its own code or training data, leading to faster progress. This has enabled major AI companies to build more powerful models internally, accelerating innovation and contributing to the current AI capability gap.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.