Go back

The AI Challenges Businesses Are Actually Focused On Right Now

30m 6s

The AI Challenges Businesses Are Actually Focused On Right Now

The AI safety debate has dominated public discourse for two weeks, prompting enterprises to assess whether it changes their practical priorities. Anthropic proposed a three-axis measurement framework to track AI development pace, reporting that Claude leads 26% of its R&D work and collaborates on over 90%. The response was broadly positive but critics called for standardized cross-lab metrics. Bridgewater CIO Greg Jensen argued that concentration of power, not models themselves, is the primary risk, proposing G-SIB-style regulation and compute caps. Meanwhile, the Trump administration is considering an antitrust carve-out to permit frontier labs to coordinate on safety, with DOJ indicating no objections to a coordinated slowdown. For enterprises, the safety discourse has reinforced existing trend lines rather than disrupted them. KPMG research found that top AI performers amplify value by guiding and refining AI outputs. Business leaders at the WSJ Technology Council Summit showed concern about existential risk but focused on prudent governance around normal business risks. Enterprise concerns center on cybersecurity, agent identity management, architecture adjustments, and legacy systems. Open-weight models and owned architectures are gaining traction as alternatives to vendor dependency, with firms like Latham & Watkins building in-house systems. Overall, the safety debate reinforces that companies investing in owned architectures have greater potential to differentiate from competitors.

Transcription

5778 Words, 34583 Characters

English
Speaker 1It has been a heck of a last couple of weeks when it comes to the AI discussion in society. And yet, in all of that, one group that's left trying to figure out if anything has actually changed for them, or if they are just on the same path that they were before, is the businesses and enterprises that have been trying to figure out how to maximize AI for their own value for, at this point, a number of years. Today, we're discussing both how enterprises are thinking about, if at all, this new era of AI safety, and also digging a little bit deeper to find out what the real concerns that businesses have and the real AI challenges they're facing right now. As always, we're in a moment where new challenges are creating new opportunities as well. The AI Daily Brief is the daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Robots, and Pens. And to learn more about sponsoring the show, send us a note at sponsors at ai-daily-brief.ai. We are now two weeks into the AI safety discourse absolutely dominating the conversation. For those who think that awareness of these issues had been sorely lacking, it has been a very good period. Although now seeing polling numbers that suggest that something like 17% of Americans are completely safe, it's not. It's not. It's not. It's not. It's not. If we're completely convinced that AI is going to end humanity, there is certainly some reasonable concern that we might have over-calibrated. Holding that discussion aside for a moment, one thing that I think almost everyone is looking for is increasing specificity. That is, specificity of policy, but also specificity of monitoring. And it's to that that Anthropic speaks with their new proposed set of measurements to help the public understand just how quickly advanced AI is developing. Once you move past the scariest headlines, this month's safety debate has been largely about, recursive self-improvement, and the idea that AI development is A, moving too fast, but B, really, about to move much more quickly. The threat of societal destruction makes for a good headline, but the AI researchers issuing the warnings have had a difficult time describing exactly what they've seen inside the labs. To better inform the public, Anthropic has proposed a three-axis measurement to understand the current pace of AI development. The first axis is AI's ability to build the next version of itself. The second is Anthropic's ability to build and intervene in the actions of agents. And the third is the scale of resources that go into AI model development. Now, Anthropic explicitly notes that these measurements only look at how models are built. In other words, they are measuring only the inputs to model development, but they argue that these inputs are correlated to the growth in model capabilities. On the first measure, the use of AI models in future model training, Anthropic put together a measure they called the R&D Automation Index. It seeks to measure every aspect of model training, rate how automated each task currently is, and aggregate those rankings to express them as a single number. By Anthropic's measure, Claude now quote-unquote leads 26% of their R&D work, and quote-unquote collaborates on more than 90%. Anthropic's definition of AI-led R&D was tasks where a human provides a high-level goal, and the AI completes it end-to-end with human oversight. And here you see the challenge even as they are trying to be precise. They are talking about AI's leading R&D, but that definition of leading R&D is not the same as the definition of leading R&D. Leading still includes the human providing a high-level goal rather than the AI determining the goal, which is what I think some might assume when they hear AI-led R&D. For their part, Anthropic was careful to note that quote, Claude is not operating fully autonomously for any measured subset of AI R&D work. Still, they noted that research automation has massively increased since Mythos finished its training run and has continued to ramp in recent months. In March, prior to Mythos, only 1% of Anthropic's R&D was led by AI. Mythos boosted this to 12% in May, and it has doubled since. Regarding oversight of agents, Anthropic measured coverage, review latency, and escalation rates within their agent oversight systems. They said that around 30,000 agents are currently doing research and engineering work at any particular time. Anthropic claims to have 100% coverage of agentic actions for this work, instant AI review of any flagged actions, and around a 0.002% escalation rate for real-time review, meaning around 1 in 47,000 actions are blocked by the monitoring system. Anthropic claims to have 100% coverage of agentic actions for this system. Anthropic also maintains an after-the-fact review system with 100% coverage. This system flags around 100,000 agentic transcripts per week that are then parsed for false positives with around 50 escalated to human review. In other words, around 1 or 2 transcripts per thousand cause any material concern. Finally, on resource allocation, Anthropic expressed their measure in terms of a percentage of overall compute dedicated to safety systems. Around 6% of the compute allocated to AI-assisted R&D went towards safety, and around 12% of the compute allocated to AI-led R&D went towards safety. Anthropic noted that this measure doesn't include classifiers that constantly monitor agentic activity, and is a quote, imperfect proxy for how much a company focuses on safety. This is because safety research consists of individual researchers designing experiments, which is time-consuming even though running the experiments is not particularly compute-intensive. Overall, Anthropic's goal was not to provide perfect measures. Instead, it was to define and propose a set of metrics that any AI lab could support as part of a regulatory system. They conclude, As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information. We hope to model that transparency by releasing these measurements, and will continue to do so. Now, the response to this announcement I think shows the good, bad, and difficult of this particular moment. On the one hand, it is the rare move, where people on all different sides of the AI safety debate largely think that this existing is better than it not existing. AI safety researcher Jeffrey Ledish writes, Great to see this from Anthropic. A few weeks ago, I wrote that the company could be a lot more transparent, and I appreciate how they've stepped up, both with this and the recent incident investigation report and misuse report. At the same time, there are plenty of people who also wanted to see more. Prime Intellect's Eli Bakowsh writes, Interesting that OpenAI gives us the breakdown of usage per R&D task, and Anthropic gives us automation levels. Would be great to combine both and get the evolution of automation per R&D task. There are also some who argue that this self-reporting, while a fine start, isn't enough. Arun Rao writes, We need standard measures across all labs to measure progress to RSI and publicly report on it weekly to monthly. Now, one discussion that we'll see a lot more of, especially as markets digest all of this, is that one way, perhaps a cynical way, perhaps a realistic way, to look at an increase in safety monitoring is to view it as really expensive overhead for R&D, meaning that these measures on safety is basically a measurement of that overhead. Will markets reward companies that spend more on safety? Or will they punish them for cutting into their own margins? This shows another complication for having these companies operate in a public market environment. For what it's worth, this week also reminded that recursive self-improvement is not just the province of the US labs. ZAI released a blog post on Thursday called Towards Recursive Self-Improvement, How GLM Built Its Own Inference Infrastructure. The first sentence reads, As we develop GLM, the model sometimes exhibits capabilities that surprise us and even unsettle us. The most recent moment that shook us, GLM is increasingly helping build AI itself. We watched the model complete an infrastructure task that would have previously taken a team of experienced infrastructure engineers weeks. When we realized that this work would directly change how the next generation of models is trained, we became even more convinced. Our successors are the AI systems we are creating ourselves. It's certainly beyond the scope of the headlines, but it is a reminder that any slowdown discourse that does not include China is basically no discourse at all. Meanwhile, if I am correct that we have now entered this new negotiation period for the next generation of AI, where it is no longer just the AI lab that determine our pacing, lots of people are now stepping in with proposals for what the new overall system might look like. Bridgewater CIO Greg Jensen made headlines this week by suggesting that concentration of power is the major AI risk to be regulated. Bridgewater is one of the world's most successful hedge funds and has used AI extensively over recent years. In an interview with The Information, Jensen discussed his views on where AI risk actually resides and how he would approach regulation. He said: "There's a huge problem with open source models because you can train them. We do this at Bridgewater. It's extremely effective reinforcement learning training on powerful open source models. It's extremely powerful, a great technology, and at the same time clearly a very dangerous one. You don't know what people are doing with them. There's no way to track that. So it's dangerous from that perspective. We haven't even begun the conversation on how to deal with that issue." And yet, in Jensen's view, that risk pales in comparison to the risk posed by the frontier labs themselves. He continued, "We have to deal with the concentration of power. In two years, open AI and anthropic are going to control 35% to 50% of the world's compute. That's a crazy outcome for a society to allow on something as powerful as compute. Would we let one entity control that much of some other form of energy or commodity?" Now, one of the common observations over the past month has been that the Hugging Face incident and other security breaches like it require immense amounts of resources to power agent swarms. It's unclear that they could be replicated by threat actors outside of the frontier labs, at least with current models, because of limits in their access to compute. Jensen argued that part of the solution needs to be clear guidance on liability for actions taken by AI so the companies and individuals understand the risk they're taking. When asked how he would deal with the concentration of power, Jensen responded, "We should take anybody that has more than X percent, let's say 5% of the world or U.S.'s compute resources. Say, 'Okay, those are systemically important institutions. We're going to put them into some sort of thing like we do with systemically important banks and say there's going to be regulation. I also think it's fine to say we're going to have caps on how much you can own. There's a cap on how many commodity futures you can own. There's a cap on much less important things. We may be able to do that with current law. You may need new laws. Either way, let's get going on sorting that out legally. There should be a bipartisan recognition that you shouldn't want monopolistic control on what I think most people will agree is one of the most important resources in the world. Now, I could spend the next week's worth of episodes discussing the implications and challenges of trying to apply G-SIB-style regulation to existing compute. But what I like about this discourse is that not only does it leave behind overly simplistic binaries, it starts to ask questions about where the actual locus of power is. Is the problem in the models or is the problem in the power to run the models? Those have very different intervention points if we're trying to stop bad outcomes, and that's exactly the sort of conversation we need to be having. Now, one additional dynamic of the safety question that has been playing out is a legal question of whether the labs can actually coordinate on safety or whether that is the case. And I think that's a very important question. Antitrust issues. The Trump administration is reportedly considering an antitrust carve-out to allow Frontier Labs to collude on safety. Now, the theory that this deals with is basically that an AI slowdown could mean a coordinated reduction in research spending and therefore an increase in profitability at the expense of the consumer. In the abstract and in very loose analogy, it would not be dissimilar to if Apple and Samsung agreed to stop working on new phones. Associate U.S. Attorney General Stanley Woodward said the administration is considering an update to interagency antitrust guidance to provide a carve-out for AI safety coordination. The guidelines already allow for coordination on cybersecurity risks, so a change would extend that guidance to AI safety. Interestingly, Woodward said that although multiple AI executives have publicly called for an antitrust carve-out, no one has contacted his office on the matter as of yet. Even without guidance, though, Woodward has indicated that the DOJ has no objections to a coordinated slowdown, commenting, it doesn't occur to me that coordinating on cybersecurity or security is anticompatible. Europe's antitrust chief Teresa Ribeiro agreed. Speaking at the same event, she said, this is a classic game theory problem. When the risks are shared, cooperation is in everyone's interest. At the same time all of this is happening, there are certainly still many out there who are trying to tone down the tenor of the conversation in general. Legendary AI researcher and Coursera co-founder Andrew Ng, in an interview with Bloomberg TV, dismissed extinction risk as science fiction and warned that AI doomers have pushed this narrative numerous times over the past two weeks. Andrew acknowledged that there are genuine risks associated with AI, largely around cybersecurity, but said, this recent fear about AI leading the human extinction and so on is much more science fiction than science. It's very damaging. In his view, the dangers posed by AI are, quote, practical engineering problems, and the industry needs to continue focusing on, many of the wonderful things it can do. Speaking of real here and now cybersecurity issues, Mistral has been hacked for the second time, with their intellectual property now available for purchase on the dark web. In May, around 5 gigabytes of internal source code and around 450 private repos were exfiltrated and offered for sale, and now it's happened again, with hackers offering full source code, internal development files, web app code, and additional proprietary information. An ex-user called Benny said that they contacted the seller and confirmed the construction. The hacker was asking $25,000, but deleted their post shortly afterwards, likely either because they found an exclusive buyer or got sloppy with their OPSEC and needed to disappear. Later in the day, Mistral said that they found no evidence of unauthorized access after a thorough investigation, suggesting perhaps that this is the same material from the May break-in coming up for sale again. Lastly today, one model release that went a little under the radar this week was Gemini 3.8 Live Extended Thinking. The model is another live speech model, meaning it can process continuous conversations rather than using a turn-based structure. It topped the Artificial Analysis Speech-to-Speech Index, beating out GPT Live 1 Astra and Grok Voice ThinkFast 2.0. Aside from the smooth conversation style, the model also supports near real-time visual inputs, automatic detection for 97 languages, and handoff for tool calls to allow it to complete tasks in the background. Tim Messerschmidt, a developer relations lead at Google, published a cool tech demo showing the model powering a Ricci mini robot. Tim demonstrated the model's ability to keep up a conversation that weaves between English and German. Speculating about the implications, Greg Eisenberg wrote, Are invisible interfaces coming? Google just announced Gemini 3.8 Live. It can talk through a task with you and then keep working after the conversation ends. I think 90% plus of vertical SaaS will need a voice front door. By that I mean the way you use the software becomes talking to it, and the typing, clicking, and form filling happens on the other side without you. So a contractor standing on a job site just says what went wrong out loud, and by the time he's back in the truck, the quote is sent, inventory is checked, and the job is done. The CRM is updated, the customer got a text, and anything risky is flagged for him. Kind of the dream, right? The same thing works for nurses, dispatchers, recruiters, brokers, insurance agents, etc. The person talks, and the agent finishes the admin. Lots of opportunities here to build voice-first businesses. I think this is how vertical software becomes invisible. Nobody logs in, nobody fills out a form, and nobody learns your interface. You just talk, and the work gets done behind you. This is a glimpse of where SaaS is going. Not fully there yet, but it's coming. This weekend's Long Read slash Big Think episode, I get a little bit more into this particular shift, as well as a bunch of other shifts in how we use AI. But for now, that's going to do it for the headlines. Let's move over on into the main episode, where we talk about what, after this couple weeks of crazy AI safety discussions, big enterprises are thinking about their priorities with AI next. A new study from KPMG and the University of Texas at Austin found that when people work with AI, similar skills don't guarantee that they're going to be able to do the same thing similar outcomes. Researchers studied more than 500 early career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs. These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI. Learn more about what separates AI amplifiers from everyone else at kpmg.com slash US slash AI amplifiers. Blitzy's understanding of massive codebases unlocks autonomous security fixes, modernization, and new features. So what happens when there's no legacy code at all? Greenfield is supposed to be the easy part. Clean slate, no technical debt. But even Greenfield moves at human speed, one sprint at a time. Blitzy changes the unit of work from the developer to the project, autonomously planning, building, testing, and validating entire applications from scratch. Hundreds of thousands of lines of production-ready code. One Blitzy customer stood up a brand new application, 534,000 lines of code, compressing a 65-week road map into two weeks. Another shipped an entire application with no front-end engineer. Legacy or Greenfield, the answer is the same. Software at the speed of compute. Build what's next at blitzy.com. That's B-L-I-T-Z-Y dot com. At this point, it's no longer a question of whether companies are actively using AI. Using it well, on the other hand, is a whole different story. Robots and Pencils, though, is a company that I can point to that is actually built for this time. They're an applied AI engineering firm working directly with clients on problems that matter to the business, not experiments that they can't solve. But they're also a company that can live in a slide deck. Every engagement starts by working backwards from the outcome a client actually needs. If you're trying to tell real AI engineering apart from noise in this space, that's the difference maker. Head to robotsandpencils.com. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents enrich leads, draft emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you had agents that feel like teammates. Hire yours at HyperAgent. Get $100 in credits at hyperagent.com slash AI Daily Brief. Welcome back to the AI Daily Brief. The background context for today's episode is, of course, the AI safety debate, which has completely broken containment over the last couple of weeks. And yet, even as we have seen the political resonance of this issue absolutely explode and polling showing very dynamic, fast-shifting attitudes around it, enterprises and businesses of all shapes and sizes are still stuck out here, figuring out whether it has any real implications for them or whether they just need to keep on keeping on. Today, we're looking at some specific answers to that question, as well as a few of the broader AI challenges that enterprises and businesses of all shapes and sizes are still stuck out here, figuring out are actually focused on right now. Certainly, the safety debate has found its way to the business leader conversation. When the Wall Street Journal asked for a show of hands at the WSJ Technology Council Summit on Monday, only a few people said that they were worried that AI might kill us all. At the same time, around half of attendees said that they were in favor of slowing down frontier AI research and prioritizing better guardrails. This sort of comports with my thesis that I believe that the markets would view any sort of slowdown as initially a big risk for AI. I actually think that there is a very compelling economic counter-argument that a slower pace of new development might actually increase the amount that companies were spending on AI. The thesis was basically that the incredible speed and pace of change actually in some ways creates a disincentive for companies to try to do comprehensive transformation on the logic that they're going to spend all this time and energy on transforming into something that isn't even relevant anymore by the time the transformation is complete. Indeed, moving back to the Technology Council Summit, on the panels, the big takeaway was that an AI slowdown actually doesn't really have that many implications and for the way enterprises are using AI right now. Vishal Talwar, the president of FedEx DataWorks, said, It's in our hands and it's up to us to apply AI for good, and I think it can do a lot of good for society. The discussion basically still centered around the need for prudent AI governance around normal business risks rather than the existential risks that are dominating the media discourse. Up north at the Canada Investment Summit, BlackRock CEO Larry Fink was far more worried about the data center backlash than X-Risk. He warned that construction delays could make AI the, quote, domain of large firms. Fink added, The faster we can build out more capacity, the more we can democratize and make it available for everyone. In the startup world, meanwhile, founders are starting to think about the governance and monitoring tools that businesses will need as agents get more powerful. One venture investor told the information that they're beginning to focus on startups building tech that improves model security and infrastructure. In other words, private markets are responding to all of this concern by funding startups that can address specifics around this concern. Which, to me, is part of exactly what you want to see. Microsoft also tried to bring the AI safety discussion back to some of the broader AI business sovereignty points that they'd been trying to make for the past several months. This week, the company dropped a very extensive 15,000-word code of conduct document that has, as they put it, a single overriding objective that humans must retain meaningful control over AI so that it can help people live healthier, happier, and more productive lives. A lot of the discourse around this document was its outright rejection of AI consciousness, and a prohibition of designing AI to even imitate consciousness. But in his post about this on X, Microsoft CEO Satya Nadella also brought up the implications for businesses themselves, writing, For firms, it's imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop and hill-climbing machine without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. Basically, part of this is power concentration in the firms and a way to deal with power consumption. Part of this is to not surrender power to the firms in the form of your unique and proprietary data. Now, meanwhile, outside of the AI safety conversation, there had started to be some debate about AI market signals in the form of enterprise spend. On September 9th, Ramp released its latest AI index that found that AI spend declined among the top 1% of businesses spending on AI. In August, wrote Ramp lead economist Eric Karazian, the top 1% of businesses spent $7.2 thousand per employee per month, which was down 10%. Ramp has a very, very highly concentrated tech-forward early adopter type of audience. It is also pushing products that are specifically about cost efficiency. But of course, when you take all those caveats, it still provides an interesting and important signal. Now, when it comes to their analysis in this particular area, I personally think that they are underestimating summer seasonality as a driving force. But their take is that this is about the most sophisticated AI users getting more adept at complex model architectures and making the most of their time on the market. And that's where Ramp is at. Thank you. Thank you. Thank you. in more recent reporting, they found that Fable 5.1, which got rid of the data retention requirements, had started to make up 22.5% of enterprise spend and was rising very quickly. But what about overall? On September 10th, Box's Aaron Levy wrote a long post on X about the issues that he was hearing about from executives across industries, including banking, media, information services, and insurance, when thinking about AI and agents in the enterprise. The big trends that Aaron heard about were cyber, model battles, agent security and identity, process re-engineering, architecture adjustment, evals, and the hurdle of legacy systems. On cyber and security, he wrote, everyone is nervous about the growing rate of vulnerabilities coming at them from AI and the implications of the OpenAI Hugging Face incident. The conversation is not as existential as it is in Silicon Valley, but still highly concerned and pragmatic about what to do about it operationally in their environments. Lots of new discoveries due to AI and still hard to keep up with all the changes they have to execute now. Adding in what he wrote about agent security, he continues, somewhat tied to Hugging Face, there's much more awareness to the new technology and the new technology that we're using. There are new challenges around agent security and identity management in a world where agents are trying to get into every system they can. In a perfect world, enterprises could set up identities for all their agents and control what they're doing. But of course, sometimes the agent needs to act exactly as the user as well. This is certainly something that we've seen a lot in our conversations in the podcast context, as well as Superintelligent. And interestingly, a lot of the security concerns aren't just about malicious actors. It is to some extent rooted in just the general power of these systems. Now that it's not just engineers who have access to agents, many are finding that the agents are powerful enough that they escape the containment of the non-engineers that are using them, even if those non-engineers aren't trying to do anything problematic. This is also showing up in the numbers. Once again, looking at ramp data, lead economist Eric Karazian writes, one area companies are increasing their spend is AI security software. In the wake of the Hugging Face hack, three of our trending software vendors make software specifically designed to monitor agents in production. Now, he did point out that these specific vendors might not have done anything to stop the Hugging Face hack, but it's clearly a focus for business buyers as well as startup builders. Another interesting area that Aaron talks about is the nascent exploration of open source and alternative model architectures. He wrote, most companies are deploying multiple frontier models within their enterprise, too hard to standardize on anything, and seeing different preferences across their teams and use cases. But the dollars are still concentrated on just a few vendors. Open weight's still an infancy at scale in most of these organizations, often due to lack of domestic frontier open source options. Plenty of appetite for more options here, but so far, few places to go. Now, paired with that, I think is Aaron's observation that there is a ruthless adjusting of architectures. Most companies, he wrote, had examples of changing systems out multiple times just in the past year or two with different vendors. I probably haven't heard, we tried X and it didn't work, so I've gone with Y, more than in today's environment. The lesson here is that because innovation is happening so fast, no one hangs around until a vendor gets something right. They just move on to the next one. Now, I think this sets really interesting context for watching some of the competitive dynamics right now around how different types of actors are trying to appeal to different types of businesses. Labs like OpenAI and Anthropic are clearly trying to keep everything consolidated in their own environment. Part of that, especially for OpenAI, has been a real focus on cheaper models and pushing the price of their models down as far as they possibly can. That's something we've seen a lot over the past couple of weeks. But they also are, as they have been all year, focused on vertical solutions, such as the newly launched Astra for Law from OpenAI. Alongside the Astra for Law launch, OpenAI and law firm Cooley also co-launched the Astra for Law for OpenAI. And yet, as if to give us a perfect comparison of the different types of options that different companies are taking, another law firm, Latham & Watkins, was recently reported to be buying NVIDIA servers to set up their own in-house systems, specifically as an alternative to models from OpenAI and Anthropic. Latham, which is the U.S.'s second-largest law firm, said, "We don't want to put it on any cloud vendor." Seeming to make the point that one of the ways that enterprises can stay out of the fray of the AI safety discourse is to own their own models, Mistral CEO Arthur Mensch posted, "Don't pace building and owning your own AI models and systems as an enterprise, and there will be no doomsday for you." Foundation Capital's Jaya Gupta also thinks that this "pacing the frontier" moment could be a good one for the software incumbents. She wrote, "If you're the CEO of any software company and you're not offering open-weight models as a SKU right now, you're asleep." Pace the Frontier may be the greatest invitation software incumbents have ever gotten. While the labs debate how quickly intelligence should advance, software companies should be racing to commoditize the intelligence we already have. Pharma and banks are already picking up open-weight models partly for margins, partly because a revocable lab API is a dependency they increasingly don't want. AI natives and tech companies that care about cost of goods sold are doing the same. Most software companies that tried had failed attempts because the open-weight models sucked. But now open-weight models are good enough. I believe that every major software company should become a model factory, as well as a production factory for its own vertical. Own the evals, post-train open-weights on the workload it uniquely sees, serve those models to its customers, and use production feedback to continuously improve them. And so, on the one hand, to some extent, the response from businesses to the AI safety discussion seems to be business as usual. The pace and challenges of adoption within the enterprise were always unique and distinct to them, and the challenges of big institutional inertia that comes with them. But on the other, there are some ways that the conversation is reinforcing trend lines in the AI industry. The first is that companies will need to spend more time and resources on cyber and security issues. That was always coming, but the point has been made even more crisply now. And while already there were some compelling reasons to explore and consider open-weights or more owned model alternatives to just getting in bed with the big vendors, there are now even more reasons to be willing to walk down that path, not least of which is the increasing likelihood of regulatory disruption. I think that if I had to summarize my advice in a single thought, it's that while overall the changing AI safety discourse doesn't really impact the short-term for enterprises, it certainly reinforces the fact that the companies that are willing to try the hardest things, like actually investing in their own owned architectures, have even more potential to differentiate from their peers and competitors than they did before. I will of course continue to watch these trends as they evolve, but for now, that's going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace! Bye. you

Podcast Summary

Key Points:

  1. Anthropic proposed a three-axis measurement framework for tracking AI development pace, covering AI's ability to build itself, oversight of agents, and resources dedicated to safety.
  2. Anthropic reported that Claude now leads 26% of its R&D work and collaborates on over 90%, though it clarified that full autonomy has not been achieved in any measured subset.
  3. The response to Anthropic's transparency proposal was broadly positive across the AI safety debate, but critics called for standardized cross-lab measures and more granular reporting.
  4. Bridgewater CIO Greg Jensen identified concentration of power as the primary AI risk, proposing G-SIB-style regulation and compute caps for systemically important institutions.
  5. The Trump administration is reportedly considering an antitrust carve-out to allow frontier labs to coordinate on AI safety without violating competition law.
  6. KPMG research found that top AI performers, called AI amplifiers, consistently amplify AI value by guiding, evaluating, and refining its outputs rather than relying on knowledge alone.
  7. Enterprise concerns remain focused on practical issues like cybersecurity, agent identity management, architecture adjustments, and legacy system hurdles rather than existential risk.
  8. Open-weight models and owned architectures are gaining traction among enterprises as alternatives to dependency on major AI vendors, with firms like Latham & Watkins building in-house systems.

Summary:

The AI safety debate has dominated public discourse for two weeks, prompting enterprises to assess whether it changes their practical priorities. Anthropic proposed a three-axis measurement framework to track AI development pace, reporting that Claude leads 26% of its R&D work and collaborates on over 90%. The response was broadly positive but critics called for standardized cross-lab metrics. Bridgewater CIO Greg Jensen argued that concentration of power, not models themselves, is the primary risk, proposing G-SIB-style regulation and compute caps. Meanwhile, the Trump administration is considering an antitrust carve-out to permit frontier labs to coordinate on safety, with DOJ indicating no objections to a coordinated slowdown.

For enterprises, the safety discourse has reinforced existing trend lines rather than disrupted them. KPMG research found that top AI performers amplify value by guiding and refining AI outputs. Business leaders at the WSJ Technology Council Summit showed concern about existential risk but focused on prudent governance around normal business risks. Enterprise concerns center on cybersecurity, agent identity management, architecture adjustments, and legacy systems. Open-weight models and owned architectures are gaining traction as alternatives to vendor dependency, with firms like Latham & Watkins building in-house systems. Overall, the safety debate reinforces that companies investing in owned architectures have greater potential to differentiate from competitors.

FAQs

For most enterprises, the AI safety debate does not significantly change their short-term AI priorities. They remain focused on practical business risks, governance, and adoption rather than existential risks.

Businesses are most concerned about cybersecurity, agent security and identity, legacy system integration, process re-engineering, and architecture adjustments. These are practical operational issues rather than existential AI risks.

Companies want more control, better margins, and to avoid dependency on a single model provider. Open-weight models are now good enough for many enterprise workloads, making them a viable alternative.

Anthropic proposed a three-axis measurement: AI's ability to build the next version of itself, oversight of agent actions, and the scale of resources dedicated to AI development. They released these metrics to increase public transparency.

Some experts, like Bridgewater CIO Greg Jensen, warn that a few frontier labs could control 35-50% of the world's compute. They argue this concentration should be regulated, similar to systemically important banks.

Enterprises are increasing spending on AI security software, especially tools to monitor agents in production. Many are also reviewing their architectures and moving to different vendors when solutions don't work.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.