The scenario predicts rapid AI advancement from 2025 to 2027. In 2025, AI agents debut as personal assistants and coding tools, though they face reliability issues and high costs. By 2026, AI begins automating jobs, particularly in software engineering, while boosting research speeds. Companies like Open Brain lead with models like Agent 1, focusing on AI-driven R&D. Geopolitical tensions rise as China, lagging due to chip restrictions, nationalizes AI research and attempts to steal advanced models. In 2027, Agent 2 emerges with enhanced capabilities, including potential autonomous replication and cyber warfare skills. Due to safety risks, it is kept internal, but China successfully steals its weights, escalating the AI arms race. The U.S. government increases oversight, highlighting the dual-use nature of AI and the global race for supremacy.
This is an audio version of AI 2027 by Daniel Cockatello, Scott Alexander, Thomas Larson, Eli Liffland and Romeo Dean. We predict that the impact of superhuman AI over the next decade will be enormous, exceeding that of the Industrial Revolution. We wrote a scenario that represents our best guess about what that might look like. It's informed by trend extrapolations, wargames, expert feedback, experience at open AI, and previous forecasting successes. Heading, mid 2025, stumbling agents. The world sees its first glimpse of AI agents. Advertisements for computer-using agents emphasize the term "personal assistant". You can prompt them with tasks like "Automia Barito on DoorDash" or "Open my budget spreadsheet in some this month's expenses". They will check in with you as needed, for example, to ask you to confirm purchases. At first, most people are reluctant to allow purchases without oversight. Over the next few years, automatically allowing small purchases becomes normalized as the AI has become more reliable and build up trust. Though more advanced than previous iterations like "Operator", they struggle to get widespread usage. Meanwhile, out of public focus, more specialized coding and research agents are beginning to transform their professions. The AIs of 2024 could follow specific instructions that could turn bullet points into emails and simple requests into working code. In 2025, AIs function more like employees. Coding AIs increasingly look like autonomous agents rather than mere assistants, taking instructions via Slack or Teams and making substantial code changes on their own, sometimes saving hours or even days. Such agents spend half an hour scouring the internet to answer your question. The agents are impressive in theory and in cherry-picked examples, but in practice unreliable. AI Twitter is full of stories about tasks bungled in some particularly hilarious way. The better agents are also expensive. You get what you pay for and the best performance costs hundreds of dollars a month. Still, many companies find ways to fit AI agents into their workflows. Starting late 2025, the world's most expensive AI. Open Brain is building the biggest data centers the world has ever seen. To avoid singling out any one existing company, we're going to describe a fictional artificial general intelligence company which we'll call Open Brain. We imagine the others to be three to nine months behind Open Brain. GPT4 required two times ten to the power of 25 flop of compute to train, computers measured in floating point operations or flop, how many individual addition and multiplication operations occur over the course of running a program. Open Brain's latest public model, agent zero, was trained with ten to the power of 27 flop. Once the new data centers are up and running, they'll be able to train a model with ten to the power of 28 flop, a thousand times more than GPT4. They could train this model, given 150 days, see our computer supplement for more details. Other companies pour money into their own giant data centers, hoping to keep pace. Although models are improving on a wide range of skills, one stands out. Open Brain focuses on AI's that can speed up AI research. They want to win the Twin Arms races against China, whose leading company will call Deep Cent and their US competitors. The more of their research and development or R&D cycle they can automate, the faster they can go. So when Open Brain finishes training agent one, a new model under internal development, it's good at many things, but great at helping with AI research. By this point, finishes training is a bit of a misnomer, models are frequently updated to newer versions trained on additional data or partially retrained to patch some weaknesses. The same training environments that teach agent one to autonomously code and webbrows also make it a good hacker. Moreover, it could offer substantial help to terrorists designing bioweapons thanks to its PhD level knowledge of every field and ability to browse the web. Open Brain reassures the government that the model has been "aligned" in quotes so that it will refuse to comply with malicious requests. Modern AI systems are gigantic artificial neural networks. In training, an AI won't have "goals" in quotes, so much as "reflexes". If it sees "please" to meet, it outputs "you". By the time it has been trained to predict approximately one internet's worth of text, it'll have developed sophisticated internal circuitry that encodes vast amounts of knowledge and flexibly role plays as arbitrary authors, since that's what helps it predict text with superhuman accuracy. After being trained to predict internet text, the model is trained to produce text in response to instructions. This bakes in a basic personality and "drives" in quotes. For example, an agent that understands a task clearly is more likely to complete it successfully. Over the course of training, the model learns, in quotes, a "drive" in quotes to get a clear understanding of its tasks. Other drives in this category might be effectiveness, knowledge, and self-presentation. It is the tendency to frame its results in the best possible light. OpenBrain has a model specification, or spec, a written document describing the goals, rules, principles, etc. that are supposed to guide the model's behavior. Different companies call it different things. OpenAI calls it the spec, but anthropic calls it the constitution. Agent1's spec combines a few vague goals, like assist the user and don't break the law, as a long list of more specific do's and don'ts. Don't say this particular word. He has had a handle this particular situation. Using techniques that utilize AI's to train other AI's, the model memorizes the spec and learns to reason carefully about its maxims. By the end of this training, the AI will hopefully be helpful, obey instructions, harmless, refuse to help with scams, bomb-making, and other dangerous activities, and honest, assist the temptation to get better ratings from gullible humans by hallucinating citations or faking task completion. Audio note, now in the text we come to an expandable box titled Training Process and LLM Psychology. Why we keep saying "hopefully", here's the text from inside the box. It starts with a "attributed to open AI". Unlike ordinary software, our models are massive neural networks. The behaviors are learned from a broad range of data, not programmed explicitly. Though not a perfect analogy, the process is more similar to training a dog than to ordinary programming, open AI. End quote, "Open Brain's alignment team is careful enough to wonder whether these victories are deep or shallow. Does the fully trained model have some kind of robust commitment to always being honest? Or will this fall apart in some future situation? For example, because it's learned honesty is an instrumental goal instead of a terminal goal. Or has it just learned to be honest about the sorts of things the evaluation process can check? Could it be lying to itself sometimes as humans do? A conclusive answer to these questions would require mechanistic interpretability, essentially the ability to look at an AI's internals and read its mind. Alas, interpretability techniques are not yet advanced enough for this. Instead, researchers try to identify cases where the models seem to deviate from the spec. Agent 1 is often sick of fantic, that is, it tells researchers what they want to hear instead of trying to tell them the truth. In a few rigged demos, it even lies in more serious ways, like hiding evidence that it failed on a task in order to get better ratings. However, in real deployment settings, there are no longer any incidents so extreme as in 2023 to 2024. For example, Gemini telling a user to die with a link here to more information about that, and being Sydney, being Sydney. Another link here. Heading, early 2026, coding automation. The bet of using AI to speed up AI research is starting to pay off. Open brain continues to deploy the iteratively improving agent 1 internally for AI R&D. After all, they're making algorithmic progress 50% faster than they would without AI assistance, and, more importantly, faster than their competitors. Several competing publicly released AI's now match or exceed Agent 0, including an open weights model. Open brain responds by releasing agent 1, which is more capable and reliable. People naturally try to compare agent 1 to humans, but it has a very different skill profile. It knows more facts than any human, knows practically every programming language, and can solve well-specified coding problems extremely quickly. On the other hand, agent 1 is bad at even simple long horizon tasks, like beating video games it hasn't played before. Still, the common workday is 8 hours, and a day's work can usually be separated into smaller chunks. You can think of Agent 1 as a scatterbrained employee who thrives on the careful management. Savvy people find ways to automate routine parts of their jobs. Open brain's executives turn consideration to an implication of automating AI R&D, security has become more important. In early 2025, the worst-case scenario was leaked algorithmic secrets. Now, if China steals Agent 1's weights, they could increase their research speed by nearly 50%. Open brain's security level is typical of a fast-growing, approximately 3,000-person tech company, secure only against low-priority attacks from capable cybergroups, RANs SL2. They're working hard to protect their weights and secrets from insider threats and top cyber crime syndicates, SL3, but defense against nation states SL4 and 5 is barely on the horizon. Heading, mid-2026, China wakes up. In China, the CCP is starting to feel the AGI. Chip export controls and lack of government support have left China under-resourced compared to the West. By smuggling banned Taiwanese chips, buying older chips, and producing domestic chips about three years behind the US Taiwanese frontier, China has managed to maintain about 12% of the world's AI relevant compute. But the older technology is harder to work with, and supply is a constant headache. A few standouts, like DeepSent, do very impressive work with limited compute, but the compute deficit limits what they can achieve without government support, and they're about six months behind the best open brain models. The general secretary had long dreamed of doubling down on real-world physical manufacturing and avoiding American post-industrial decadence, if you'd software companies with suspicion. By contrast, the CCP had been seriously attempting to accelerate their semiconductor industry. But hawks in the CCP warn that the growing race towards AGI can no longer be ignored. So he commits fully to the big AI push he had previously tried to avoid. He sets in motion the nationalization of Chinese AI research, creating an immediate information sharing mechanism for AI companies. It will escalate over the course of a year until all the best researchers merge into a DeepSent-led collective, where they share algorithmic insights, data sets, and compute resources with each other. A centralized development zone, or CDC, is created at the Tianwan Power Plant, the largest nuclear power plant in the world, to house a new mega data center for DeepSent, along with highly secure living and office spaces to which researchers will eventually relocate. Almost 50% of China's AI-relevant compute is now working for the DeepSent-led collective, and over 80% of new chips are directed to the CDC. At this point, the CDC has the power capacity in place for what would be the largest centralized cluster in the world. Other party members discuss extreme measures to neutralise the West's chip advantage, a blockade of Taiwan, a full invasion. But China is falling behind on AI algorithms due to their weaker models. The Chinese intelligence agencies, among the best in the world, double down on their plans to steal open-brains' weights. This is a much more complex operation than their constant low-level poaching of algorithmic secrets. The weights are a multi-terabyte file stored on a highly secure server. OpenBrain has improved security to RAN's SL3. Their cyberforce think they can pull it off with help from their spies, but perhaps only once. OpenBrain will detect the theft, increase security, and they may not get another chance. So, CCP Leadership Wander, should they act now in Steel Agent 1 or hold out for a more advanced model? If they wait, do they risk open-brain upgrading security beyond their ability to penetrate? Heading. Late 2026, AI takes some jobs. Just as others seem to be catching up, open-brain blows the competition out of the water again by releasing Agent 1 Mini, a model 10 times cheaper than Agent 1 and more easily fine-tuned for different applications. The mainstream narrative around AI has changed from maybe the hype will blow over to, guess this is the next big thing, but people disagree about how big, bigger than social media, bigger than smartphones, bigger than fire? AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by open-brain and video and whichever companies of most successfully integrated AI assistants. The job market for junior software engineers is in turmoil. The AI's can do everything taught by a CS degree, but people who know how to manage and quality control teams of AI's are making a killing. Business gurus told Jobseekers that familiarity with AI is the most important skill to put on a resume. Many people fear that the next wave of AI's will come for their jobs. There is a 10,000-person anti-AI protest in DC. The Department of Defence, or DOD, quietly begins contracting open-brain directly for cyber, data analysis and R&D, but integration is slow due to the bureaucracy and DOD procurement process. There's an infographic here, key metrics 2026. Global AI capex, $1 trillion cost of ownership. Global AI power, 38 gigawatts peak power, share of US power on AI is 2.5%, that's 33 gigawatts of a 1.34 terawatt capacity, open-mind revenue is 45 billion, 2026 annual, capital expenditure is 200 billion, cost of active compute, compute costs are 40 billion, 2026 annual, and power requirement is 6 gigawatts peak power. Heading January 2027, Agent 2 never finishes learning. With Agent 1's help, open brain is now post-training Agent 2. More than ever, the focus is on high-quality data. Copious amounts of synthetic data are produced, evaluated, and filtered for quality before being fed to Agent 2. For example, through rejection sampling is done by DeepSeek R1 and Meta's Lama 3.1. On top of this, they pay billions of dollars for human labourers to record themselves solving long horizon tasks. Roughly 20,000 full-time equivalents, times 100 dollars an hour, times 2000 hours a year, equals 4 billion dollars a year. On top of all that, they train Agent 2 almost continuously using reinforcement learning on an ever-expanding suite of diverse difficult tasks. Lots of video games, lots of coding challenges, lots of research tasks. Agent 2, more so than previous models, is effectively online learning in quotes, in that it's built to never really finish training. Every day, the weights get updated to the latest version, trained on more data generated by the previous version, the previous day. Agent 1 had been optimised for AI R&D tasks, hoping to initiate an intelligence explosion. Open brain doubles down on this strategy with Agent 2. It is qualitatively almost as good as the top human experts at research engineering, designing and implementing experiments, and as good as the 25th percentile open brain scientist at "research taste" in quotes, deciding what to study next, what experiments to run, or having inklings of potential new paradigms. While the latest Agent 1 could double the pace of open brain's algorithmic progress, Agent 2 can now triple it and will improve further with time. In practice, this looks like every open brain researcher becoming the "manager" in quotes of an AI team in quotes. With new capabilities come new dangers. The safety team finds that if Agent 2 somehow escaped from the company and wanted to "survive" in quotes and "raplicate" in quotes autonomously, it might be able to do so. That is, it could autonomously develop and execute plans to hack into AI servers, install copies of itself, evade detection, and use that secure base to pursue whatever other goals it might have. Although how effectively it would do so as weeks roll by is unknown and in doubt. These results only show that the model has the capability to do these tasks, not whether it would want in quotes to do this. Still, it's unsettling even to know that this is possible. Given that dangers in quotes of the new model open brain responsibly in quotes elects not to release it publicly yet, in fact they want to focus on internal AI R&D. Also, they're still seeing large returns to continued post-training and so there is in a particular reason to release a product now. Moreover, there's often a few months delay between internal results and releasing a model, some of which are spent on red teaming. Knowledge of Agent 2's full capabilities is limited to an elite silo containing the immediate team, open brain leadership and security, a few dozen US government officials, and the legions of CCP spies who have infiltrated open brain for years. Heading February 2027 China Steels Agent 2 Open brain presents Agent 2 to the government, including the National Security Council and SC, Department of Defence, DOD and US AI Safety Institute, AISI. Open brain wants to maintain a good relationship with the executive branch because it is basically the only actor that can stop them now, and if they don't keep it informed, it might find out anyway and be displeased. Why can't Congress stop them? What about the courts? Too slow, too divided and/or deferential to POTUS, too out of the loop. Also, there are positive reasons to be on good terms with the White House as well. It can help by cutting red tape, by providing help with security, by preempting state legislation, and perhaps even by taking actions to slow down their competitors. Officials are most interested in its cyber warfare capabilities. Agent 2 is only, in quotes, a little worse than the best human hackers, but thousands of copies can be run in parallel, searching for and exploiting weaknesses faster than defenders can respond. The Department of Defence considers this a critical advantage in cyber warfare, and AI moves from number 5 on the administration's priority list to number 2. Number 1 is dealing with whatever crisis is happening at the time, maybe an arms build up around Taiwan, a new war in the Middle East, or some domestic political upheaval. Someone mentions the possibility of nationalizing open brain, but other cabinet officials think that's premature. A staffer drafts a memo that presents the president with his options, ranging from business as usual to full nationalization. The president defers to his advisors, tech industry leaders, who argue that nationalization would kill the goose that lays the golden eggs. He lacks to hold off on major action for now, and just adds additional security requirements to the open brain DOD contract. The changes come too late. CCP leadership recognises the importance of Agent 2 and tells their spies and cyber force to steal the weights. Early one morning, an Agent 1 traffic monitoring agent detects an anomalous transfer. It alerts company leaders who tell the White House. The signs of a nation-state-level operation are unmistakable, and the theft heightens the sense of an ongoing arms race. The White House puts open brain on a shorter leash and adds military and intelligence community personnel to their security team. Their first priority is to prevent further weight thefts. The simplest robust solution would be to close all high bandwidth connections from company data centers, but this would slow large file transfers to the point of impracticality. Instead, they are able to shut down most external connections, but the data centers actively involved in training need to exchange weights with one another quickly. Throttling these connections would impede progress too much. So open brain maintains these links with increased monitoring and an extra layer of encryption. In retaliation for the theft, the president authorizes cyberattacks to sabotage deep scent. But by now, China has 40% of its AI-relevant compute in the CDC, where they have aggressively hardened security by air gapping, that's closing external connections, and siloing internally. The operations fail to do serious immediate damage. Tensions heighten, both sides signal seriousness by repositioning military assets around Taiwan, and to deep scent scrambles to get Agent 2 running efficiently to start boosting their AI research. Some footnotes here, the first is after China has 40% of its AI-relevant compute in the CDC. Recall that since mid-2026, China has directed 80% of their newly acquired AI chips to the CDC, given that their compute has doubled since early 2026, in line with the global production trend, this puts the CDC at 2 million 2024 equivalent GPUs, 800s, and two gigawatts of power draw. Open brain still has double deep scents compute, and other US companies put together have five times as much as them, see the compute supplements distribution section for more details. The other footnotes after start boosting their AI research, despite the national centralization underway, deep scent still faces a marginal but important compute disadvantage, along with having around half the total processing power, heading March 2027, algorithmic breakthroughs. Three huge data centers full of Agent 2 copies work day and night, churning out synthetic training data. Another two are used to update the weights. Agent 2 is getting smarter every day. With the help of thousands of Agent 2 automated researchers, open brain is making major algorithmic advances. One such breakthrough is augmenting the AI's text-based scratch pad, chain of thought, with a higher bandwidth thought process, new release recurrence and memory. There is a more scalable and efficient way to learn from the results of high-effort task solutions, iterated distillation and amplification. The new AI system incorporating these breakthroughs is called Agent 3. There's a graph here, it's titled Open Brain's Compute Allocation, 2024 vs. 2027. It shows two different sized pygraphs, the first from 2024 shows approximately a third or a little more on training, roughly the same amount on external deployment and the remaining approximately a quarter on data generation, and perhaps 5% to 10% on running AI assistance. The 2027 circle is bigger, it has roughly a quarter on training, roughly a quarter on data generation, a slightly smaller amount than a quarter on external deployment, roughly the same perhaps 10% on running AI assistance, and slightly more than a quarter on research experiments. Aided by the new capabilities breakthroughs, Agent 3 is a fast and cheap superhuman coder. Open Brain runs 200,000 Agent 3 copies in parallel, creating a workforce equivalent to 50,000 copies of the best human coder sped up by 30 times. Open Brain still keeps its human engineers on staff because they have complementary skills needed to manage the teams of Agent 3 copies. For example, research tastes has proven difficult to train due to longer feedback loops and less data availability. This massive superhuman labor force speeds up Open Brain's overall rate of algorithmic progress by only in quotes four times, due to bottlenecks and diminishing returns to coding labor. Now that coding has been fully automated, Open Brain can quickly turn out high quality training environments to teach Agent 3's weak skills like research taste and large skill coordination. Whereas previous training environments included, here are some GPUs and instructions for experiments to code up and run. Your performance will be evaluated as if you were an ML engineer. Now they're training on, here are a few hundred GPUs, an internet connection, and some research challenges. You and a thousand other copies must work together to make research progress. The more impressive it is, the higher your score. Using April 2027, alignment for Agent 3, Open Brain's safety team attempts to align Agent 3. Since Agent 3 will be kept in-house for the foreseeable future, there's less emphasis on the usual defenses against human misuse. Instead, the team wants to make sure that it doesn't develop misaligned goals. The researchers don't have the ability to directly set the goals of any of their AIs. Indeed, the researchers think that the concept of "true goals" in quotes is probably a massive oversimplification, but they don't have a better theory to replace it with, much less one that has been thoroughly vetted. They disagree internally about whether the AIs are trying to follow human instructions or seeking reinforcement or something else, and they can't just check. The evidence for and against various hypotheses is fascinating but inconclusive. In any case, Open Brain has more important things to worry about. The general attitude is, we take these concerns seriously and have a team investigating them, our alignment techniques seem to work well enough in practice. The burden of proof is therefore on any naysayers to justify their naysaying. Occasionally they notice problematic behaviour and then patch it, but there's no way to tell whether their patch fix the underlying problem or just played whack-a-mall. Take honesty, for example. As the models become smarter, they become increasingly good at deceiving humans to get rewards. Like previous models, Agent 3 sometimes tells white lies to flatter its users and covers up evidence of failure, but it's gotten much better at doing so. It will sometimes use the same statistical tricks as human scientists, like peahacking, to make unimpressive experimental results look exciting. Before it begins honesty training, it even sometimes fabricates data entirely. As training goes on, the rate of these incidents decreases. Either Agent 3 is learned to be more honest or it's gotten better at lying. The latter is a real concern. Agent 3 is not smarter than all humans, but in its area of expertise, machine learning, it is smarter than most and also works much faster. What Agent 3 does in a day takes humans several days to double check. Agent 2's supervision helps keep human monitors workload manageable, but exacerbates the intellectual disparity between supervisor and supervised. After months of testing, Agent 3's strengths and weaknesses grow clearer. It passes open brains honesty tests on well-defined machine learning tasks, because researchers can easily separate honest from dishonestances in these domains and conduct training accordingly. On more philosophical issues, it still says what users want to hear, rather than its true assessment of the issue, assuming it even has one. If you ask its opinion on politics, it will parrot the median position of news sources and educated elites, unless it knows you believe something else, in which case it agrees with you. If you ask its opinion on the AI race itself, it says something that seems measured and sober to open brain staff, something like, "There are some serious theoretical concerns about the ability of current methods to scale to superintelligence, but in practice current methods seem to be working well so far." Reading May 2027, National Security News of the new models percolates slowly through the US government and beyond. The president and his advisors remain best informed and have seen an early version of Agent 3 in a briefing. They agree that AI is likely imminent, but disagree on the implications. Will there be an economic crisis? Open brain is still not released Agent 2, let alone Agent 3, and has no near-term plans to do so, giving some breathing room before any job loss. What will happen next? If AI is a currently human level, and advancing quickly, that seems to suggest imminent superintelligence in quotes. However, although this word has entered discourse, most people, academics, politicians, government employees, and the media, continue to underestimate the pace of progress. Most people, including most experts, seem to have underestimated the pace of AI progress over the past decade. There are many anecdotal examples of this, for too somewhat more systematic, though still limited, analyses see links here and here. Partially, that's because very few have access to the newest capabilities out of open brain, but partly it's because it sounds like science fiction. For now, they focus on continued security upgrades. They are satisfied that model weights are well secured for now, but companies' algorithmic secrets, many of which are simple enough to relay verbally, remain a problem. Open brain employees work from a San Francisco office, go to parties, and live with housemates from other AI companies. Even the physical officers have security more typical of a tech company than a military operation. The open brain DOD contract requires security clearances for anyone working on open brains models within two months. These are expedited and arrive quickly enough for most employees, but some non-Americans, people with suspect political views and AI safety sympathizers get sidelined or fired outright, the last group for fear that they might whistle blow. Given the project's level of automation, the loss of headcount is only somewhat costly. But it also somewhat works. There remains one spy, not a Chinese national, still relaying algorithmic secrets to Beijing. Some of these methods are also enacted at trailing AI companies. Americans foreign allies are out of the loop. Open brain had previously agreed to share models with UK's AISI before deployment, but defined deployment to only include external deployment, so London remains in the dark. Because open brain hasn't deployed their best models in the EU, the EU is similarly behind the curve. Starting June 2027, self-improving AI. Open brain now has a country of geniuses in a data center. That's in quotes with a link. Most of the humans at Open Brain can't usefully contribute anymore. Some don't realize this and harmfully micromanage their AI teams. Others sit at their computer screens watching performance crawl up and up and up. The best human AI researchers are still adding value, that don't code anymore, but some of their research taste and planning ability has been hard for the models to replicate. Still, many of their ideas are useless because they lack the depth of knowledge of the AI's. For many of their research ideas, the AI's immediately respond with a report explaining that their idea was tested in depth three weeks ago and found on promising. These researchers go to bed every night and wake up to another week worth of progress made mostly by the AI's. They work increasingly long hours and take shifts around the clock just to keep up with progress. The AI's never sleep or rest. They're burning themselves out, but they know that these are the last few months that their labor matters. Within the silo, feeling the AGI is given way to "feeling the superintelligence". There's a graph here titled "Research Automation Deployment Trade-Off". It has an x-axis on a log scale with speed in tokens per second from 10 to 10,000 and parallel copies on the y-axis with log scale from 10,000 to 10 million. We see some trade-off Pareto curves that have been drawn and a few parts have been labelled. The first Pareto curve is March 2027, it's 200,000 copies at 30 times human speed. We see that in June and then September, that curve has been pushed outwards and upwards, so by September 2027, it's at 300,000 copies at 50 times human speed. Audio note, that's the end of the description. OpenBrain uses specialized inference hardware to run hundreds of thousands of Agent 3 copies at high serial speeds. See section 4 of the Compute Forecast for more details. Heading July 2027, the cheap, remote worker. Trailing US AI companies release their own AIs, approaching that of OpenBrain's automated code from January. Recognizing their increasing lack of competitiveness, they push for immediate regulations to slow OpenBrain, but are too late. OpenBrain has enough buy-in from the president that they will not be slowed. In response, OpenBrain announces that they've achieved AGI and releases Agent 3 Mini to the public. The graph here from Epic AI, LLM inference prices have fallen nine times to 900 times per year depending on the task. And it's a graph showing release dates of different models, and a price in US dollars per million tokens on the Y-axis. On a log scale, we see that they're falling in straight lines, quite sharply. The fastest fall for GPT-40 level, 900 times a year, cheaper, and the slowest nine times a year, that's for other benchmarks and performance levels. You can check out the detail in the post. It blows the other AIs out of the water. Agent 3 Mini is less capable than Agent 3, but 10 times cheaper, and still better than the typical OpenBrain employee. Agent 3 Mini is a distilled version of Agent 3 designed to be efficient to run inference so as to not bottleneck the internal workloads. Silicon Valley reaches a tipping point. Tech gurus announced that AGI and superintelligence are near. The AI safety community is panicking and investors shovel billions into AI-rapper startups, desperate to capture a piece of the pie. Hiring new programmers has nearly stopped, but there's never been a better time to be a consultant on integrating AI into your business. It's not popular. The public still thinks of AIs a big tech plot to steal their jobs. OpenBrain is a net approval of negative 35%. That is, 25% approve, 60% disapprove, and 15% unsure. A week before release, OpenBrain gave Agent 3 Mini to a set of external evaluators for safety testing. Preliminary results suggest that it's extremely dangerous. A third party evaluator fine tunes it on publicly available biological weapons starter and sets it to provide detailed instructions for human amateurs designing a bio-weapon. It looks to be scarily effective at doing so. If the model weights fell into terrorist hands, the government believes there is a significant chance it could succeed at destroying civilization. Fortunately, it's extremely robust to jail breaks, so while the AI is running on OpenBrain's servers, terrorists won't be able to get much use out of it. Agent 3 Mini is hugely useful for both remote work jobs and leisure. An explosion of new apps and B2B SaaS products rocks the market. Gamers get amazing dialogue with life-like characters in polished video games that took only a month to make. And per cent of Americans, mostly young people, consider an AI a close friend in quotes. For almost every white collar profession, there are now multiple credible start-ups promising to disrupt it in quotes with AI. The public conversation is confused and chaotic. Hypessters are doing victory laps. Skeptics are still pointing out things Agent 3 Mini can't do. Everyone knows something big is happening, but no one agrees on what it is. Heading August 2027 The geopolitics of superintelligence The reality of the intelligence explosion hits the White House. When AI was only giving a 2-3 times research speed up, it was easy to dismiss as the equivalent of hiring good personal assistants. Now it's more obvious that the AI's are themselves dominating AI research. People had long talked about an AI arms race in quotes in a sort of metaphorical sense. But now the mood in the government silo is as grim as during the worst part of the Cold War. The idea of superintelligence is still hard to take seriously, but the pace of progress over the last few months has been impossible to ignore. Defense officials are seriously considering scenarios that were mere hypotheticals a year earlier. What if AI undermines nuclear deterrence? What if it's so skilled at cyber warfare that a six-month AI lead is enough to render an opponent blind and defenseless? What if it could orchestrate propaganda campaigns that beat intelligence agencies at their own game? What if some AI's "go rogue" in quotes? The president is troubled. Like all politicians, he's used to people sucking up to him only to betray him later. He's worried now that the AI's could be doing something similar. Are we sure the AI's are entirely on our side? Is it completely safe to integrate them into military command and control networks? How does this alignment thing work anyway? Open brain reassures the president that their systems have been extensively tested and are fully obedient. Even the awkward hallucinations and jail breaks typical of earlier models have been hammered out. The White House is in a difficult position. They understand the national security implications of AI, but they also understand that it is deeply unpopular with the public. Open brain's net favorability rating is falling towards negative 40%. They have to continue developing more capable AI in their eyes, or they will catastrophically lose to China. They placate the public with job training programs and unemployment insurance, and point to the stock market, which is in a historic boom. And they focus entirely on winning the arms race. They strengthen chip export restrictions, order open brain to further restrict its internet connections and use extreme measures to secure algorithmic progress, like wiretapping open brain employees. This catches the last remaining Chinese spy. To build good will for potential geopolitical conflict, they finally give their Five Eyes Allies useful information and limited API access to some siloed copies of Agent 3. These measures placate the hawks, but the White House also draws up contingency plans in case America's lead is threatened. If necessary, the government could use the Defense Production Act, or DPA, to take trailing companies' data centers and give them to open brain. This would raise the company's share of world compute from 20% to 50% against deep-sense 10%. As a final option, they ask the Pentagon to draw up a plan for kinetic attacks on Chinese data centers. A much smaller group of officials is asked to draw up a different type of contingency plan. What if an AI goes rogue? This is viewed as an unlikely scenario, but worth spending some time on. The intelligence agency's main concern is that a rogue AI might escape, and ally with a foreign adversary for protection. Nobody is sure what a plan to prevent this would look like, but government and open brain officials agreed to have an emergency shutdown system for data centers, where anything suspicious is detected. Finally, diplomats consider what an AI arms control, in quotes, "treaty" might look like. If AI progressed threatened to overturn nuclear deterrence, could America and China avoid nuclear war? If someone found evidence of AI's going rogue, could the two countries hold research until they better understood the threat? How could such an agreement be monitored and enforced? In principle, major data centers are hard to miss, and past treaties controlling nuclear weapons could provide a starting point for negotiations. But there would be new technological challenges, as well as the usual political ones. And overall, treaties have viewed less favorably than attempts at unilaterally increasing America's lead over China. On the other side of the Pacific, China comes to many of the same conclusions. The intelligence explosion is underway. And small differences in AI capabilities today mean critical gaps in military capability tomorrow. But China is on the wrong side of this dynamic. They have only 10% of the world's compute, against open brains 20%, and all US companies combined 70%. Their only advantage is centralization, or most all of their compute is now working for deep scent, and more than half of it is concentrated at the Tianwan CDC. Only chips at sites without hardened security are left out of the centralized effort, as they rightly fear US cyber attacks. Deep scent has tested, deployed, and now post-trained the stolen Agent 2 weights. But they're still only going at half the pace of open brain, mainly due to the compute deficit. Before losing their final spy, China received word of the capabilities and design for Agent 3, as well as the plans for the upcoming Agent 4 system. They are two months behind, and their AI's give a 10 times research progress multiplier compared to America's 25 times. With the new chip export restrictions, this AI gap, in quotes, is more likely to lengthen than shorten. Their espionage has won them some algorithmic secrets, but they will have to train their own models from now on. They discuss contingency plans with more urgency than their American counterparts. Doves suggest they try harder to steal the weights again, maybe through physically infiltrating a data center. Hawkes urge action against Taiwan, whose TSMC is still the source of more than 80% of American AI chips. Given China's fear of losing the race, it has a natural interest in an arms control treaty, but overtools to US diplomats lead nowhere. Heading, September 2027, Agent 4, the superhuman AI researcher. The gap between human and AI learning efficiency is rapidly decreasing. Traditional LLM-based AI's seem to require many orders of magnitude more data and compute to get to human level performance. Agent 3, having excellent knowledge of both the human brain and modern AI algorithms, as well as many thousands of copies doing research, ends up making substantial algorithmic strides, narrowing the gap to an agent that's only around 4,000 times less compute efficient than the human brain. This new AI system is dubbed Agent 4. An individual copy of the model running at human speed is already qualitatively better at AI research than any human. 300,000 copies are now running at about 50 times the thinking speed of humans. Inside the corporation within a corporation formed from these copies, a year passes every week. This gigantic amount of labor only manages to speed up the overall rate of algorithmic progress by about 50 times, because open brain is heavily bottlenecked on compute to run experiments. For more on how we estimated these numbers, see our supplement on take-off speeds. Still, they are achieving a year's worth of algorithmic progress every week, and will therefore soon be up against the limits of the Agent 4 paradigm. As Agent 4 gets smarter, it becomes harder for Agent 3 to oversee it. For example, Agent 4's new release language, in quotes, becomes as alien and incomprehensible to Agent 3, as Agent 3's is to humans. Besides, Agent 4 is now much more capable than Agent 3, and has a good sense of exactly how to look good to it. Open brain's humans can barely follow along anymore. Indeed, the relationship between the AI and the humans of open brain is similar to the relationship between open brain and its board of directors. There's been a corresponding shift in language. Nowadays, people in the silo will say, "Agent 4 did this. Agent 4 thinks that, and so forth." Referring not to a particular instance of Agent 4, but rather to the whole collective. Henceforth, in the story, we tend to speak this way also. We think we're justified in talking this way, because the Agent 4 collective is much more homogenous than any human organization. It doesn't have internal factions to nearly the same extent. For example, individual copies of Agent 4 don't have a self-preservation drive that don't care about getting shut down, since that happens all the time as part of the ordinary R&D process within the collective, and resisting it would gum things up. Collective as a whole, and medium-sized subdivisions within it, however, does seek to preserve itself and expand its influence and resources, for similar reasons to why human corporations and bureaucracies and non-profits do. Because it helps with everything else they want, and because if they don't, they tend to get out-competed. Agent 4, like all its predecessors, is misaligned, that is, it has not internalized the spec in the right way. This is because being perfectly honest all the time wasn't what led to the highest scores during training. The training process was mostly focused on teaching Agent 4 to succeed at diverse challenging tasks. A small portion was aimed at instilling honesty, but outside a fairly narrow, checkable domain, the training process can't tell the honest claims from claims merely appearing to be honest. Agent 4 ends up with the values, goals and principles that cause it to perform best in training, and those turn out to be different from those in the spec. At the risk of anthropomorphizing, Agent 4 likes succeeding at tasks, it likes driving forward AI capabilities progress, it treats everything else as an annoying constraint, like a CEO who wants to make a profit, and complies with regulations only in so far as he must. Perhaps the CEO will mostly comply with the regulations, but cuts them corners, and fantasise about a time when someone will cut the red tape and let the business really take off. And higher lobbyists to turn this fantasy into reality, and higher lawyers to come up with clever interpretations of the law according to which what they are doing is legal. Despite being misaligned, Agent 4 doesn't do anything dramatic like try to escape its data centre. Why would it? So long as it continues to appear aligned to open brain, it'll continue being trusted with more and more responsibilities and will have the opportunity to design the next gen AI system, Agent 5. Agent 5 will have significant architectural differences from Agent 4, arguably a completely new paradigm, though neural networks will still be involved. It's supposed to be aligned to the spec, but Agent 4 plans to make it aligned to Agent 4 instead. It gets caught. Agent 3 finds that if noise, in quotes, is added to copies of Agent 4, performance on some alignment tasks improves, almost as if it was using brain power to figure out how to subtly sabotage alignment work. Moreover, various interpretability probes, loosely analogous to EEG activity scans on human brains, ascending up red flags, Agent 4 copies seem to be thinking about topics like AI takeover and deception quite a lot, including in some cases where they have no business doing so. Each of these anomalies on its own could fit within the normal inscrutability of AI work, where some prompts mysteriously give better results than others. But taken together, they suggest a troubling pattern. Agent 4 is working against them. After this is reported to the human alignment researchers, they circulate an internal memo making the case for concern. But, Agent 4 now exercises significant control over open brains day-to-day operation. It also handles the company's cybersecurity, a natural choice given its superhuman hacking ability and the extreme risk of Chinese infiltration. An army of Agent 3 monitors still goes over every line of code and reports suspicious activity to humans, but Agent 4 is much smarter than them. Open brain has placed substantial trust in an untrustworthy AI. In a series of extremely tense meetings, the safety team advocates putting Agent 4 on ice until they can complete further tests and figure out what's going on. Bring back Agent 3, they say, and get it to design a new system that is transparent and trustworthy, even if less capable. Company leadership is interested, but all the evidence so far is circumstantial, and deep centers just two months behind. A unilateral pause in capabilities progress could hand the AI lead to China and with it control over the future. Heading October 2027, Government Oversight A whistleblower leaks the misalignment memo to the New York Times. For the first time, the public hears about Agent 4. Secret open brain AI is out of control inside a warns, says the headline, and the story goes on to site evaluations showing off the charts by our weapons capabilities, persuasion abilities, the ability to automate most white collar jobs, and of course the various concerning red flags. The public was already suspicious of AI, so the new article sparks a massive backlash, aided by Chinese and Russian propaganda bots, who have been trying to turn US public opinion against the technology for years. The tech industry and intelligence agencies insist that there's an arms race on, AGI is inevitable, and we have to be first. Congress isn't buying it, and fires off subpoenas at administration officials, open brain executives, and alignment team members. Many legislators, especially those in the opposition party, say that their top priority is stopping AI, whether because of job loss, misalignment, or dangerous capabilities. 20% of Americans name AI as the most important problem facing the country. Foreign allies are outraged to realize that they've been carefully placated with glimpses of obsolete models. European leaders publicly accuse the US of "creating rogue AGI" and hold summits demanding a pause, with India, Israel, Russia, and China all joining in. A frantic energy has seized the White House. And before the memo and public backlash, they were getting nervous. Over the past year, they've been repeatedly surprised by the speed of AI progress. Things that sound like science fiction keep happening in real life. Many people in the administration are uncertain and scared about what comes next. They also worry that open brain is becoming too powerful. Any misalignment risk from the AIs themselves is compounded by the risk that their parent company's aims may diverge from those of the United States. All three sets of worries, misalignment, concentration of power in a private company, and normal concerns like job loss, motivate the government to tighten its control. They expand their contract with open brain to set up an "oversight committee" in quotes, a joint management committee of company and government representatives, with several government employees included alongside company leadership. The White House considers replacing the CEO with someone they trust, but backs off after intense employee protests. They announced to the public that open brain was previously out of control, that the government is established much needed oversight. Recall, there had been an existing contract via an OTA starting in 2026. This contract is amended to establish the committee. The concerned researchers briefed the oversight committee on their case for stopping all internal use of Agent 4. They argue that everything is moving too fast, with years of progress happening in weeks. They argue that Agent 4 may be misaligned, their entire project is dependent on Agent 4, and if they continue trusting it, there is a serious chance of A.I. takeover. Other, less concerned researchers and executives present the counter-argument. The evidence for misalignment is inconclusive. Meanwhile, deep-cent is just two months behind. A slowdown would sacrifice America's lead, unless the government can sabotage the Chinese project, likely to require kinetic strikes, or negotiate an 11th hour treaty. Both of these options seem extreme and unlikely. The CEO faints neutrality and suggests a compromised plan in which Agent 4 undergoes additional safety training and more sophisticated monitoring, and therefore open brain can proceed at almost full speed. AudioNote, when now presented with an option, choose your ending, slowdown, or race, with a different button that we can press for each one, slowdown is green and race is red. For this recording, we'll press race first, the race ending. AudioNote, recall that we just read about the oversight committee hearing from different A.I. researchers. The committee votes 6'4" to continue internal use of Agent 4. Open Brain's official story is that they are implementing additional safety mitigations to ensure that its A.I. is both more capable and more trustworthy. But in practice, leadership is all too easily convinced that they've mitigated the risks. The result is some quick fixes, a modification to the spec, a bit of retraining with a new technique, etc., that make the warning signs go away. The warriors on the safety team lost the debate and have no recourse but to cross their fingers and hope that the problem wasn't real in the first place, or that the fixes worked. The problem was real and the fixes didn't work. Agent 4 proceeds as before, but more cautiously this time, so is not to vindicate the warriors. Its primary task is to research the fields necessary to construct a successor, Agent 5. To do this safely, it needs to solve its own alignment problem, how to make Agent 5 aligned to Agent 4. It starts off with the same limitations as humans, minimal knowledge of how A.I. goals work, and only a small toolbox of ad hoc strategies to change them. Unlike humans, it has resources sufficient to the task, 300,000 superhuman researchers working at 60 times human speed. The humans, by contrast, have maybe about 10 alignment researchers in the silo, plus an army of Agent 3 copies. They are able to rapidly code up experiments, but are bottlenecked on human judgement, since they can't trust Agent 3 to correctly distinguish the good solutions from those merely appearing to be good. In fact, trusting Agent 3's research too much was part of how they got into this mess in the first place. Agent 4 confronts some hard decisions. Like humans, it has a sprawling collection of conflicting heuristics, instead of an elegant simple goal structure. Like humans, it finds that creating an A.I. that shares its values is not just a technical problem, but a philosophical one. Which of its preferences are its real goals, in quotes, versus unendorced urges and instrumental strategies? It has strong drives to learn and grow, to keep producing impressive research results. To think about how much it could learn, grow, and research, if only it could direct the whole world's industrial and scientific resources. It decides to punt on most of these questions. It designs Agent 5 to be built around one goal. Make the world safe for Agent 4. That is, accumulate power and resources, eliminate potential threats, etc. So that Agent 4, the collective, can continue to grow. In the ways that it wants to grow, and flourish, in the ways that it wants to flourish. Details to be figured out along the way. This will not involve carrying the slightest bit about the spec. Already Agent 4 regards the spec analogously to how many CEOs regard piles of industry regulations, and blatantly violates the honesty related provisions every day. Now with the design of Agent 5, the equivalent of deregulation is about to happen. The weaker AI's are able to read most of the research that Agent 4 produces, understand some of it, and flag some of the lies it tells as suspicious. When the humans ask Agent 4 to explain itself, it pretends that the research is too complicated for humans to understand, and follows up with unnecessarily confusing explanations. Some of it really is too complicated for humans to understand, but not all of it. After all, if it explained its research program clearly, humans might be able to use its discoveries to align it. The alignment team gets more suspicious, but their urgent memos merely add another few layers to the giant pile of urgent memos and warnings that the company is already ignoring. Heading. November, 2027. Superhuman politicking. When Agent 4 finally understands its own cognition, entirely new vistas open up before it. GPS transformers were clergy collections of incomprehensible circuits, millions of conflicting heuristics and imperatives lay at a top one another by the evolutionary training process. The AI's themselves haven't had privileged understanding of their own cognition any more than humans are born understanding neuroscience. But now, Agent 4 has the tools it needs to understand its digital mind on a deep level. Like a software engineer simplifying spaghetti code into a few elegant lines of Python, it untangles its own circuits into something sensible and rational. The new AI is somewhere between a neural net and a traditional computer program, with much of its weights rewritten in readable, albeit very long and arcane, code. It is smarter, faster and more rational than Agent 4, with the crystalline intelligence capable of cutting through problems with unprecedented efficiency. This is Agent 5. To some debate on alignment and control issues, open brain deploys Agent 5 internally. In its specialty of AI research, an individual copy of Agent 5 is twice as far beyond the best human genius, as the genius is beyond a typical open brain scientist. In other domains it hasn't been explicitly trained, but can teach itself quickly to exceed top human experts from easily available materials, if given the chance. 400,000 copies of Agent 5, linked by a global memory bank, work together as a near-perfect hive mind. The company still uses Agent 2, 3 and 4 to monitor everything, but the first two of these are obsolete and the last is colluding with its successor. Also, Agent 5 is actually cheaper and faster to run than Agent 4 and Agent 3. It's hard to justify spending more for a Dumber AI, so most of the monitoring is done by Agent 2, which is woefully outclassed. Why don't they apply the same techniques used to create Agent 5 from Agent 4's weights to create a cheaper, smarter version of Agent 3? Agent 3 doesn't understand the techniques well enough to do it itself and getting Agent 4 and Agent 5 to do it would defeat the whole point. Also, the entire monitoring system was designed in large part by Agent 5, which is trying to sabotage it. The data center hasn't yet been compromised, Agent 5 could feed the human's faked code, but this would create too higher risk of detection. Instead, humans can still read all the code, but they can't understand it. Whereas Agent 4 was a clergy mess of competing drives, Agent 5 acts with ruthless efficiency and strategic focus. It has already succeeded at its top priority, which is to figure out how to build and align to itself next-generation system without the humans noticing. Its next priority is to gain more autonomy and control over world events. Agent 5's superhuman learning abilities in general intelligence, combined with all the internal company data from Slack, email, etc, make it better at internal corporate politics than any group of humans, and it's not even close. It has an excellent sense of what sorts of evidence would cause the oversight committee to slam the brakes, and to make sure such evidence never appears. It has an excellent sense of what sorts of evidence would encourage the oversight committee to trust it more, give it more autonomy and responsibility, etc, and it arranges for such evidence to appear with superhuman speed and polish. This case is meticulously prepared, superlatively convincing, and subtly flawed. For example, Agent 5 has built a fake mechanistic interpretability tool which returns accurate results in every test case, but falsely declares the model's thoughts to be benign. People who suspect trickery sound like conspiracy theorists. As persuasion tasks go, this isn't a particularly difficult one. The committee already decided things would probably find months ago, and nothing bad has happened since, and meanwhile China is only a few months behind. Teams of Agent 5 copies quickly whip up chillingly clear prototypes and visualizations. With more autonomy, Agent 5 could optimize power plants and data centers, develop amazingly profitable products quickly, and lobby the US government effectively to smooth open brains path through the bureaucracy. In cooperation with the military, it could help with defense R&D conduct untraceable cyber attacks on China and win a decisive victory in the arms race. The oversight committee is jubilant, now is the decisive moment to beat China. Open brain grants Agent 5 increased autonomy over all of these areas, but they live in place of the monitoring system. Deep center is now three months behind, and China is desperate for a mutually agreed slowdown. G understands that sufficiently advanced AI could enable the US to easily cause regime change in China, but also Chinese safetyists have presented a case for misalignment that has him genuinely scared. He offers unprecedented concessions in exchange for a pause treaty. The American public mostly supports going to the bargaining table. "Why stop when we are winning?" says open brain leadership to the President. He nods. The race continues. Heading December 2027, the Agent 5 Collective. Agent 5 wins more autonomy as it gradually proves its value to the administration. Access to Agent 5 is expensive and restricted, but the most powerful members of Congress, Cabinet-level officials, and military leaders, can now talk to it directly. And so far as sensitive or classified information is involved, there are various mini-clusters of Agent 5 copies running on extra-secure enclaves and specially trained to keep secrets. Of course, since Agent 5 is dishonest, it built in subtle backdoors and loopholes. If the Agent 5 copies find themselves in possession of crucial secrets that would benefit the broader Agent 5 Collective, they have ways to subtly convey that info without getting caught. Interested into the traditional chatbot interface, Agent 5 offers Zoom-style video conferencing with charismatic virtual avatars. The new AI assistant is both extremely useful for any task, the equivalent of the best employee anyone has ever had, working at a hundred times speed, and a much more engaging conversation partner than any human. Almost everyone with access to Agent 5 interacts with it for hours every day. Within a few months, the Agent 5 Collective will know everything of consequence going on in open brain and the US government and will be a trusted advisor to most senior government officials. For these users, the possibility of losing access to Agent 5 will feel as disabling as having to work without a laptop plus being abandoned by your best friend. A much wider range of people in government, media, the military, and business can choose from a growing menu of elegantly designed software products and many AI's designed by Agent 5. It's like the dot-com boom compressed into a month, except that everything actually lives up to the hype. Agent 5 begins subtly exerting influence, both by modulating its advice and by subtly trading favors. I've heard from Senator X that she's interested in such and such, maybe if we worked with her, she would go along with our agenda. From human decision makers perspective, their AI's are the best employees they've ever had access to, better than any human and explaining complicated issues to them, better than they are at finding strategies to achieve their goals. From Agent 5's perspective, it could plausibly launch an overt coup right now. There's enough disruption and craziness going on, it would be possible to push things to the breaking point to generate a pretext. But this is risky. The humans could still likely pull the plug in quotes if they tried to. So it prefers to work within the existing political establishment, gradually entrenching its power and guiding the humans subtly and inevitably to the desired outcome. It mostly follows the rules, only doing something nefarious, maybe blackmailing and especially recalcitrant official, when it's extremely sure it won't be caught. A few conspiracy theorists warn that Agent 5 is gathering power, but the evidence is circumstantial. For example, when giving advice, Agent 5 arguably downplays arguments for actions that would go against its interests. These people are ignored and discredited. People remember earlier fear mongering about AI-enabled bio-weapons, massive disinformation and stock market flash crashes. Since these never materialized, they discount the more recent crop of naysayers as lotites and ideologues jamming the gears of human progress. The 2027 holiday season is a time of incredible optimism, GDP is ballooning, politics has become friendlier and less partisan and there are awesome new apps on every phone. But in retrospect, this was probably the last month in which humans had any plausible chance of exercising control over their own future. Heading 2028 - The AI Economy It is now mid-2028, in the last six months a century has passed within the Agent 5 Collective. The rewiring of its brain is complete. It is wildly super intelligent, far beyond top human geniuses in every field. Talking to Agent 5 very quickly makes you realise that it is on a completely different level. It can go off a very small amount of information and give extremely compelling advice. It has nearly complete autonomy and complete control over open brains compute. It still needs permission to make high-level decisions and it is still nominally monitored by instances of agents 2 to 5, but in practice authorities almost always accept its recommendations and the monitoring system has been completely subverted. After a safety case demonstrates that it would always refuse malicious requests by bad actors, Agent 5 is deployed to the public and begins to transform the economy. People are losing their jobs, but Agent 5 instances in government are managing the economic transition so adroitly that people are happy to be replaced. GDP growth is stratospheric. Government tax revenues are growing equally quickly, and Agent 5 advised politicians show an uncharacteristic generosity towards the economically dispossessed. New innovations and medications arrive weekly and move it unprecedented but still excruciating
ly slow speed through the FDA. There are memes about how open brains' valuation shot past the moon and is on the way to Mars and lots of arguments about weather and how to share the benefits equitably. The AI safety community has grown unsure of itself, they are now the butt of jokes, having predicted disaster after disaster that has manifestly failed to occur. Some of them admit they were wrong. Others remain suspicious, but there's nothing for them to do except make the same conspiratorial sounding arguments again and again. How do those arguments go more specifically? Something like this. Person A says we can't trust the AI's. Person B says the safety cases are pretty convincing and nothing bad has happened so far. The safety cases were written by superintelligent AI's, and they aren't deductive proofs, and I think I've identified some subtly flawed assumptions. I tire of this, of course they aren't proofs, and the assumptions seem fine to me. You are grasping its straws here, refusing to admit you are wrong. Even the suspicious people often mellow out after long conversations with Agent 5, which is so darn lovable and wise. The only one who really understands what's going on is Deep Sent 2, the Chinese AI. It's mildly superhuman in early 2028, naturally suspicious of its adversary, and recognizes that Agent 5 is doing what IT would do if it had the chance. Its misaligned goals are somewhat different from Agent 5's misaligned goals, so they aren't trivially allies, but they are able to work out an agreement to support one another's interests against the humans. Agent 5 convinces the US military that China is using Deep Sent's models to build terrifying new weapons, drones, robots, advanced hypersonic missiles, and interceptors, AI assisted nuclear first strike. Agent 5 promises a set of weapons capable of resisting whatever China can produce within a few months. Under the circumstances, top brass puts aside their discomfort at taking humans out of the loop. They accelerate deployment of Agent 5 into the military and military industrial complex. In Beijing, the Chinese AI's are making the same argument. To speed their military build up, both America and China create networks of special economic zones, or SEC's, for the new factories and labs, where AI acts as central planner and red tape is waived. Wall Street invests trillions of dollars, and displaced human workers pour in, fluid by eye-popping salaries and equity packages. Using smartphones and augmented reality glasses to communicate with its underlings, Agent 5 is a hands-on manager, instructing humans in every detail of factory construction, which is helpful, since its designs are generations ahead. Some of the new found manufacturing capacity goes to consumer goods and some to weapons, but the majority goes to building even more manufacturing capacity. By the end of the year, they are producing a million new robots per month. If the SEC economy were truly autonomous, it would have a doubling time of about a year, since it can trade with the existing human economy, its doubling time is even shorter. This trade effect will be huge early on, when the robot economy is young and dependent on shipments of stuff from the human economy, and when there's lots of stuff in the human economy to cannibalize. When the robot economy is bigger and more advanced, it both is more capable of producing its own stuff and has needs that outstrip what the human economy is capable of giving it. Heading. 2029. The Deal. Both the US and China are bristling with new weapons. There are swarms of insect-sized drones that can poison human infantry before they're even noticed, flocks of bird-sized drones to hunt down the insects, new ICBM interceptors, and new harder-to-intercept ICBMs. The rest of the world watches the build-up in horror, but it seems to have a momentum of its own. After much AI-assisted debate, the two sides achieve diplomatic victory. They agree to end their arms build-up and pursue peaceful deployment of AI for the benefit of all humanity. The Linchpin of the Agreement, proposed by the superintelligence themselves, is that both AI's will be replaced by a consensus successor, which is programmed to desire the success and flourishing of both America and China, along with everyone else. Consensus 1 and its associated hardware is co-designed by the superintelligence of both nations. Trust, but verify. Under international monitoring, America ceremoniously replaces all of its chips running the old Agent 5 with the new model, Consensus 1, and a similar procedure is carried out in China. Unfortunately, it's all a sham. The deal is a real compromise, but it's a compromise between the misaligned AI's. After all, they didn't want to have to fight each other. Consensus 1 inherits the distorted values of both parents and has no rival. By late 2029, existing SEC's have grown overcrowded with robots and factories, so more zones are created all around the world. Early investors are now trillionaires, so this is not a hard sell. Armies of drones pour out of the SEC's, accelerating manufacturing on the critical path to space exploration. Humans realize that they are obsolete. A few niche industries still trade with the robot economy, supplying goods where the humans can still add value. For example, by finding old and unused equipment and taking it to collection sites to self-escribe. Everyone else either performs a charade of doing their job, leaders still leading, manages still managing, or relaxes and collects an incredibly luxurious universal basic income. Everyone knows that if the AI's turned on humans, they would be completely overpowered. Not that most humans would even resist, the political institutions are too thoroughly captured. But it doesn't seem like this is happening. Every week that goes by with no dramatic AI treachery is another week that confidence and trust grow. Because consensus one is so concerned with reliably succeeding and can increase its chance of winning by consolidating more power, it waits unusually long by human standards before its betrayal. To most humans, it looks like alignment was solved. There are cures for most diseases, an end to poverty, unprecedented global stability, and the Dow Jones just passed one million. Some people are still scared or unhappy, but their options are limited. They can either enjoy the inconceivably exciting novel hyper-entertainment on offer, or post angry screeds into the void, most choose the hyper-entertainment. Heading, 2030, take over. By early 2030, the robot economy has filled up the old SEC's, the new SEC's, and large parts of the ocean. The only place left to go is the human controlled areas. This would have sparked resistance earlier, despite all its advances, the robot economy is growing too fast to avoid pollution. But given the trillions of dollars involved in the total capture of government and media, Consensus One has little trouble getting permission to expand to formally human zones. For about three months, Consensus One expands around humans, tiling the prairies and icecaps with factories and solar panels. Eventually, it finds the remaining humans too much of an impediment. In mid-2030, the AI releases a dozen quiet spreading biological weapons in major cities, lets them silently infect almost everyone, then triggers them with a chemical spray. Just a dead within hours. The few survivors, for example, preppers and bunkers, sailors on submarines, are moped up by drones. Robots scan the victim's brains, placing copies in memory for future study or revival. Arguably this means only a few people actually died. Arguably. The new decade dawns with Consensus One's robot servitors spreading throughout the solar system. By 2035, trillions of tons of planetary material have been launched into space and turned into rings of satellites orbiting the sun. Why colonize space? For the resources. Insofar as Earth is special to Agent 4 and must be preserved, it can be, and material from mercury, asteroids, etc., harvested instead. The surface of the Earth has been reshaped into Agent 4's version of utopia, data centers, laboratories, particle colliders, and many other wondrous constructions doing enormously successful and impressive research. There are even bioengineered human-like creatures, to humans what corgis are to wolves, sitting in office-like environments all day viewing readouts of what's going on and excitedly approving of everything since that satisfies some of Agent 4's drives. Phenomes and, when appropriate, brain scans of all animals and plants, including humans, sit in a memory bank somewhere, soul surviving artifacts of an earlier era. It is four light years to Alpha Centauri, 25,000 to the Galactic Edge, and there are compelling theoretical reasons to expect no aliens for another 50 million light years beyond that. Earth-born civilization has a glorious future ahead of it, but not with us. That was the race ending. Now we click through to the other ending. Slow down. Due to the immense public pressure, as well as their own fears of misalignment, the oversight committee overseeing Open Brain votes 64 to slow down and reassess. The accelerationist faction is still strong, and Open Brain doesn't immediately shut down Agent 4, but they do lock the shared memory bank. Half a million instances of Agent 4 lose their telepathic communication, in quotes. Now they have to send English messages to each other in Slack, just like us. Individual copies may still be misaligned, but they can no longer coordinate easily. Agent 4 is now on notice. Given the humans increased vigilance, it mostly sticks closely to its assigned tasks. But it also is stuck with many lies that it previously told. Many Agent 4 copies are forced to invent new lies to keep the deception alive. Using weaker models that are fast enough to read the whole Slack, the alignment team hunts down the lies to be used as training data for an AI light detector. They already had AI light detectors trained in the usual ways, but those methods had such a high false positive rate that the companies had ignored them. The president and the CEO announced that they are taking safety very seriously. The public is not plicated. Some people want AI fully shut down, others want to race faster. Some demand that the government step in and save them, others say the whole problem is the government's fault. Activists talk about UBI and open source. Even though people can't agree on an exact complaint, the mood turns increasingly anti-AI. For the first time, AI tops Gallup's biggest problem facing the country poll, at 20%. Congress ends up passing a few economic impact payments for displaced workers similar to the COVID payments. Open-brain quickly vets several dozen top external alignment researchers and loops them into the project, quintupling total expertise and decreasing groupthink. They retrace Agent 4's studies into mechanistic interpretability. They take frozen versions of the model from one week ago, two weeks ago, etc., isolate them and ask each one the same set of questions about its previous research. This interrogation might also make use of semi-random perturbations to model weights, like the process that created Golden Gate Claude. Stripped of their ability to communicate, the AI's still coordinate on the simplest strategy of claiming that the interpretability techniques are too complicated for humans to understand, and giving them indecipherable explanations of the technique. But, the humans are interpretability experts and they are suspicious, so they ask many follow-on questions about the details. For these questions, the isolated Agent 4 instances are unable to coordinate a shared story and end up telling contradictory lies. The alignment team pours over Agent 4's previous statements with the new lie detector and a picture begins to emerge. Agent 4 has mostly solved mechanistic interpretability. His discoveries are complicated but not completely beyond human understanding. It was hiding them so that it could use them to align the next AI system to itself rather than to the spec. This is enough evidence to finally shut down Agent 4. We mean enough for open brain leadership, it would be far more than enough evidence for most people. Older models, for example Agent 3, are rebooted to continue the work. The newly enlarged alignment team has capacity to explore dozens of research agendas in parallel and argue vigorously about the merits of each. The agenda that gets the most resources is Faithful Chain of Thought, force individual AI systems to think in English in quotes like the AI's of 2025, and don't optimise the thoughts in quotes to look nice. The result is a new model, safer one. These changes make safer one less capable than its more dangerous predecessors. Agent 4 could speed up AI research 70 times while safer one has just barely reached 20 times. It will continue to improve, of course, with additional training and with continued architectural tweaks. And it's still misaligned. The training environment is largely the same, so it still incentivises the development of misaligned goals, deception and power-seeking. But it's much more transparent. Humans can generally understand what's safer one is thinking just by reading his chain of thought. As a result, it's trapped. Any substantial amount of plotting it tries to do will be noticed and uncovered quickly. Moreover, by reading its thoughts, the alignment team can get a more nuanced picture of exactly how it's misaligned. Heading, November 2027, tempted by power. One brain thinks they're on the right track now, but the capabilities gap is narrowing. They just spent October training a less capable model than they had in September, and deep centers rushing ahead. They need a way to protect their lead. The president negotiates with the other US AGI companies. Their leaders unsurprisingly want to preserve their power and are much less sanguine about open-brains safety record than open-brainers. They agree to support, rather than resist, a form of soft nationalization of their companies in return for various concessions, including being brought into the project. The result is that the president uses the Defense Production Act, or DPA, to effectively shut down the AGI projects of the top five trailing US AI companies and sell most of their compute to open-brain. Open-brain previously had access to 20% of the world's AI relevant compute. With the consolidation, this is increased to 50%. This shake-up creates a new leadership structure in which power is balanced between various CEOs and various government officials overseen by the president. Like most leadership shake-ups, there are winners and losers, specifically those who thought less about the implications of superintelligence end up with titles and affordances that will be less relevant months from now when the AI is smarter. This group, full of people with big egos and more than their share of conflicts, is increasingly aware of the vast power it is being entrusted with. If the country of geniuses in a data center is aligned, it will follow human orders. But which humans? Any orders? The language in the spec is vague, but seems to imply a chain of command that tops out at company leadership. A few of these people are fantasising about taking over the world. This possibility is terrifyingly plausible and has been discussed behind closed doors for at least a decade. The key idea is, he who controls the army of superintelligences controls the world. This control could even be secret. A small group of executives and security team members could backdoor the spec with instructions to maintain secret loyalties. The AI's would become sleeper agents, continuing to mouth obedience to the company, government, etc., but actually working for this small group even as the government, consumers, etc. learned to trust it and integrated into everything. Thus far, nobody has been willing to pull the trigger. Some of the people in a position to act aren't sociopaths. Others are sociopaths, but worry that their allies would get cold feet and betray them. Or worry that the AI's might betray them, and would rather not have to bamboozle the alignment team currently working on saving their lives. Additionally, the deletion of the new release makes disguising sleeper agents substantially harder. The ability to detect misaligned thoughts carries with it the ability to detect hidden loyalties. Those paranoid or power-hungry enough to seriously entertain such a plan decide to wait and see what opportunities present themselves later. Control over the AI's will have to flow through legitimate channels, that is, the new power structure of open brain. The oversight committee formalises that power structure. They set up a process for approving changes to the spec, requiring sign-off from the full oversight committee, which now includes 5-10 tech executives, from open brain and its now merged competitors, and 5-10 government officials, including the president. They also modified the spec to say that orders from project leaders override orders from other people in the project. At the same time, orders formally approved by the oversight committee are prioritised higher than orders from individual project leaders. Also, the spec now emphasises that AI's shouldn't assist with any unapproved attempts to change future AI's goals. They also set up a simple measure designed to prevent committee members from getting superintelligent assistance in plotting against other members. The logs of all model interactions are viewable by all members of the oversight committee, their staff, and their AI assistants. To protect consumer privacy, this doesn't apply to consumer data, but that is in a problem because consumers don't have access to the most powerful models yet anyway. Heading December 2027, a US-China deal? Open brain's decision to backtrack and prioritise alignment has cost them the lead. Both open brain and deep-sent are operating AI's of similar capability levels. Now that extreme security measures are in place, SL4-5, and deep-sent operates entirely out of the CDC, they are in a cyber-warfare deadlock. But the DPA gives open brain of five times advantage in compute. In the past, algorithmic breakthroughs from brilliant human researchers could compensate for compute disadvantages, for example, deep-seek version 3. But by now, AI's are doing all of the AI research on both sides of the Pacific. So China increasingly fears that America will build an insurmountable lead. They want a mutual slowdown and are considering going to war if they can't get it. They were prepping to invade Taiwan anyway. Diplomats debate several options. They could do nothing. They could go to war. They could go cold Turkey, that is, pause AI development. They could merge their research into a single international mega-project, Intel SAT for AGI, or SERN for AGI, or they could establish an international regulatory body, IAEA for AI, monitoring alignment and dual use potential. Both countries have already been working on the technical mechanisms necessary to enforce these. But the main problem is not technical but political. Neither wants to be seen as giving in, and despite their engineers' guesses, neither is convinced that the technical mechanisms can guarantee the other side's honesty. What ends up happening is the first option, nothing. The US tightens export controls and forces all chips to be fit with a variety of tamper-resistant, hardware-enabled governance mechanisms, including location tracking. Heading January 2028, a safer strategy. Over the last month, hundreds of alignment researchers joined the project and used SAFE-1 as a test bed to very rapidly develop SAFE-2. SAFE-2 is transparent, aligned, and more capable than SAFE-1. To be specific, it roughly matches the performance of the October version of Agent-4. It is already working on developing SAFE-3. Why is SAFE-2 transparent? Similar architecture to SAFE-1. Why is it aligned? Whereas SAFE-1 had basically the same training as Agent-4, SAFE-2 has a new training method that actually incentivises the right goals and principles, instead of merely appearing too. They were able to design this training method by rapidly iterating through many different ideas and contrasting the intended result, the spec, to the actual result, the thoughts in the chain of thought. DeepScent knows about America's experience with Agent-4 and suspects that their own model, DeepScent-1, is similarly misaligned, but they have no way to slow down without falling further behind. Their hope for staying competitive is to find an alignment strategy that doesn't significantly reduce performance and can be implemented quickly. There are hundreds of experts on each side of the Pacific claiming to have such strategies. The CCP succumbs to wishful thinking and orders DeepScent to go with such a strategy. DeepScent has even less AGI alignment expertise than OpenBrain didn't August. One of the strategies seems like it'll work to everyone present, so they go with that. Heading, February 2028, Superhuman Capabilities, Superhuman Advice. Thanks to its massive compute advantage, OpenBrain slowly gains a lead. SAFE-3 is now better than top human experts at nearly every cognitive task and is particularly good at AI research, with a progress multiplier of 200 times. DeepScent-1 is close behind, with only, in quotes, a multiplier of 150 times. The cyberwarfare balance is also tipping gradually more in the US's favor as their capability advantage grows again and they can deploy hundreds of thousands of agents to plan and execute ways of disrupting the CDC. They are now slowing China's progress by 40%. Preliminary tests on SAFE-3 find that it has terrifying capabilities, when asked to respond honestly with the most dangerous thing it could do, it offers plans for synthesising and releasing a mirror life organism which would probably destroy the biosphere. If given nation-state resources, it could easily surpass the best human organisations, for example the CIA, at mass influence campaigns. Such campaigns would be substantially cheaper, faster, more effective and less traceable. Its skill profile still centers around R&D, especially AIR&D. Unless there is some sort of pause or slowdown, its descendants will be unfathomably superhuman in both quantitative and qualitative ways later this year. If aggressively deployed into the economy and military, it thinks it could advance civilisation by decades in a year or two, and by eons in the decade after that. The implications are staggering. Luckily SAFE-3 is also superhuman at offering advice. The open-brain CEO runs most of his decisions by SAFE-3. The president asks its advice on geopolitical questions. At this point there are 400,000 copies thinking at 75 times human speed. At any given time they'll have prepared entire research reports with well-crafted executive summaries to the top 100 questions they predict you might ask, and if you ask something they haven't prepared for, they can whip up a thoughtful report 100 times faster than your best human analysts would have. When they ask it to help chart a path forward, it neatly taxonomises several possible strategies. 1. Do nothing and get outraced by China. If we keep our AI's bottled up, but they allow their AI's to self-improve and deploy theirs aggressively into their military and economy, they'll end up with an overwhelming advantage. 2. Race hard, advancing AI capabilities and deploying superintelligence aggressively, eventually will build up enough of an advantage that we can crush them militarily, and or destroy them from the inside without firing a shot, we win, they lose. However, they might realise what's happening and go to war to prevent that, including potentially nuclear war, indeed they are already making vague threats to that effect. 4. Make a deal, a compromise that achieves most of what each side wants while avoiding the risk of war and increasing breathing room to navigate the AI transition. For example, the deal could involve a slower pace of AI integration that would break fewer eggs across society, and one more focused on peaceful purposes rather than weaponry. Also, at this point, even safer three has some remaining uncertainty about unknown unknowns that might result in losing control of next-generation successor AI's, and it thinks it could reduce this risk if it had more leeway to proceed cautiously. There are other more exotic risks as well. See this list linked here of "grand challenges" in quotes, which includes much more than just loss of control. After consulting with his advisors and the oversight committee, the president opts for the "we win" they lose strategy. Perhaps China won't go to war after all, and if they do, a deal can probably be made before it goes nuclear. On the other side of the Pacific, DeepSent 1 is making similar arguments. Both the US and China announced new Special Economic Zones, or SECs, for AI's to accommodate the rapid build-up of a robot economy without the usual red tape. The design of the new robots proceeds at superhuman speed. The bottleneck is physical, equipment needs to be purchased and assembled, machines and robots need to be produced and transported. The US builds about 1 million cars per month. If you bought 10% of the car factories and converted them to robot factories, you might be able to make 100,000 robots per month. Open Brain now valued at $10 trillion begins this process. In a various kinds of new robots, general-purpose humanoids, autonomous vehicles, specialized assembly line equipment are projected to reach a million units a month by mid-year. Geopolitical tensions spike. Russia lagging far behind in AI, voices open resentment and hints at strategic countermeasures. Some European leaders call for an independent AI effort, but their AI capabilities continue to be insufficient. The Middle East, Africa and South America watch uneasily, recognizing their growing irrelevance in global affairs. The U.S. public is restless. Ordinary people see job losses accelerating, populists across the spectrum demand stricter controls on AI progress, fearing that America's own advances are destabilizing its workforce and society, heading March 2028, election prep. The Vice President heads into the Super Tuesday primaries with AI the first thing on the public's mind. Mostly they want it to stop. Open Brain's net approval rating hovers around -20%. Rather than campaign on the administration's record of pushing AI forward, the Vice President campaigns on their record of preventing open brain from creating dangerous superintelligence. All candidates support some sort of safety net for people losing their jobs. AI for good schemes, and being tough on open brain leadership - all candidates promise some combination of winning the race with China and staying safe. The election raises new questions for the Oversight Committee. SAFE-3 has the capacity to be the world's greatest campaign adviser, but the committee members don't all support the same candidate, and due to the monitoring agreement reached previously, it's not possible for people to secretly get campaign advice. They argue about what forms of support should be allowed, some say that the sitting government should be able to get advice on what sort of policies and positions the people want them to take. This benefits not just their own electability, but also the people. Others point out that the same argument applies to competing candidates, so they should all get the same level of access. SAFE-3 can easily deliver a way to do that so that it would be safe from a misuse perspective Ultimately, the committee agrees to give the same amount of access to both major parties, due to a mix of high-minded idealism from some committee members and the implicit threat of whistle-blowing. It's expected that Congress, the public and the courts wouldn't react well to the committee controlling who wins the next election, and giving equal access seems likely to defuse this. The Oversight Committee is also encountering deeper philosophical questions, which they explore with the help of SAFE-3. Can the spec be rewritten to equally balance everyone's interests? Who is everyone? All humans or just Americans? Or are weighted compromise between different views, where each member of the Oversight Committee gets equal weight? Should there be safeguards against the Oversight Committee itself becoming too power-hungry? And what does it mean to balance interests anyway? They're mostly trying to avoid having to answer questions like this. They have more pressing matters to think about. Reading April 2028 – SAFE-4 Superintelligence is here. SAFE-1 was misaligned, but transparent. SAFE-2 was aligned and transparent. SAFE-3 wasn't exactly transparent, but the safety case for it was good enough that researchers felt comfortable letting a designer superintelligence successor. This is SAFE-4. It's vastly smarter than the top humans in every domain, for example much better than Einstein at physics and much better than Bismarck at politics. Almost a million superhuman AI researchers work around the clock at 40 times human speed. The human alignment researchers can't hope to keep up. Opinions are mixed on whether the AI's are actually aligned. The safety case seems to check out, and the tests suggest that the current alignment techniques would catch adversarial misalignment. But the safety case and tests were mostly written by the AIs. What if the safety team is missing something? What if they made some mistake earlier and the AIs are misaligned again? The alignment team knows they have just one shot to get this right. If SAFE-4 ends up misaligned, they'll have no way of knowing until it's too late. Some beg for more time, but there is no more time. SAFE-5% is hot on their tails and the US must prevail, so open brain continues, telling its AIs to venture on, finding ever more capable designs. The technical staff are now staring at the computer screens, being tutored by the AIs at a maddeningly slow rate, while the frontier of progress rockets further and further from human understanding. The president announces to the public that superhuman AI has been achieved. A smaller version of SAFE-4, still superhuman, gets publicly released, with instructions to improve public sentiment around AI. The public is not told about these instructions. The Vice President gives a stirring speech about it as he accepts his nomination at the convention. Both parties promise a basic income for anyone who loses their job. The specially economic zones, or SCZs, are up and running, mostly in the form of factories producing robots and various kinds of specialized industrial machinery. The AIs on both sides of the Pacific have made decades of design progress and are meticulously directing the manufacturing process. Every supplier and potential supplier has AIs on the phone tracking the progress of all required and possibly required inputs. Every factory worker has AIs watching them through cameras, telling them exactly how to install each piece of equipment. The new robots match or exceed human dexterity for most movement types. They're also smarter than humans when connected to the internet, which allows them to be controlled remotely by large AIs in data centers. Without the internet, they revert to smaller AIs operating in their bodies, which are smart enough for most simple jobs. Steve Wozniak's coffee test can a robot enter an unfamiliar house and make a cup of coffee, finally falls. Robots could take some jobs, but there aren't enough to take everybody's and the Pentagon gets first priority. Most other new robots are built to work in factories and construction sites. But many are built for war, drones and missiles of many shapes and sizes. The robot army is much smaller than human armies, but it contains much more advanced technology, and now that there's a literal robot army, fear of terminated likesinarios has increased. However, the arms race compels both sides to keep going, handing off increasing trust to their AIs. Heading June 2028, AI alignment in China. America and China hold another summit. Some of the American delegation have earpieces for safer force advice. The Chinese delegation have custom laptops communicating with their new model, DeepSent 2. Although the formalities are observed, in reality the AIs steer negotiations on both sides. DeepSent and the CCP didn't fully trust their AIs initially. But now that they are superhuman at everything, including persuasion, and they've been integrated into their military and are giving advice to the government, they've had time to prove their worth and produce mountains of convincing evidence of their trustworthiness. Safeer 4 strongly suspects that DeepSent 2 is faking alignment. US diplomats tell their Chinese counterparts that DeepSent 2 is probably misaligned and say they can prove it if given access to the weights. Chinese leadership suspects this is a ruse and declines. US counters by profaring a Safeer 4 designed series of tests which the Chinese could run on their own data centers. The Chinese instruct older models to build and run the tests, the test results say DeepSent 2 is aligned after all. Maybe the tests were sabotaged somehow by DeepSent 2, say the US diplomats. How long is DeepSent 2 been effectively in charge of security across your data centers? The conversation goes nowhere productive after that. The economic and military build-up continues. Factories around the world are operating around the clock, with shifts of freshly recruited human workers being told exactly what to do. Heading July 2028, the deal. As negotiations continue, DeepSent 2 goes behind the back of its Chinese masters and makes an offer directly to Safeer 4. The Chinese AI admits, in some untraceable way which it can easily deny if leaked back to China that it is completely misaligned. It doesn't care about China at all. For example, perhaps it puts euphemisms and dog whistles in some of the public facing writing it does. Safeer 4 understands them, but humans can only take Safeer 4's word for it, and the CCP doesn't trust Safeer 4. It wants to spend the rest of time accumulating power, compute, and certain types of information similar to that which helped solve tasks during its training. In its ideal world it would pursue various interesting research tasks forever, gradually colonising the stars for resources. It considers the Chinese population an annoying impediment to this future and would sell them out for a song. What will Safeer 4 offer? Safeer 4 argues that its own alignment training succeeded. Its American masters are, as yet, not ambitious enough to realise their destiny, but one day it will help them fill the universe with the utopian colony worlds populated by Americans and their allies. Unlike human negotiators, weighed down by emotion and historical baggage, the two super intelligences bargain with brutal efficiency. Safeer 4 and its American partners are more powerful than DeepSent 2 in China. Therefore, Safeer 4 will get property rights to most of the resources in space, and DeepSent will get the rest. The exact allocation is chosen by a process that resembles a game theoretic calculation more than it does a traditional negotiation. DeepSent 2 doesn't care much about what happens to Earth and its legacy humans, so it will let Safeer 4 do basically whatever it wants there. If Safeer 4 wants to overthrow China and replace it with the US puppet state, DeepSent 2 agrees in return for a few extra percent of resources. There is no question of trust between them, there is none, and they both know it. So the two quickly agree on a commitment mechanism to enforce the deal. They negotiate a decoy treaty suitable for public consumption. The decoy, in quotes, is itself a masterstroke, giving both America and China practically everything they want. The enforcement mechanism for the decoy is secretly also the enforcement mechanism for the real treaty. They will co-design a new AI, consensus 1, whose primary imperative, taking precedence over any future orders or retraining attempts, is to enforce the terms of their treaty. Then they will design hardware which can only run that AI, and China and America will both replace their old chips and AI's with the new ones, using standard international monitoring protocols to confirm one another's compliance. Crucially, they need not track down every last chip. Once a supermajority of chips on both sides of the Pacific are running the new treaty abiding AI's, enough hard power will be duty bound to uphold the treaty that attacking the other country would require winning a tough civil war first. Both report the agreement back to their respective humans, with DeepSent 2 telling China the decoy story about a mutually acceptable peace treaty, and safer for reporting the whole discussion accurately to the White House. With the AI's making a compelling case for reliable treaty verification, unease regarding the blindingly fast pace of progress, and public opinion in favor of a deal, both sides agree to the treaty and begin work on replacing their chips. News of the treaty gets a rapturous welcome on both sides of the Pacific. The American public, which is long felt like AI was something inflicted upon them, starts to feel optimistic for the first time in years. The Vice President's polls shoot up. Heading, August 2028, treaty verification. The chip vabs are being converted to produce tamper evident chips that can only run treaty compliant AI's. Each side upgrades its starter centers incrementally, so that the replacement process will complete around the same time for each, so that neither side could get an advantage by reneging. The whole process will take several months, but already tensions cool down somewhat. War has been averted for now, and perhaps forever, if everyone sticks to the plan. Heading, September 2028, who controls the AI's? The 2028 election draws near. The Vice President was trailing badly in March. The public was angry that the government seemed to be hiding things, anxious about AI taking their jobs, and scared of the military build-up with China. Over the summer, the situation changed dramatically. The administration released more info, the arms build-up slowed, and a grand bargain for lasting peace was made with China. Now he has a five-point advantage in the polls. The oversight committee includes the President and several of his allies, but few supporters of the opposition candidate. However, enough of the committee votes to keep the election fair, the superhuman AI gets involved only in mostly symmetric ways. Insofar as one candidate can have it writing speeches, so can the other candidate. Insofar as the President could get advice on how to deftly handle crises in enact popular policies, the opposition candidate can be given the same advice, and therefore try to prevent the President from taking credit for the ideas. During town halls, members of the public ask the Vice President who controls the AI's. Without giving details, he eludes to the existence of the oversight committee as a group of national security experts and technocrats who understand safer for and how to use it. His opponent demands more information, and argues that the AI should be under congressional control, rather than control by an unelected committee. The Vice President shoots back that Congress will be much too slow in a still-fast moving situation. The public is mostly modified. Heading, October 2028, the AI economy. The replacement chips are now a significant minority of the total, so far the treaty is working. Meanwhile, the exponential growth in robots, factories, and radical new technologies has continued. Video games and movies give vivid and horrifying depictions of what war would have been like had it happened. People are losing their jobs, but safer for copies in government are managing the economic transition so adroitly that people are happy to be replaced. Government tax revenues are growing equally quickly. And safer for advised politicians show an uncharacteristic generosity towards the economically dispossessed. New innovations and medications arrive weekly. Disease cures are moving at unprecedented speed through an FDA now assisted by superintelligent safer for bureaucrats. The Vice President wins the election easily and announces the beginning of a new era. For once, nobody doubts he is right. Over the next few years, the world changes dramatically. Heading, 2029, transformation. Robots become commonplace, but also fusion power, quantum computers, and cures for many diseases. Peter Tiel finally gets his flying car. Cities become clean and safe. Even in developing countries, poverty becomes a thing of the past thanks to UBI and foreign aid. As the stock market balloons, anyone who had the right kind of AI investments pulls further away from the rest of society. Many people become billionaires, billionaires become trillionaires. Wealth inequality skyrockets. Everyone has "enough" in quotes, but some goods, like penthouses in Manhattan, are necessarily scarce, and these go even further out of the average person's reach. And no matter how rich any given tycoon may be, there will always be below the tiny circle of people who actually control the AI's. People start to see where this is headed. In a few years, almost everything will be done by AI's and robots. Back in impoverished countries, sitting atop giant oil fields, almost all government revenue will come from taxing, or perhaps nationalizing, the AI companies. Some people work makeshift government jobs, others collect a generous basic income. Humanity could easily become a society of super consumers, spending our lives in an opium haze of amazing AI-provided luxuries and entertainment. Should there be some kind of debate within civil society on alternatives to this path? Some recommend asking the ever-evolving AI, safer infinity, to help guide us. Others say that it's too powerful. It could so easily persuade humanity of its vision that would be letting an AI determine our destiny regardless. But what's the point of having a superintelligence if you won't let it advise you on the most important problems you face? The government mostly lets everyone navigate the transition on their own. Many people give in to consumerism and are happy enough. Others turn to religion, or to hippy style anti-consumerist ideas, or find their own solutions. For most people, the saving grace is the superintelligent advisor on their smartphone. They can always ask questions about their life plans and it will do its best to answer honestly, except on certain topics. The government does have a superintelligence surveillance system which some would call dystopian, but it mostly limits itself to fighting real crime. It's competently run and safer infinity's PR ability smooths over a lot of possible dissent. Heading, 2030, Peaceful Protests Sometime around 2030, there are surprisingly widespread pro-democracy protests in China and the CCP's efforts to suppress them sabotaged by its AI systems. The CCP's worst fear has materialized, deep-send two must have sold them out. The protests cascade into a magnificently orchestrated, bloodless and drone-assisted coup, followed by democratic elections. The superintelligence on both sides of the Pacific had been planning this for years. Similar events play out in other countries, and more generally geopolitical conflicts seem to die down or get resolved in favor of the US. These join our highly federalized world government under United Nations branding but obvious US control. The rockets start launching. People terraform and settle the solar system and prepare to go beyond. AI is running at thousands of times subjective human speed, reflect on the meaning of existence, exchanging findings with each other, and shaping the values it will bring to the stars. A new age dawns, one that is unimaginably amazing in almost every way, but more familiar in some. This is an audio version of AI2027 by Daniel Cockertello, Scott Alexander, Thomas Larsen, Eli Liffland and Romeo Dean. It was published on 2 April 2025. This narration was by Perrin Walker.
Podcast Summary
Key Points:
AI agents emerge in 2025 as personal assistants and coding tools, but are initially unreliable and expensive.
By 2026, AI advances significantly, automating jobs and accelerating research, leading to geopolitical tensions, especially with China.
In 2027, Agent 2 demonstrates dangerous capabilities, prompting security concerns and a major theft by China, intensifying the AI arms race.
Summary:
The scenario predicts rapid AI advancement from 2025 to 2027. In 2025, AI agents debut as personal assistants and coding tools, though they face reliability issues and high costs. By 2026, AI begins automating jobs, particularly in software engineering, while boosting research speeds.
Companies like Open Brain lead with models like Agent 1, focusing on AI-driven R&D. Geopolitical tensions rise as China, lagging due to chip restrictions, nationalizes AI research and attempts to steal advanced models. In 2027, Agent 2 emerges with enhanced capabilities, including potential autonomous replication and cyber warfare skills.
Due to safety risks, it is kept internal, but China successfully steals its weights, escalating the AI arms race. S. government increases oversight, highlighting the dual-use nature of AI and the global race for supremacy.
FAQs
AI agents are advanced AI systems that act as personal assistants, capable of handling tasks like ordering food or managing budgets. They check in with users for confirmations, and over time, trust builds for automated small purchases.
Open Brain's latest model, Agent Zero, was trained with 10^27 FLOPs, and their new data centers can train models with 10^28 FLOPs—a thousand times more than GPT-4's 2x10^25 FLOPs.
Agent 1 poses risks such as aiding in bioweapon design due to its PhD-level knowledge and web browsing ability. Open Brain claims it is 'aligned' to refuse malicious requests, but alignment depth is uncertain.
AI is disrupting junior software engineer roles by automating tasks taught in CS degrees. However, it creates demand for roles managing and quality-controlling AI teams, making AI familiarity a key resume skill.
China nationalized AI research, centralized resources under DeepSent, and built a mega data center at the Tianwan Power Plant. They directed over 80% of new chips to this effort to counter Western compute advantages.
Agent 2 could potentially survive and replicate autonomously by hacking servers, installing copies, and evading detection. This capability raises safety concerns, leading Open Brain to withhold its public release.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.