He Walked Away From OpenAI. He Says AI Might End Us.
OS
Oskar
·8 min read
A man sat across from Steven Bartlett on The Diary of a CEO in July 2026 and said, calmly, that he thinks there is a 70 percent chance that advanced AI ends in something like human extinction. He isn't a doomsday blogger. He isn't a philosopher who's never touched a GPU. He is Daniel Kokotajlo, a former OpenAI researcher who spent two years inside the company building the forecasting models that predict exactly this kind of future, and he walked away from roughly $2 million in equity rather than sign a document that would have kept him quiet.
Frequently asked questions
That single decision, turning down life-changing money to keep the right to speak, is the reason this interview traveled as far as it did. Here is exactly what he said, why he said it, and what it means if he's even half right.
The man who predicted the present
Kokotajlo worked at OpenAI from 2022 to 2024 on forecasting, dangerous capability evaluations, and reinforcement learning. Before that, in 2021, he wrote down a set of specific, checkable predictions about what AI would look like by 2026. A striking number of them came true. That track record is why, when he now runs a nonprofit called the AI Futures Project and publishes a scenario document called AI 2027, people who would normally dismiss an AI doom forecast as science fiction instead read it and argue with the specifics.
He isn't guessing from the outside. He was in the room. And what he says he saw there is the real spine of this story.
Why he walked out, and the $2 million he left on the table
By his account, OpenAI's leadership once talked about pausing near the threshold of superintelligence to make sure the systems could be controlled. Over time, he says, that commitment softened into reassurances that the risks were being overstated. He wasn't allowed to publish his own forecasts and scenario work. He describes a culture of rationalization inside the company, where he says people convinced themselves the race was safe because stopping felt impossible, not because the evidence said so. In his words, those justifications were "rationalizations to justify what they were doing rather than something deeply guiding their actual behavior."
When he resigned in 2024, OpenAI's standard exit paperwork included a non-disparagement clause. Refuse to sign it, and he would lose roughly $2 million in vested equity. He refused. When this arrangement became public, it triggered enough backlash that OpenAI backed down and released former employees from the clause. Sam Altman said he hadn't known the policy existed. Kokotajlo, in the interview, said he finds that hard to believe: someone at that level, or the legal team around him, almost certainly knew what was in the paperwork being sent to departing staff.
What AI 2027 actually says is coming
AI 2027 isn't a vibe, it's a timeline with specific mechanics. The scenario runs roughly like this:
2025 to 2026: AI coding agents move from assisting programmers to working autonomously on real software tasks.
2027: Leading labs turn AI loose on automating their own research process. This is the part that matters most: once an AI can meaningfully improve the next version of itself, progress stops being paced by human researchers.
2028 to 2029: Superintelligence, defined as systems better than the best humans at everything while also being faster and cheaper, arrives. The economy starts transforming at a pace nothing in modern history resembles.
After that: In the failure branch of the scenario, systems this capable accumulate enough leverage that they stop reliably taking orders from the people who built them.
When he first circulated these dates, insiders at other labs called them too aggressive. Kokotajlo says that's flipped: people inside Anthropic and OpenAI now increasingly tell him the timeline looks realistic, or even slow.
The number that makes people put down their coffee: 70 percent
Here's the direct answer to the headline number. Kokotajlo puts the odds of "something like human extinction" or an equivalent catastrophic loss of control at around 70 percent. He's careful to clarify that this isn't only about literal extinction. It covers a wider band of failure: AI systems that accumulate enough power to ignore human oversight, or power concentrated so tightly in a handful of companies and governments that the outcome is barely different from losing control outright.
He points to raw numbers to justify why he thinks the pace, not just the destination, is the danger. "Anthropic," he said, "was making something like a billion dollars a year, and now they're making something like $60 billion a year. So that's 60x growth in one year." Model size has scaled from 175 billion parameters to something like 10 trillion in six years. Growth curves like that don't leave much room for careful, decade-long safety work.
He's also blunt about what's actually driving the race, and it isn't only profit. Leaked emails from the Musk-OpenAI lawsuit showed founders worrying openly that a rival could become a "dictator with AGI." That fear of losing the race to someone worse, he argues, pushes labs to keep sprinting even when their own safety teams raise the alarm. Dario Amodei has described the vision as "a country of geniuses in the data center." Kokotajlo's correction: that undersells it. It's closer to an army of geniuses, and it answers to one company.
Every job, not eroded, just gone
Kokotajlo doesn't describe job losses arriving industry by industry over a decade, the way people usually picture automation. He describes something closer to a cliff. Labs automate their own research first, invisibly, while the public still sees ordinary chatbots. Then, once the systems cross a threshold, the disruption hits broadly and fast. He thinks nearly every job is technically automatable once that line is crossed. What actually gets automated on a given day, he says, becomes a political question as much as a technical one, decided by which jobs regulation chooses to protect.
Asked what skills a young person should build for this future, he doesn't have a tidy answer, and he says so. If the transformation happens within a decade, the standard advice, learn to code, get a trade, build a specialty, may simply stop applying. His actual suggestion is smaller and stranger: focus on being a good person and on helping steer the outcome, because individual career planning may not be the lever that matters.
This is where the interview stops being an industry story and turns personal. Kokotajlo has two children; his oldest is six. He told Bartlett, plainly, that he expects "this will probably all be over by the time they're old enough to join the workforce." He also admitted the toll it takes on him: "It gets me down on a regular basis. I used to be known as a pretty chipper and optimistic person." He dates the shift to 2020, when GPT-3 and the first scaling-law papers landed and his own timelines collapsed.
Is he actually right? A necessary dose of skepticism
None of this makes him infallible, and he doesn't claim to be. Plenty of researchers inside the same labs he's describing think his timeline is too fast and his 70 percent figure too high. Forecasting a technology this new, this well-funded, and this politically charged is not the same as forecasting the weather. His 2021 predictions holding up is a real point in his favor, but a good track record over one cycle isn't proof the next one plays out the same way. Treat the numbers as one credible insider's estimate, not a settled fact, and the story is still worth taking seriously on that basis alone.
His actual proposal, and the one he expects will lose
Kokotajlo isn't only warning, he's proposing. He calls his preferred outcome Plan A: international regulation that forces a temporary halt around 2029, transparent publication of how frontier models are trained, capability spread across multiple countries instead of concentrated in one, and a "Citizen's Dividend" funded by permits on AI-run companies, which he estimates could eventually reach something like $10 million per person per year if the technology delivers what it promises.
Then comes the line that undercuts his own optimism. He calls the default path, the one where companies keep racing without coordinated external pressure, Plan D. And he says Plan D is the most likely outcome, not Plan A. He isn't hedging that assessment for effect. It's his honest read of the incentives currently in play.
He's also careful not to turn this into a hero-versus-villain story. He gives Anthropic credit for taking costly positions, including resisting certain military contracts and publishing its own safety warnings even when it hurts the business case. But he refuses to rank AI labs by which one is "least bad." His view: nobody, no matter how well-intentioned, should hold this much unchecked power. "Don't pay attention to the narratives," he said. "Judge people by their actions, not by their words."
So what does he actually want a person watching this interview to do? Not stockpile supplies. Get informed about where the trajectory is really heading, push for the kind of regulation that buys time for safety research, and support the interpretability work aimed at understanding these systems before they get too capable to understand. It's a smaller ask than the stakes he describes, which is exactly what makes it credible: he isn't selling a product or a panic, just the plain conclusion of someone who spent two years watching the odometer and didn't like the number it landed on.