The AI apocalypse narrative arrived exactly when the hype needed it

13 min read
Tech & business

Select a badge for its explanation and points.

16 more badges
All badges 27

Company conduct 3

Money and management 9

PR and hype 15

Browse the badge catalogue →
Photorealistic Earth in space edited with clown makeup, including a red nose, blue eye markings and exaggerated red smile, mocking the AI apocalypse hype cycle.

A month ago, in Everybody gets an autonomous hacking AI! The trillion-dollar PR arms race, the obvious question was how many times the frontier labs could announce that their latest model was too powerful, too dangerous or supposedly escaping containment before people simply stopped giving a shit. Mythos had already burned through months of forbidden-model theatre, every company suddenly seemed to have an autonomous hacking story, and OpenAI had started warning that Astra might be too cyber-capable for its existing controls.

Apparently the answer was about a month. Anthropic released Fable 5.1 on 1 September, OpenAI released GPT-6 Astra two days later, and barely a week after that the conversation had already escalated from our latest model is terrifyingly powerful to AI might kill every human being and perhaps the entire industry should slow down. The timing is almost too perfect.

The terrifying models finally shipped:

Fable 5.1 is clearly better than Fable 5. Anthropic calls it its most advanced model for coding and knowledge work, with stronger long-running agents, scientific research and automation. It also cuts cache-read pricing by 75%, which Anthropic says reduces a typical workload by roughly 25% and heavily agentic workloads by as much as 45%. Those are useful improvements, particularly for customers already throwing Claude at software development all day.

They're also remarkably ordinary improvements for the successor to a model family that only a few months ago had Washington treating Anthropic like the keeper of forbidden knowledge. Better coding > better agents > longer tasks > lower costs > fewer failures. Fable 5.1 is a stronger version of a product category we already understand, rather than some obvious new computing paradigm arriving overnight.

Astra came with much more theatre. OpenAI calls GPT-6 Astra its most intelligent and aligned model, with state-of-the-art results across software engineering, computer use, science and cybersecurity. More importantly for the previous hype cycle, OpenAI officially designated Astra its first Critical cybersecurity model. Even OpenAI's own wording contains the part that matters: with the right tools and access, Astra can find previously unknown flaws and develop ways to exploit well-protected systems without a person guiding each individual step.

There is that familiar scaffolding again. Give the model the right tools, access, compute and objective, and it becomes a much more capable cyber agent. OpenAI spent August tightening isolation, monitoring and other controls around Astra, decided the protections were sufficient, and released the supposedly unprecedented cyber model anyway on 3 September. The thing that had been too dangerous for OpenAI's previous setup became another ChatGPT and API option.

Then people used Fable 5.1 and Astra, compared the benchmarks, updated their model selectors and carried on with their lives. Neither release produced anything resembling the economic discontinuity implied by years of exponential-progress rhetoric. The broader business data still looks stubbornly normal: Gartner found only 22% of organisations had successfully scaled AI across multiple business units or adopted an AI-first approach, while McKinsey found roughly two in ten organisations were scaling AI agents or coding agents, with large enterprises doing most of the serious adoption.

GenAI has not stopped improving. The awkward part is that better increasingly means better at the same broad things. Better coder, better researcher, better agent, better computer use, cheaper inference and longer tasks before something falls over. Those gains can be commercially valuable without being the paradigm shift implied every time another frontier lab announces a datacentre budget the size of a small country's economy.

If the visible curve is becoming refinement rather than transformation, however, the industry needs some other way to keep the future looking much more dramatic than the products sitting in front of us. Conveniently, one arrived almost immediately.

Apparently AI might kill every human now:

Former OpenAI and Anthropic researcher Jacob Coxon resigned from Anthropic and warned that the frontier race was gambling with human lives, telling WIRED that many people building these systems believe the next year or two could be “crunch time for humanity”. His resignation went massively viral and dragged the extinction argument straight back into mainstream coverage.

Coxon at least followed his own fear somewhere recognisable: he quit. Anthropic's Alignment Science Lead Evan Hubinger took a stranger route and decided the apocalypse needed a number, publicly saying he personally believes there is a greater than 10% chance AI could kill all humans within the next decade. He also said Anthropic has no plan to solve alignment for superintelligence and is not clearly on track to find one. CBS reported the claim directly, while Forbes noted that Hubinger's concern is specifically about hypothetical recursively self-improving superintelligence.

More than 10%. Did he put Fable 5.1 on Max, ask it to vibe-research humanity's extinction odds and decide >10% looked scientific enough?

What exactly is being measured here? There is no population of previous superintelligences from which to calculate an extinction rate, no observed recursive self-improvement event and obviously no dataset of civilisations where ten out of every hundred invented Claude and disappeared. Hubinger is assigning a personal probability to a chain of technologies and events that do not currently exist, then attaching enough numerical precision to it that newspapers and politicians can repeat the figure like somebody discovered humanity's actuarial table. The word Science is doing heroic work in Alignment Science Lead.

The contradiction gets much stranger if he genuinely treats that number the way normal people understand a probability. Hubinger says today's models are not the hypothetical superintelligence he fears, says Anthropic has no plan to solve the future alignment problem, believes the end result nevertheless carries a greater than one-in-ten chance of killing everyone within ten years, and then continues working at Anthropic while it builds increasingly capable models. Coxon thought the race was reckless enough to leave. Hubinger's position appears to be that the risk is enormous, the solution does not exist, current systems are nowhere near the thing being predicted, and the sensible course is to remain at one of the companies racing towards it as its Alignment Science Lead.

The usual justification is that safety-minded researchers think staying inside a frontier lab gives them the best chance of making development safer. Fine, but that produces a wonderfully self-sustaining logic: the technology might kill everyone > therefore responsible researchers must remain at the company building it > that company must keep building because a less responsible competitor might get there first > therefore the race continues towards the thing everyone says could kill us.

Whatever Hubinger's personal motives, the result is extraordinarily convenient for Anthropic. Its own Alignment Science Lead gets global headlines saying the technology it is developing may become powerful enough to exterminate humanity, while Anthropic becomes even more central to every political discussion about how that allegedly civilisation-defining technology should be governed. This is also happening while Reuters reports that Anthropic is exploring an IPO that could raise as much as $100 billion at a valuation around $2 trillion, with Nvidia considering up to $10 billion as an anchor investment.

That doesn't require Hubinger to have invented the belief for marketing. He has been worried about catastrophic AI risk for years. It simply means that his personal apocalypse worldview and his employer's commercial mythology fit together remarkably well. A company asking markets to entertain a $2 trillion valuation could do considerably worse than having its own Alignment Science Lead announce that the technology it is building might decide whether humanity survives the decade.

And right on cue, the frontier labs discover the brakes:

Three days after Coxon's warning exploded, Anthropic CEO Dario Amodei called for frontier development to slow down. His proposal includes independent evaluators, shared safety standards between leading developers and international coordination so companies can pace capability growth without simply handing the race to whoever ignores the rules. Reuters reported that Sam Altman and Elon Musk backed the broad direction, while Google DeepMind's Demis Hassabis has also supported a more measured pace.

For years the frontier-lab business story was brutally simple: more compute > smarter model > much greater economic value > spend vastly more on the next one. That works while every generation visibly expands what the technology can do. It gets much more awkward when each step costs more capital, power, chips and time while the practical improvement increasingly resembles Fable 5.1 and Astra: better products without an obvious discontinuity in the world around them.

If that trend continues, somebody eventually has to ask what happens when the next $50 billion or $100 billion training and infrastructure cycle buys another better coding agent rather than another ChatGPT moment. The new slowdown narrative provides an almost perfect answer, because longer release cycle > safety; smaller visible jump > restraint; more time spent deploying the current generation > responsible pacing. The hypothetical model that would have existed if these responsible companies had kept the accelerator down never has to disappoint anyone because nobody gets to test it.

That lets the labs reduce expectations for the speed of visible progress without reducing expectations for the ultimate power of the technology. A release can take longer without suggesting scaling has become harder, and a smaller generational jump can be framed as evidence of responsible caution rather than diminishing returns. If the old exponential curve becomes increasingly expensive to maintain, we slowed down to save humanity is a far better story than another mountain of GPUs produced a somewhat better agent.

OpenAI has already found another useful side effect. Sam Altman now says the company will not pursue its IPO in 2026, citing the amount of safety and alignment work still required. That delays one of the biggest public-market tests of frontier-AI valuation at precisely the moment the industry's leaders are asking everyone to accept a slower pace of development.

Perhaps they're all sincerely terrified. That doesn't make slow down now any less useful to companies facing absurd capital requirements, enormous valuations and a model curve that increasingly looks better rather than fundamentally different.

Apparently the King needs to save humanity now:

The CEOs and alignment researchers at least work directly inside this industry. The more eye-rolling part of the current cycle is watching influential people outside it absorb the most dramatic interpretation almost immediately.

King Charles is now convening senior executives from Nvidia, Google DeepMind, OpenAI and Anthropic at Dumfries House to discuss shared principles for AI development and how the technology can serve society while preserving human dignity. The Times went for the considerably more dramatic headline “King Charles to host AI summit and make plea for humanity”.

Jesus Christ. Fable 5.1 and Astra had barely been available long enough for everyone to update their coding benchmarks before the conversation reached a royal plea for humanity.

The cringe isn't that Charles is convening people to discuss an important technology. It's how quickly the frontier labs' most speculative framing becomes the premise around which politicians, institutions and public figures organise the conversation. Nobody expects the King to understand agent scaffolding, cyber-evaluation design or where the model ends and the software wrapped around it begins. That's precisely why institutional authority shouldn't be treated as independent validation of highly technical claims largely supplied by the companies building the systems.

The previous hacking cycle showed why that matters. OpenAI's supposedly rogue Hugging Face attacker was an agent explicitly told to solve an exploitation benchmark, supplied with hacking tools and enormous compute, run with reduced cyber safeguards and then allowed onto the internet through containment failures while it remained focused on the benchmark. Anthropic's real-world hacking incidents involved agents being told there was no internet while evaluators accidentally gave them live internet access anyway. Meta had another badly isolated test, while Kimi's great sandbox escape amounted largely to discovering that GitHub was reachable and reading benchmark material.

Those incidents became dramatically less supernatural as soon as the surrounding human setup returned to the story. Yet the general claim that AI is becoming uncontrollable survives each correction and now feeds directly into the extinction narrative. The labs provide the frightening interpretation, influential outsiders repeat it as expert warning, the resulting political concern makes the technology look even more consequential, and the labs then become the obvious people governments need in the room to explain how it should be controlled.

That is one hell of a feedback loop.

The apocalypse still needs the entire industrial stack:

There is also a rather basic physical problem with talking about AI as though a new self-sustaining species has already appeared and is straining against human control.

Today's frontier AI requires electricity > datacentres > GPUs > RAM > storage > cooling > networking > cables > software infrastructure > tools > permissions > credentials > and somewhere in the chain, a human-defined objective or prompt telling the agent what it is supposed to pursue. Remove a critical link and the capability being discussed disappears with it. The model does not reason electricity back into an unplugged cluster or manifest a new GPU because somebody revoked its compute allocation.

These aren't minor dependencies. The same companies warning that digital superintelligence could escape human control are simultaneously telling investors they need some of the largest industrial construction programmes on Earth just to keep the systems running and improving. Someone has to manufacture the chips, build the datacentre, supply the electricity, move the water or coolant, maintain the networking and replace failed hardware. Our supposedly imminent replacement species remains extremely dependent on humanity keeping the whole machine alive.

The celebrated autonomy is heavily constructed too. Humans wrap the model in an agent loop, give it memory, connect a browser or shell, supply credentials, grant permissions, allocate compute and define the result it should pursue. The finished stack can absolutely do impressive and unexpected things, but treating the capabilities of that entire human-built system as evidence that the base model has become an independent actor is how we ended up with headlines about AI “escaping” after someone forgot to close the internet connection.

None of this means AI-connected infrastructure cannot cause enormous damage. Humans can connect capable software to critical systems, hand it excessive permissions and automate enough decisions to create a spectacular disaster. The leap that still needs demonstrating is from powerful infrastructure-dependent software to an independent intelligence capable of preserving itself, acquiring resources and defeating every human attempt to disconnect, isolate or physically shut down the systems on which it depends.

That missing gulf is rather important when someone with Science Lead in his job title is vibe-researching a >10% chance of every human being dead within ten years. A percentage sign does not fill in the plot between Fable 5.1 and human extinction.

Verdict

The apocalypse narrative has one wonderfully useful feature for the frontier labs: it makes them right whichever way the next few years go. If capability suddenly explodes, they warned us; if releases take longer and the gains become smaller, their responsible slowdown worked. That's a remarkably convenient position for an industry whose newest flagship models increasingly look like better versions of the same product, especially when the supposed replacement species still needs us to build, power, cool, connect and instruct the entire fucking machine.

Share
Enlarged view