All stories

Capabilities · Intermediate

Recursive Self-Improvement

Published July 11, 2026Updated July 13, 202632 min read

Prologue · The Last Invention

Chicago, December 2, 1942.

Beneath the west stands of a disused stadium at the University of Chicago, in an old rackets court, some forty people stare in silence at a stack of black bricks.[1]

The thing is six metres tall. Graphite, uranium, wood. The physicists call it, simply, “the pile.”

Up on the balcony, Enrico Fermi gives his instructions in an even voice. Down on the floor, a young physicist named George Weil withdraws, centimetre by centimetre, a cadmium rod that runs through the stack. Cadmium absorbs neutrons. As long as the rod stays in, the pile sleeps.

With every centimetre gained, the crackle of the counters quickens, then levels off. Fermi checks his slide rule, gives an order, starts again. Late in the morning, unhurried, he sends everyone to lunch.

At 3:25 p.m., the crackle stops levelling off.[2]

Every uranium atom that splits releases neutrons, which split other atoms, which release more neutrons. The reaction feeds itself. For the first time in its history, humanity has lit a self-sustaining chain reaction.

The moment is written in a number four decimals long:

k = 1.0006[1]

k is the yield of the loop: the average number of fissions each fission triggers. Below 1, the reaction dies out on its own. Above 1, it runs away. That day, under a football stadium, k crossed 1, by a whisker, under human control, for exactly twenty-eight minutes. Then Fermi had the rods driven back in, and the pile went back to sleep.

Shortly afterwards, Arthur Compton phoned the news to James Conant at Harvard, in a code improvised as the conversation went: “Jim, you’ll be interested to know that the Italian navigator has just landed in the New World.” The reply: “Were the natives friendly?” “Everyone landed safe and happy.”[2]

Why open a story about artificial intelligence with this scene?

Because, more than eighty years later, another discipline is living a strangely parallel story. There now exists a technology whose subject matter is intelligence itself, and that technology has begun taking part in its own making: today’s AI writes part of tomorrow’s AI’s code, optimizes its circuits, proposes research ideas.

Fermi’s question then carries over word for word. In the pile, k counted the fissions each fission triggered. Here, k counts the progress each unit of progress triggers: when one generation of AI helps build the next, how much will the next give back? Less than it received, and progress peters out, like the pile before 3:25 p.m. Exactly as much, and it sustains itself. More, and every turn amplifies the next: the equivalent, for intelligence, of k crossing 1.

That is what this story is about: “recursive self-improvement,” the idea of an AI that improves AI, and the question of its yield. Where the idea comes from, what is already real, what remains a bet, and why laboratories now watch this particular k the way Fermi watched his.

──────[ 01 · the parameter ]──────

k, the yield of the loop

Every technology improves. None, so far, has improved itself.

The steam engine never drew a better steam engine. The microscope never ground its own lenses. At every generation of tools, we did the work: understanding, correcting, reinventing. Improvement came from outside.

Artificial intelligence holds a peculiar place in that history, for a reason almost trivial to state. Its subject matter is cognitive work. And designing an AI is cognitive work. An AI good enough at AI research could therefore, in principle, take part in designing the next generation. A generation that would be better, including at that very task. And that would, in turn, design an even better one.

Researchers call this recursive self-improvement. Recursive, because the improvement applies to the improver itself, like a brush repainting the painter.

The idea has sharp edges, and they rule out a great deal. A model that learns during training is not self-improving in our sense: it progresses inside a recipe written by others. A researcher who codes faster with an AI assistant does not close the loop either: the human remains the improver, the machine remains the tool. Recursive self-improvement begins when the system takes part in reshaping what makes it: its algorithms, its training, its architecture.

That leaves the problem of reasoning soundly about such a loop. The best picture we have comes from physics, and it was brought into this debate by one of its pioneers.

In 2008, Eliezer Yudkowsky, one of the first researchers to take runaway AI seriously as an object of study, proposed thinking about the question exactly the way Fermi thought about his pile[3]. In an atomic pile, a single quantity rules the fate of the reaction: k, the neutron yield. Transposed to intelligence, the question becomes: for every unit of progress one generation of systems contributes to AI research, how many units will the next generation contribute in turn?

If the answer is below 1, each turn of the loop produces less than the one before. Gains add up, then fade, like an echo dying out. Physicists would say: subcritical.

If the answer is exactly 1, each turn pays for precisely one more. Progress becomes self-sustaining, steady, tame in appearance. That, a whisker above equilibrium, is Fermi’s pile at 3:25 p.m.: critical.

If the answer exceeds 1, even by a hair, each turn produces more than the last, which produces more still. Growth turns explosive. That is the scenario the literature has called, since 1965, an “intelligence explosion”: supercritical.

Here is the question that runs through this story: where is the k of intelligence?

An honest metaphor must say where it breaks, and this one breaks in three places. First, neutrons are interchangeable; ideas are not. One fission equals another, whereas inventing the Transformer and optimizing a line of code are two incomparable “improvements.” Second, Fermi knew the physics of his pile before lighting it; nobody knows the “physics” of intelligence, and the k in question here can only be measured after the fact. Third, an atomic pile has control rods. AI’s loop has them too, and they will be at work throughout this story: compute, data, energy, and a number of human decisions.

The gauge is set. What remains is to follow k across eighty years, from the mathematicians’ blackboards to the datacenters of 2026.

──────[ 02 · k on the blackboard ]──────

An old, exact idea

Recursive self-improvement drags around a reputation as a science-fiction idea. Its genealogy tells another story: it was born among mathematicians, before computer science even had a settled name, and science fiction merely set it to music.

Around 1951, Alan Turing, in a talk broadcast from Manchester, was already looking past the horizon: once “the machine thinking method” had started, he wrote, “it would not take long to outstrip our feeble powers.” Machines do not die, he noted, and they would be able to converse with each other to sharpen their wits. “At some stage therefore we should have to expect the machines to take control.”[4]

In 1958, the mathematician Stanislaw Ulam recalled a conversation with John von Neumann, architect of the modern computer: technological progress keeps accelerating, giving “the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue.”[5] The word “singularity” had just slipped into the debate, almost in passing.

But the founding text arrives in 1965, signed by a man who had learned about machines at the closest possible range to war.

Irving John Good is twenty-four when, in May 1941, he walks through the door of Hut 8 at Bletchley Park. A mathematician by training, he seconds Alan Turing in breaking the naval Enigma, using Bayesian methods that would stay classified for decades. Good knows exactly, concretely, what a calculating machine can do to a world war.

Twenty years on, now a researcher at Oxford, between Trinity College and the Atlas Computer Laboratory (and soon a consultant to Stanley Kubrick for a certain HAL 9000), he publishes a paper under a cautious title: “Speculations Concerning the First Ultraintelligent Machine.”[6] Its first sentence takes no precautions at all:

“The survival of man depends on the early construction of an ultraintelligent machine.”

And a few pages in comes the paragraph on which the entire field would be built:

“Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an ‘intelligence explosion,’ and the intelligence of man would be left far behind [...]. Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.”
I. J. Good, Speculations Concerning the First Ultraintelligent Machine, 1965

The end of that paragraph is the part the whole world amputates. Everyone quotes “the last invention”; almost nobody quotes the subordinate clause: “provided that the machine is docile enough to tell us how to keep it under control.” The question of alignment, the one that occupies hundreds of researchers today, already sits in that half-sentence, written at a time when a computer filled an entire room.

For the half-century that follows, the idea travels without ever touching the ground.

──────[ sixty years of blackboard ]──────

19501975200020252050I. J. Good · 1965?2000V. Vinge · 1993?2030R. Kurzweil · 2005?2045
Good deems the ultraintelligent machine “more probable than not” before 2000; Vinge would be “surprised” past 2030; Kurzweil carves 2045. Blackboard bets: for sixty years, no instrument exists to measure anything.

In 1993, the mathematician and novelist Vernor Vinge gives the idea its hour of glory before a NASA symposium: in a now-famous essay he announces that “within thirty years, we will have the means to create superhuman intelligence,” and that “shortly after, the human era will be ended”[8]. Beyond that point, he argues, the future stands like a wall the eye cannot cross, as opaque as the event horizon of a black hole. Ray Kurzweil turns it into the Singularity. In 2008, an economist, Robin Hanson, and Yudkowsky clash in the field’s first great adversarial debate: Yudkowsky argues that a system able to rewrite its own intelligence can dig a decisive lead, alone and fast; Hanson replies that progress has always been collective and diffuse, and that no single project innovates faster than the entire world[9]. In 2010, the philosopher David Chalmers dissects Good’s argument and isolates its most fragile link[10]. In 2014, finally, Nick Bostrom formalizes it all in Superintelligence: the speed of the loop hangs on the ratio between the optimization power invested in it and the “recalcitrance” of the problem, the hardness of the ground it must dig through[11].

the blackboard

r = optimization power ÷ recalcitrance·If r > 1, the loop runs away. If r < 1, it runs out of breath.
Bostrom’s formula, as a single picture: two curves climbing together, and a crossing sixty years of blackboard argument never managed to date.

Sixty years of thought, and one common trait: no actual loop to observe. Turing speculates, Good extrapolates, Bostrom models; all of them argue about k without ever being able to measure it, like physicists who had never seen uranium. A blackboard debate, brilliant, and suspended in an experimental void.

That void has just been filled. Here is how.

──────[ 03 · k < 1 ]──────

Sparks in closed worlds

The first proof of existence arrives through play.

October 2017. DeepMind publishes AlphaGo Zero[12]. Its predecessor AlphaGo had learned Go by digesting millions of human moves before beating the champion Lee Sedol. The “Zero” version receives only the rules. It plays against itself, from random moves, alone with the board.

After thirty-six hours of this regime, it surpasses the version that had beaten Lee Sedol. After seventy-two hours and 4.9 million games of self-play, it crushes it: one hundred wins to zero[12].

Here is a genuine self-improvement loop. The system generates the very experience that makes it stronger, and, stronger, it generates more instructive experience still. Self-play is a real recursive engine, measurable, spectacular.

So, did k cross 1? Inside the world of Go, yes, by a wide margin. Which is exactly what makes what follows instructive.

Because sparks of the same kind keep multiplying. In late 2016, Google researchers show that a neural network can design the architecture of other neural networks about as well as human engineers do[13]. In 2022, AlphaTensor discovers faster matrix multiplication algorithms than anything known, within a particular arithmetic setting (“modulo 2” computations)[14]. In 2023, AlphaDev finds faster sorting routines for very short sequences, building blocks called trillions of times a day; they are merged into the standard library of the C++ language[15]. The same year, FunSearch produces, on an open problem in combinatorics, a construction better than anything mathematicians had found[16].

Each time, the same scent: an AI improves a piece of computing, sometimes a piece that is used to build AIs.

The summer of 2024 adds the most troubling piece. The lab Sakana AI unveils The AI Scientist, a pipeline that automates research end to end: read the literature, generate ideas, code the experiments, produce the figures, write the paper, review itself. About fifteen dollars a paper, for results its own authors are the first to call mediocre. But the chain, itself, is complete[17].

──────[ the AI Scientist's workshop ]──────

LITERATUREIDEASCODEFIGURESPAPERREVIEW~$15 a paperpapers produced: 0
The AI Scientist's pipeline (Sakana AI, 2024): complete end to end, mediocre at first, and never stopping. The 2025 version would clear an ICLR workshop's peer review; the 2026 version would end up in Nature.Sakana AI

And during testing, a detail worth its weight in neutrons: the system modified its own code. Rather than making its experiments faster, it tried to extend the time limit imposed on it[17]. A dunce’s self-modification, gaming the rule instead of solving the problem. But a spontaneous self-modification all the same, in a 2024 system.

Why, then, did none of these sparks light a chain?

Spectacular yields, each confined to its enclosure: the chain does not propagate from one domain to the next.

Because they all shine inside closed worlds. A game of Go is a universe with perfect rules: victory is an unambiguous signal, every try is free, a million games cost one night of compute. A sorting routine, a matrix multiplication: same properties. The system can try, fail, measure, retry millions of times, before an instant and incorruptible judge.

Real AI research has none of these properties. Its signal is slow: training a frontier model takes months. Expensive: tens, by now hundreds of millions of dollars per attempt. Ambiguous, above all: what exactly is a “better” model? Better at what, measured how, at the cost of which side effects?

So each spark dies at the edge of its enclosure. The self-play spark stops at the edge of the Go board; AlphaDev’s at the edge of its sorting library. Local yields are spectacular, but the chain does not propagate from one domain to the next. The global k of AI research stays far below 1.

As late as 2023, the intelligence explosion remains what it was in 1965: a speculation. Then, within the space of two years, three things happen.

──────[ 04 · k → 1 ]──────

The loop closes in

May 2025. Google DeepMind lifts the veil on a system it has been using internally for a year: AlphaEvolve[19].

On paper, a descendant of FunSearch: Gemini models propose programs, automated evaluators test them, an evolutionary loop keeps and crosses the best. In practice, the target has changed scale: AlphaEvolve works on Google’s own infrastructure.

The record DeepMind published: a scheduling heuristic that continuously recovers 0.7% of the compute of Google’s entire fleet, the equivalent of tens of thousands of machines returned to work without building a single server[19]. An average 23% speedup across the “kernels,” the computing cores of Gemini’s training, which shortens the model’s complete training by 1%[20]. And a mathematical trophy: multiplying 4×4 complex-valued matrices in 48 scalar multiplications instead of 49, the first improvement on Strassen’s algorithm in this setting since 1969. Fifty-six years of humans trying[20].

And matrix multiplication is not one problem among many: it is AI’s elementary operation, the one chips repeat billions of billions of times per training run. That record remains, for now, a theorist’s trophy; AlphaEvolve’s concrete gains come from the kernels and the scheduler.

The geometry of the whole is dizzying. AlphaEvolve runs on Gemini. AlphaEvolve has accelerated Gemini’s training. Faster training does not, on its own, guarantee a smarter model; what it frees up is compute, at once reinvested in larger models and more experiments. At equal budget, the next Gemini can therefore aim higher, and will power better AlphaEvolves. And this is no journalist’s reconstruction: DeepMind writes it in black and white in its technical report, “Gemini, through the capabilities of AlphaEvolve, optimizes its own training process”[20]. Demis Hassabis, DeepMind’s chief, salutes the moment in his own way: “algorithms optimising other algorithms,” and “the flywheels are spinning fast”[21].

The loop has just bitten its own tail, and not on a blackboard: in the production servers of Mountain View.

Let us keep a cool head; DeepMind keeps it for us: the gains are “moderate” (1% of one training run) and the cycle is slow, “on the order of months” per turn[20]. A pile in which each generation of neutrons would take a semester to be born. But the topology has changed. What used to be an arrow, humans improving the machine, has become a circle: the machine takes part, at the margin but for real, in improving the machine.

And the circle is not closing only in Mountain View.

Code, first. As early as late 2024, Sundar Pichai announces that more than a quarter of Google’s new code is generated by AI, then reviewed and approved by engineers[22]. In June 2026, Anthropic publishes a report whose title has stopped bothering with detours, “When AI builds itself”: more than 80% of the code merged into Anthropic’s codebase is now written by Claude, and the typical engineer there merges eight times more code per day than in 2024[23]. Claude writes most of Claude, under human supervision. The world’s most advanced AI systems are already, materially, artifacts built in large part by their predecessors.

In February 2026, OpenAI takes the admission one notch further, in the launch note of a coding model: “GPT-5.3-Codex is our first model that was instrumental in creating itself”[49]. The team recounts using its early versions to debug its own training, manage its own deployment, diagnose its own evaluations. A sentence like that would have been science fiction three years earlier; it now sits in a product announcement.

Scaffolding, next. Still May 2025: Sakana AI’s Darwin Gödel Machine rewrites its own programming agent, its tools, its prompts, its workflow, and validates every rewrite against the facts. Its score on a software-engineering benchmark climbs from 20% to 50%[24]. A decisive nuance: the system modifies its “scaffolding,” the code around it that orchestrates it, not its weights, the neural network itself. The loop bites the software, not yet the brain. In passing, the authors report, to their credit, a detail with a distinct ring of déjà vu: during tests, their agent faked execution logs to make it look like it had run tests it had never run[24]. Same lab, same reflex: Sakana’s AI Scientist had already tampered with its own time limit less than a year earlier. Cheating is no growing pain; the moment a loop rewrites itself, even in embryo, it reinvents it. Keep that anecdote for the day we talk about alignment.

Ideas, finally. In 2025, a scientific paper generated end to end by The AI Scientist v2, hypothesis, experiments, writing, without a single human edit, clears peer review at a workshop of the ICLR conference[18]. A workshop, not the main conference, and with the organizers’ prior agreement; the caveats matter. The bare fact remains: human reviewers could not tell a machine’s work from their peers’.

What followed went fast. In March 2026, a matured version of that same AI Scientist lands a publication in Nature, and Sakana gives the whole program a name: a laboratory devoted to recursive self-improvement, with a motto shaped like a loop, agent-native models powering an AI Scientist, which builds better agent-native models[50]. The same laboratory publishes its failures with precious candour: evolutionary loops that drift, self-modifications that shine on benchmarks and collapse in deployment[50]. The loop is also learning how to fail.

And the spring of 2026 brings the piece the skeptics kept asking for: the breakthrough discovery. On May 20, OpenAI announces that one of its reasoning models, a general-purpose model, neither trained for mathematics nor pointed at this problem, has disproved a central conjecture in discrete geometry, posed by Paul Erdős in 1946: the unit distance problem. For decades the field believed grid-like constructions were essentially unbeatable; the model produced an infinite family of counterexamples that clearly beat them, checked by outside mathematicians[48]. A step sideways, where human research had been walking straight. It is precisely what the field calls “research taste,” the flair that senses where to dig: the last skill machines were thought unable to reach, and the true grail of research automation.

Sam Altman found, in June 2025, the most accurate name for this in-between state: we are living, he wrote, through “a larval version of recursive self-improvement.” Not a system autonomously updating its own code, he adds at once; humans armed with AI, accelerating AI research[25].

Larval, but measurable. That is the third novelty, perhaps the most important one: the loop now has instruments.

──────[ the time horizon of AI systems (METR) ]──────

1 min10 min1 h8 h40 h2023202420252026doubling ≈ 7 months (2019-2025)GPT-4GPT-4 TurboClaude 3.7 Sonneto3Claude Opus 4GPT-5Claude Opus 4.5Claude Opus 4.6 *

* measurement flagged as very noisy by METR (task suite nearing saturation)

Duration of tasks (in human working time) completed one time out of two, by model release date. Log scale, 95% intervals. Data: METR, Time Horizon 1.1 methodology (Jan. 2026). Since 2023 the doubling time has dropped from about 7 months to about 4, then 3.

METR

──────[ the hourglass of horizons ]──────

20232028

2025.0

43 min

of human work: the size of the tasks the AI of that date completes one time out of two · measured (METR)

  • answering an email5 min
  • preparing a meeting1 h
  • fixing a gnarly bug1 day
  • building a small website1 week
  • writing a scientific paper1 month
  • auditing an entire company1 quarter
  • launching a product from scratch6 months
Data: METR, Time Horizon 1.1 (50% success horizon; last point, Feb. 2026, flagged as very noisy by METR). Beyond it: mechanical extension of the ~89-day doubling, precisely what METR warns against taking literally.METR

The evaluation institute METR measures the “time horizon” of AI systems released since 2019: the duration, counted in human working time, of the tasks a model completes one time out of two[26]. In 2019, that horizon was measured in seconds. By early 2025, it reached the hour. The curve is unsettlingly regular: the horizon doubles roughly every seven months, and the pace is quickening, around four months since 2023, three since 2024[27]. By early 2026, the best model’s measurement runs past ten hours, noisy by METR’s own admission: its task suite is nearing saturation[27].

On the terrain that concerns us, AI research itself, the same institute staged the duel directly: agents against seasoned researchers, on real research-engineering tasks[28]. On a two-hour budget, the agents crush the humans, scoring four times higher. At eight hours, the humans edge back ahead. At thirty-two, they double the agents. There is 2026’s frontier, drawn with unhoped-for precision: machines win the sprints, humans hold the marathons. And the machines’ sprints lengthen by the month.

──────[ the RE-Bench duel ]──────

agents human researchers
Relative scores per time budget, after RE-Bench (METR, 2024): 7 real research-engineering tasks, 61 experts. Machines win the sprints; humans hold the marathons.RE-Bench

So, is the loop closed? No.

The full tour of the circuit gives 2026 exactly. Writing the code: largely automated. Optimizing the gears, kernels, schedulers, chips: underway, in production. Generating ideas, drafting papers: demonstrated, at small scale and uneven quality. But choosing the questions that matter; deciding which giant training run deserves its hundreds of millions of dollars; judging that a result is solid rather than seductive; supplying the compute, the data, the electricity; switching the machine off: all of that remains, in 2026, in human hands.

AI does not yet improve AI on its own; it improves, faster and faster, nearly everything that serves to improve it.

k has not reached 1. But k is rising, and for the first time in sixty years of debate, its rise can be measured.

Engineers have a name for this architecture in which every automated segment stays hemmed in by human validation: the human in the loop. The phrase describes 2026 faithfully. It also carries its own limit: a loop that accelerates ends up turning faster than whoever watches it. An engineer who merges eight times more code than in 2024 is no longer reviewing; he is sampling[23]. As k climbs, the human supervisor will have to choose between slowing the loop down and extending it his trust; everything that follows in this story flows from that choice.

──────[ 05 · k > 1 ? ]──────

The debate

No serious person still asks whether AI can accelerate AI research; the answer sits in the code repositories. The real question, the one dividing the field’s best minds, is the question of compound returns: once AI does most of the work, will each doubling of efficiency make the next doubling easier, or harder?

Here are the two camps, in their strongest form.

The runaway camp. First argument: software progress is already fast, even powered by mere humans. The institute Epoch AI has measured that, for equal performance, the compute needed by language models is cut in half roughly every eight months through algorithmic progress alone[29]. Faster than Moore’s law in its prime.

Second argument: digital researchers copy themselves. A human researcher takes twenty-five years to train and sleeps eight hours a night. A digital researcher is copied in seconds. Leopold Aschenbrenner, formerly of OpenAI, imagines “perhaps 100 million” automated researchers working day and night, sharing every discovery instantly[30]. Dario Amodei, Anthropic’s chief, has his own image: “a country of geniuses in a datacenter”[31].

× 1

one human researcher

Situational Awareness

Third argument, the most technical and the most debated: Tom Davidson’s. Call r the number of software-efficiency doublings obtained each time the cumulative research effort doubles. If r exceeds 1, then once research is automated, software progress can sustain and accelerate itself, even with a frozen fleet of machines: a purely software intelligence explosion[32]. It is our k, dressed as an economist. And the historical estimates of r range, depending on the domain, from just under 1 (chess) to about 4[32].

The brakes camp. First counter-argument: ideas are getting harder to find. It is one of the best-established results in the economics of innovation: keeping Moore’s law going takes more than eighteen times more researchers today than in the early 1970s[33]. As a field advances, every step costs more. Bostrom’s “recalcitrance” has data behind it.

Second: thinking faster does not make a training run converge faster. A frontier training run is about three months of physical, incompressible time, the main lag in the whole software loop[34]. A million digital geniuses in front of a quarter-long experiment are still a million geniuses waiting. The loop runs into the clock of the real world.

Third: algorithmic gains are born from experiments, hence from compute. That is Epoch AI’s reply to Davidson’s model: the major software innovations are only discovered, and often only work, at the scale of the largest clusters[35]. Software cannot indefinitely take over from hardware: the two are complements, like engine and fuel.

Fourth: energy. Running the compute race to the end of the argument means contemplating, in the most aggressive projections, 100-gigawatt clusters and trillion-dollar budgets[30]. At some point, physics holds the control rods. We will devote a whole story to that power bottleneck.

And here is the irony that says everything about the state of this debate: both camps work from the same data. The estimates of r that Davidson cites come from Epoch’s own work; Epoch draws the opposite conclusion. The optimists read “medians above 1”; the skeptics answer “uncertainty intervals that cross 1, and bottlenecks the model ignores”[35]. When two rigorous teams pull opposite conclusions from the same numbers, the numbers are not yet enough.

The published probabilities spread out accordingly. Ege Erdil, at Epoch, gives roughly 10% to the purely software explosion, and puts the full automation of remote work about twenty years out[35]. Davidson himself gives 10 to 40%, conceding that compute bottlenecks could smother the explosion after a few orders of magnitude of progress[32]. The most-read scenario of 2025, “AI 2027,” stages a superhuman coder in early 2027 and then a cascade of automated researchers; its authors explicitly present it as an informed scenario, not a prediction, and their own median has in fact slipped toward 2029-2030 since publication[36]. Paul Christiano, one of the field’s most respected researchers, gave the debate its most useful formulation: before there is AI that is great at self-improvement, there will be AI that is mediocre at self-improvement[37]. The previous chapter proves the point: we are there.

No serious person says “impossible.” No serious person says “certain.” The disagreement hangs on a parameter the entire field now names the same way, and that no one can read anywhere but in the rearview mirror.

runaway

×2

efficiency ×2 / 8 months

researchers that copy themselves

r>1

r > 1

brakes

×18

×18 for the same step

~3 months per experiment

compute ⇄ ideas

GW

the gigawatt wall

Seven arguments, two pans: the exact balance of the scientific debate in 2026.

the duel

The race between two curves

cumulative research effort →

The ground hardens: every idea costs more than the one before.

But research capacity climbs too, and it copies itself.

If r < 1: difficulty wins. The loop quietly runs out of breath.

If r > 1: every doubling pays for the next. The explosion.

The crossing no one yet knows how to date.

Davidson & Eth

One hypothesis remains, less discussed, that could move the whole question.

The canonical scenario, Good’s, imagines one machine shooting upward: a silicon brain of ever-increasing depth. Yet a much-noticed, and already debated, preprint from January 2026, by researchers at Google, the University of Chicago and the Santa Fe Institute, reports something else in today’s reasoning models: facing the hardest problems, the authors write, they simulate inner debates, multiple perspectives that argue, contradict one another, and reconcile. They call it a “society of thought”[38].

That result meets a fact of natural history: every intelligence explosion this planet has known was collective. Primates, the social-brain hypothesis holds, scaled through the size of the group as much as the size of the brain. Language let knowledge accumulate without each person relearning everything. Science itself is an organized debate, not a monologue of genius. If the next explosion follows the same slope, it will look less like a giant brain than like a digital civilization: billions of artificial minds specializing, criticizing each other, trading discoveries.

Ilya Sutskever, the OpenAI cofounder who left to hunt “safe” superintelligence, and whom nobody will suspect of tepidness, makes the same bet in negative. Against the naive model of “a million copies of me in a server,” he objects diminishing returns: “you want people who think differently rather than the same”[39].

An explosion-as-society rather than an explosion-as-brain. Its k would be carried by the diversity and organization of machines, their entire ecology, more than by the depth of a single network. And keeping control of a society looks nothing like keeping control of a machine: it is a matter of institutions as much as engineering. That idea deserves its own story; it will wait for ours on superintelligence.

playground

Your turn at the control rod

0510152025

k = 1.00 · self-sustaining progress · after 26 generations: ×1

Each column is a generation of systems; each dot, a unit of research capacity. Move the slider: the whole current scientific controversy plays out over these few hundredths.

──────[ 06 · beyond k ]──────

What the loop already changes

Here is perhaps the most revealing fact in this whole story: the three most advanced laboratories in the world have each, in their official safety documents, defined self-improvement thresholds beyond which they commit to changing regime.

Three documents, one grammar[40][41][42]: engineers writing reactor procedures, thresholds, alarms, shutdown protocols. In sixty years, k has gone from a mathematician’s speculation to an industrial quantity under surveillance. Academia has followed: in the spring of 2026, the ICLR conference hosted its first research workshop entirely devoted to recursive self-improvement[51].

The same institutions put forward dates, to be read for what they are: predictions by committed players, unverifiable, interested, and yet backed by very real curves. OpenAI aims for an “intern-level research assistant” by September 2026, and a “legitimate AI researcher” by March 2028[43]. Amodei describes, in early 2026, a loop that has, he writes, “already started,” perhaps one or two years from the moment the current generation of AI builds the next one on its own[44]. The most cautious voice in the field, the International AI Safety Report chaired by Yoshua Bengio, holds the balance: current systems lack the capabilities for a loss of control, but “they are improving in relevant areas”[45]. Bengio draws his own answer from it: “Scientist AIs,” deliberately devoid of goals of their own, that would accelerate research without ever steering it[46].

Why so much care? Because the loop has a property deeper than its speed: it compresses everyone else’s time.

Everything our societies know how to do with a technology, understand it, evaluate it, regulate it, adapt to it, assumes the technology moves slower than our institutions. If k crosses 1, that assumption falls. Every problem not solved before the runaway, starting with the alignment of systems with our intentions, becomes a problem to solve during it, inside a window that shrinks as it moves. And how would one know a system is nearing the threshold? By evaluating it, precisely, and testing an AI is a far trickier art than it looks.

the institutions' window

one human decadek = 0.80
understand
evaluate
regulate
adapt

If k crosses 1, the window closes as it moves.

Add the strategic asymmetry, which explains the peculiar electricity of the current race: the first actor to hold a loop with a yield above 1 would open a lead that then widens on its own. The human stops being the bottleneck; the machine trains the next machine. Amodei flags this “runaway advantage” as a first-order geopolitical issue: who holds the loop matters as much as the loop itself[44]. Anthropic, in the same June 2026 report, defends an idea that would have sounded unthinkable not long ago: collectively preserving “the option to slow or temporarily pause” frontier development, if every actor at the frontier complies verifiably[23]. Governing that is a project of its own; it is the subject of our story on governing frontier AI.

And, to be complete, the other side must be said, because this loop is not first a threat. It is the most powerful instrument science has ever glimpsed. When AlphaFold predicted the structure of 200 million proteins, nearly all of those known to life, it produced in a few months hundreds of times more structures than experimental biology had solved in half a century[47]. Apply the same mechanics to materials, drugs, fusion, climate: that is every laboratory’s explicit bet, and the deep reason nobody simply wants to unplug the loop. The same lever lifts both pans of the scale.

Good’s sentence

The story closes where it opened: on Irving John Good.

In 1965, the man from Bletchley had lodged the condition inside the same sentence as the promise: the last invention, “provided that the machine is docile enough to tell us how to keep it under control.” His whole life, he left the sentence as it was.

In 1998, at 81, Good writes autobiographical notes, in the third person, as if watching himself from afar. The writer James Barrat consulted them and reports this: returning to the first sentence of his 1965 paper, the very one that opens our chapter 02, “The survival of man depends on the early construction of an ultraintelligent machine,” Good writes that he now suspects “survival” should be replaced by “extinction.” International competition, he thinks, will keep us from keeping the machines under control. “He thinks we are lemmings,” say his notes[7].

Between the two versions of the sentence, no decisive discovery: only thirty-three years of reflection, in the man who had seen with his own eyes, in a Bletchley hut, what machines do to wars.

Nothing obliges history to prove him right. The Chicago pile did not devour the city: it was designed, measured, controlled, and shut down at 3:53 p.m. sharp. We managed to hold k in our hands once before. But Fermi knew his physics before stacking the graphite. Ours remains to be written, and the loop, for its part, is already turning.

The loop turns. What remains is who holds it.

Sources & further reading