learn · advanced
How to think about what comes next
After this page: You can weigh a confident AI forecast, name the walls scaling has hit before and what happened to each, and watch the few indicators that actually move before the headlines do.
Assumes Why bigger kept getting better, What a benchmark measures, and what it does not, The safety arguments, steelmanned, What AI does to work
In August 2021 Jacob Steinhardt, a Berkeley statistician, paid professional forecasters to predict progress on four AI benchmarks, with a $5,000 prize pool for each benchmark so that accuracy would cost something to get wrong. Professional forecasters are people whose living depends on being calibrated rather than interesting. On MATH, a competition-mathematics benchmark where the best system of 2021 scored 6.9 per cent and most humans stay below 50, their prediction for June 2022 was 12.7 per cent.
The grading date arrived. The actual score was 50.3 per cent.
Steinhardt had thought the predictions overheated, and said so at the time. His grade on himself, once the results were in: "I clearly thought the forecasts on MATH were aggressive in terms of how much progress they predicted, whereas it turned out they weren't aggressive enough." The objection that the stakes were too small to reward real effort has an answer in the same post. A second group he commissioned, chosen for its track record and paid at a high hourly rate, had put a probability distribution on MATH performance in 2026. The 2022 result landed at the 75th percentile of it, and "progress still outpaced the forecast."
You have the whole ladder behind you now, and this page will not add a forecast to it. It is about how to weigh everyone who offers one.
Neither reflex has the better record
An episode like that invites a cheap conclusion: the doubters are always wrong, believe the hype. The field's own history blocks it, because the doubters have been paid out twice. The first reckoning came as budgets: funding for machine translation collapsed after a 1966 government review, and a wider withdrawal followed the British survey published early in 1973. The second came in the 1980s, when researchers warned in 1984 that expectations had run ahead of the results and the industry those expectations were attached to turned a couple of years later. Both freezes lasted years, and the record of the winters rewards reading in full, not least because each document behind the folklore is narrower and stranger than "they overpromised": one asked whether the demand for the product even existed, and the other graded two of the field's three categories respectable while failing the claim that connected them into a field.
So the two reflexes on offer at every AI headline, "they always overhype" and "they laughed at this too," are both true as history and both empty as arguments. A record that contains the winters and the 12.7 per cent contains a precedent for every possible next outcome. What AI does to work showed the sharpest instance of how that record gets misused: a number built as an engineering projection, containing no economics and no clock, circulated for a decade as an employment forecast. The failure was never optimism or pessimism. It was quoting a claim of one kind as if it were a claim of another kind, and the same field that punishes enthusiasm without evidence punishes cynicism without evidence, because they are the same defect wearing different moods.
The silence a forecast inherits
Why bigger kept getting better ended somewhere deliberate: the field can predict the loss of a model nobody has built yet, and cannot explain the abilities of the models it already has. That page banked a scepticism. This is where it gets spent.
A confident capability forecast is usually a scaling law extended forward, plus a translation from loss into abilities. Both steps leak. The extension assumes the inputs keep arriving, which is a question about money and electricity rather than about learning. The translation crosses exactly the gap the field cannot explain, the one where the emergence dispute lives. Extrapolation, that page said, reports that a curve has continued, never why, and so never what could make it stop. A forecast built on it inherits that silence and papers over it with confidence, which is why "the scaling laws say" begins a strong sentence about loss and a weak sentence about the future.
The symmetrical mistake gets less scrutiny and deserves the same amount. A wall is a forecast too.
The walls, and what happened to each
The data wall's serious version was published in October 2022, when researchers at Epoch, a group that measures the inputs to AI progress, estimated the stock of text on the internet against the growth of training datasets and concluded that "the stock of high-quality language data will be exhausted soon; likely before 2026." In 2024 the same team re-measured, and the revised paper finds instead that models will be trained on datasets roughly the size of "the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained." The wall is real on both fits. What moved is where it stands, by as much as six years, and the movement is the lesson. A wall argument is built from the same material as the curve it blocks, fitted trends extended forward, and fitted trends have moved before under better measurement.
The power wall has the most physical substance, because its inputs are turbines, transformers and permits, none of which learns from examples. Its serious version is not a slogan but an audit. In August 2024 Epoch worked through the constraints on training runs one at a time, electric power, chip manufacturing, data and latency, and concluded that runs of up to 2e29 FLOP "would be feasible by the end of the decade." A FLOP is one arithmetic operation, and that many of them is a scale the report puts at roughly 10,000 times the models of its own year. Read as a wall or read as a runway, the report is being read wrong either way. Its value is that every step is quantified and dated, so a reader a few years on can check which constraint actually bound first. That is what a wall argument looks like when it is built to be graded.
The money wall is argued from inside the money. In September 2023 David Cahn, a partner at the venture firm Sequoia, asked one question of the AI infrastructure build-out, "Where is all the revenue?", and computed the annual gap between what the spending implied and what the products earned. Nine months later he ran the arithmetic again and his $200 billion question had grown, on his own method, to a $600 billion one. He also reports a major rebuttal to the earlier piece, that the spending on AI chips "is like building railroads" and the trains will eventually come, and grants part of it before answering that a railroad enjoyed monopoly pricing and a commodity metered by the hour does not.
Two things set this wall apart, and both are historical. It is the only one that has actually fallen on the field before: that is what a winter was, capital leaving on a stated disappointment. And its fall settled nothing about capability, which kept its questions and answered some of them decades later. The winters entry keeps the decomposition that makes a collapse forecast weighable at all: capability stalling, expectations correcting and capital leaving are three different claims, and the record shows them coming apart. The data wall moved when re-measured. The power wall has never yet bound, and now has a dated audit anyone can grade. The money wall has bound twice, and settled nothing either time. Any confident sentence about the walls that does not survive contact with those three histories was not built out of them.
Two curves, years apart
What AI does to work ended on a distinction built for one subject that generalises to all of this one. A projection argues from what looks possible and can be built at a desk before anything has happened. A signal is a count of something that already occurred, in a named group, over a stated period, that would have come out differently if the claim were false. Underneath that distinction sat a measured fact: what a technology can do and what the world has absorbed are different curves, years apart, moved by different forces.
Generalised, that fact reorganises most arguments about what comes next. Capability changes on the day a training run finishes. Absorption moves at the pace of procurement, retraining, regulation, workflow redesign and trust, the pace the gap between a demo and a deployment already showed you at the scale of a single product. The two curves can carry opposite headlines in the same week and both be right. The safety arguments keep a version of the same split: how far capability goes and whether institutions absorb it on schedule are separate priors, and evidence bears on them separately.
So the first question for any claim about what comes next is which curve it is about. A capability claim can be true for years before anything in your life registers it. A diffusion claim can rearrange work for years with the frontier frozen, because the distance between what has been built and what has been absorbed is already years wide. Confusing the two curves is how a person ends up certain that nothing is happening while everything is being installed, or certain that the world just changed when only a leaderboard did.
The watchlist
Refusing to forecast earns nothing on its own. The confident forecast and the shrug are the two failures on offer at every headline, and they are both ways of not watching. The alternative to a forecast is not silence; it is a watchlist. Four instruments move before the headlines do, and each carries the caveat that keeps it honest.
The compute doubling time. Before about 2010, training compute grew in line with Moore's law, doubling roughly every 20 months. Since the deep-learning era began it has doubled about every six months, a rate measured in 2022 across the notable models of three eras. That number is the input side of every scaling story, and it is a spending curve as much as an engineering one, which makes it the money wall's own early indicator. Watch whether the doubling time holds or stretches. It moves before capabilities, and years before consequences.
The benchmark replacement cycle. In 2019 the authors of GLUE, then the standard suite of language-understanding tests, reported that performance had "recently surpassed the level of non-expert humans, suggesting limited headroom for further research" on a benchmark "introduced a little over one year ago", and published its harder successor themselves. A test being worn out and replaced within a year is itself a measurement. Everything you know about scores still applies: near the ceiling a score measures the test, a jump can live in the metric rather than the model, and a flat line is weak evidence that nothing is improving underneath. The replacement cycle survives all three caveats because it does not require trusting any single score. It asks only how fast the field is wearing out its rulers.
The open-to-frontier gap. Some builders publish their models' trained weights for anyone to download and run, while the most capable models, the frontier, mostly stay behind paid interfaces. The distance between the two is tracked: one measurement published in May 2026 put the average lag at four months since January of that year. Movement in either direction is informative. A widening gap says the next increment of capability is getting expensive, which is the walls starting to sort builders by their budgets. A narrowing gap says frontier capability becomes copyable soon after it exists, which matters for everything the safety arguments attach to a release that cannot be recalled.
Prices. Diffusion runs on them, because a capability becomes a consequence at a price point, not at a press event. The measured fall is steep and uneven: one tracker found, as of its November 2025 data update, the price of matching a fixed level of performance on PhD-level science questions falling 40-fold per year, with the rate across other tasks ranging from 9-fold to 900-fold, and attached its own caveat that the steepest drops were recent and the least certain to persist. A price series catches what capability forecasts miss entirely: the world changing because last year's ability became too cheap to ignore, with the frontier not moving at all.
A date, a number, an owner
What the four instruments share is what the opening contest had: a date, a number, and an owner who can be seen to be wrong. The safety arguments gave you the question that sorts a disagreement: what would count as finding out. Pointed at the future, it becomes three questions to put to any confident claim about what comes next. What is it extrapolating? Which curve is it about? And what has its author agreed to be wrong about, by when?
Most claims fail the third question, and the failure is itself the finding, because it means the claim was built so that no arriving fact could embarrass it. An AGI timeline is the clearest case: the term fixes no threshold anyone has agreed on, which makes "what would you accept as arrival?" the first question to put to its author rather than the last. Steinhardt's forecasters were wrong by a factor of four and still produced more knowledge than a year of unfalsifiable commentary, because the miss had a size, a direction and a date, and all three became public property the day the grade came in.
The gradeable claims are the ones worth your attention, and grading is also the method of the site you are reading: a fact carries its date and its source, a value that vanished keeps its last-known reading with the date showing, and what was written stays checkable against what happened. Not because the method predicts anything. Because it is the only way to be usefully wrong about a subject that has embarrassed confident people in both directions for seventy years.
When the next headline arrives about a system that does not exist today, doing something nobody can currently demonstrate, you will not be able to tell whether it is true, and you do not need to. Ask what it extrapolates, which curve it lives on, and what its author would accept as a miss, then file whatever survives onto the watchlist. The future of this field has never been legible in advance. It has been legible in arrears every time, and the craft is to be the person who wrote down, before the fact, what to check.