the record of what stopped being hard

Impossible → Routine

Each pair below is one capability with two dates: the day it was a research result, and the day it became something anyone could buy. Both ends carry a source. An end without one does not publish — the build refuses it.

27 dated pairs, spanning to . Every end carries a source.

Sorted by the date it became routine, newest first. Nothing on this site is ordered by payment.

A forecast that was learned, not simulated

Forecasting global weather with a learned model instead of simulated atmospheric physics.

Impossible

GraphCast, a research model, beats the industry gold-standard physics forecast on more than 90% of tested variables.

10-day forecast in under a minute on one TPU machine

source

2 years, 9 months

Routine

WeatherNext 3 becomes the new leader on Brightband's Operational WeatherBench, a live third-party comparison of AI and physics global medium-range models, edging out its predecessor WeatherNext 2.

lowest 2m-temperature error on 26 of the last 30 days in August

source

Real work on a real desktop

Completing multi-step tasks on a real computer — opening applications, clicking through interfaces, finishing the job.

Impossible

The OSWorld benchmark arrives and the best model completes 12.24% of its 369 real-computer tasks.

12.24% on OSWorld, the original 369-task set

source

2 years, 3 months

Routine

Qwen3.8-Max posts 86.1% on OSWorld-Verified, the repaired 2025 revision of the benchmark, with competing agents from two other labs already above 83%.

86.1% on OSWorld-Verified, the 2025 revision

source

Video with sound from a sentence

Generating a short video with matching audio from a written prompt.

Impossible

Google publishes Imagen Video and states it has decided not to release the model or its source code.

source

2 years, 9 months

Routine

Veo 3 generates video with dialogue, effects and music in one pass, sold by the second on a public API.

$0.75 per second of output

source

Machine-checked proofs

Writing formal, machine-verifiable proofs for olympiad-level mathematics.

Impossible

The miniF2F benchmark debuts and its strongest baseline, a GPT-f/PACT prover, closes 29.2% of the olympiad-level test problems in Lean.

29.2% of test problems at Pass@8; 24.6% at Pass@1

source

3 years, 7 months

Routine

DeepSeek-Prover-V2-671B, released as open weights, proves 88.9% of the miniF2F test set at Pass@8192 — and 61.9% of it at Pass@1.

88.9% at Pass@8192; 61.9% at Pass@1

source

Real bugs in real repositories

Handing a model an open bug report from a real repository and getting back a patch that passes the tests.

Impossible

SWE-bench reports that the best model it tested resolves a mere 1.96% of real GitHub issues.

1.96% of issues resolved

source

1 year, 4 months

Routine

A model sold by the million tokens reports 62.3% on the same benchmark's human-verified subset.

62.3% on SWE-bench Verified

source

Competitive programming

Solving timed competitive programming problems, judged by hidden test cases.

Impossible

AlphaCode, a purpose-built research system, averages a ranking in the top 54.3% — mid-pack among humans — across Codeforces contests with more than 5,000 entrants.

top 54.3% of entrants

source

2 years, 11 months

Routine

o3, a general-purpose reasoning model, reaches a 2724 Codeforces rating — the 99.8th percentile — and a gold-medal score at the 2024 International Olympiad in Informatics.

Codeforces 2724, 99.8th percentile

source

Competition mathematics

Solving competition mathematics problems and showing the working.

Impossible

The MATH paper concludes that scaling to 40% accuracy would need around 10^35 parameters, which it calls impractical.

GPT-3, 175B parameters: 5.2%

source

3 years, 10 months

Routine

A model published with open weights reports 97.3% on MATH-500 and is distilled into versions from 1.5B parameters up.

97.3% pass@1 on MATH-500

source

Google-proof questions

Answering graduate-level science questions written so that even unrestricted web search does not help.

Impossible

GPQA debuts with domain PhDs scoring 65%, skilled non-experts with full web access 34%, and the strongest GPT-4 baseline 39%.

best model 39%; PhD experts 65%

source

1 year, 2 months

Routine

DeepSeek-R1, its weights free to download, reports 71.5% on GPQA Diamond, the benchmark's hardest split.

71.5% pass@1

source

Conversation at speaking speed

Holding a spoken conversation with a model at the pace people actually talk.

Impossible

ChatGPT gets voice as a relay — Whisper transcribes the speech, the text model composes a reply, and a separate synthesizer reads it out.

source

1 year, 1 month

Routine

GPT-4o's system card records the model answering speech directly in 320 milliseconds on average — the speed of human conversational turns.

232 ms fastest, 320 ms average

source

Frontier weights you can download

Downloading the weights of a largest-class language model and running them yourself.

Impossible

OpenAI withholds the full GPT-2, publishing a much smaller version and keeping back the datasets and training code.

source

5 years, 5 months

Routine

Meta publishes Llama 3.1's weights for anyone to download, customise and fine-tune.

405 billion parameters

source

How much a model reads at once

Putting a long document into a model's input and asking questions about all of it.

Impossible

GPT-3 is published with a fixed context window of 2,048 tokens for every model size.

2,048 tokens

source

4 years

Routine

Google opens a two-million-token context window to every developer on its public API.

2,000,000 tokens

source

A capable model in a phone

Running a genuinely capable language model on a phone, offline.

Impossible

GPT-3, served from a datacenter at 175 billion parameters, scores 43.9% on the new 57-subject MMLU exam.

175B parameters, 43.9% MMLU

source

3 years, 7 months

Routine

Phi-3-mini, quantized to 1.8 GB, runs fully offline on an iPhone 14 at more than 12 tokens per second and scores 69% on MMLU.

3.8B parameters on a phone, 69% MMLU

source

A whole song, generated

Generating a complete song, music and singing, as audio.

Impossible

OpenAI's Jukebox generates music with singing in the raw audio domain, at about nine hours of sampling per minute of song.

about 9 hours of compute per minute of audio

source

3 years, 10 months

Routine

Suno's v3 model makes full two-minute songs in seconds and is available to all users.

a two-minute song in seconds

source

Fine-tuning on one GPU

Fine-tuning a large language model into your own specialized version on a single machine.

Impossible

Microsoft researchers measure full fine-tuning of GPT-3 at 1.2 TB of GPU memory and call deploying independent fine-tuned copies prohibitively expensive.

1.2 TB of GPU memory

source

1 year, 11 months

Routine

QLoRA fine-tunes a 65-billion-parameter model on a single 48 GB GPU in 24 hours, reaching 99.3% of ChatGPT's level on the Vicuna benchmark.

one 48 GB GPU, 24 hours

source

Transcription at human accuracy

Transcribing ordinary recorded conversation about as accurately as a person does.

Impossible

Microsoft researchers report the first human-parity error rate on the standard Switchboard task.

5.9% word error rate

source

6 years, 1 month

Routine

OpenAI releases Whisper's models and inference code, reporting accuracy and robustness approaching human transcribers.

trained on 680,000 hours of audio; weights public

source

An image from a sentence

Producing an original image from a written description of it.

Impossible

DALL-E is published as a research paper: a transformer trained to generate images from text captions.

source

1 year, 5 months

Routine

Stable Diffusion's weights are published for anyone to download and run on one consumer graphics card.

6.9 GB of VRAM

source

The shape of a protein

Getting the three-dimensional shape a protein folds into from its sequence of amino acids.

Impossible

AlphaFold's CASP14 result answers what the assessors had called a 50-year grand challenge in biology.

median 92.4 GDT across all targets

source

1 year, 7 months

Routine

Predicted structures for over 200 million proteins are free to search and to bulk-download.

200 million structures, no cost

source

Translation at human accuracy

Translating news text from one language into another as accurately as a human translator.

Impossible

Microsoft reports human parity on Chinese-to-English news translation, for one language pair on one test set.

newstest2017, judged by bilingual evaluators

source

4 years, 3 months

Routine

Meta open-sources NLLB-200, one model translating 200 languages, with its benchmark and training code.

200 languages; 44% better BLEU than the previous state of the art

source

A car with nobody in it

Riding in a car on public roads with no human driver aboard.

Impossible

No vehicle finishes DARPA's Grand Challenge; the top-scoring one travels 7.5 miles of the course.

7.5 miles of a 142-mile course

source

16 years, 6 months

Routine

Waymo opens its fully driverless service to the general public in metro Phoenix.

service area larger than San Francisco

source

No-limit poker

Beating professional players at no-limit Texas hold'em poker.

Impossible

Libratus beats four heads-up specialists by $1,766,250 in chips over 120,000 hands, drawing on about 600 nodes of the Bridges supercomputer.

roughly 600 of Bridges' 846 compute nodes

source

2 years, 5 months

Routine

Pluribus beats professionals at six-player no-limit hold'em after computing its strategy in eight days and 12,400 core hours, playing live on 28 cores.

12,400 core hours to train; 28 cores in live play

source

Training ImageNet

Training an image classifier to state-of-the-art accuracy on ImageNet.

Impossible

AlexNet wins ILSVRC-2012 with a 15.3% top-5 test error, against 26.2% for the runner-up, after five to six days of training on two GTX 580 GPUs.

five to six days on two consumer GPUs

source

5 years, 8 months

Routine

fast.ai trains ImageNet to 93% accuracy in 18 minutes on 16 rented public cloud instances, for about $40.

18 minutes, about $40

source

Professional-level Go

Playing Go at the level of professional players.

Impossible

AlphaGo beats the European champion 5–0, a feat its Nature paper says was previously thought to be at least a decade away.

5–0; first program to beat a professional at full-size Go

source

2 years, 3 months

Routine

Facebook releases ELF OpenGo's trained model and code under a BSD license, after a 14–0 record against four top-30 professionals playing on a single GPU.

14–0 versus top-30 professionals; one GPU

source

Diabetic eye screening

Detecting diabetic eye disease from a photograph of the retina.

Impossible

A deep network detects diabetic retinopathy in retinal photographs on par with the ophthalmologists it was measured against, in a JAMA paper.

F-score 0.95 versus a median 0.91 for eight ophthalmologists

source

1 year, 4 months

Routine

The FDA permits marketing of IDx-DR, the first device authorized to use AI autonomously to detect a medical condition — its results do not require review by a specialized clinician.

first autonomous AI diagnostic authorized by the FDA

source

Speech built one sample at a time

Generating human-sounding speech directly as a raw audio waveform.

Impossible

DeepMind publishes WaveNet and calls building up samples one step at a time computationally expensive.

source

1 year

Routine

The production version generates the Google Assistant voices for US English and Japanese on all platforms.

50 ms of compute per second of speech, 1,000x faster than the research model

source

Which bird is that

Naming the species of bird in a photograph.

Impossible

xkcd uses checking whether a photo is of a bird as its example of a task needing a research team and five years.

source

8 months

Routine

Cornell's free Merlin Bird Photo ID names the species in an uploaded photo.

400 species, free

source

Reading handwriting

Reading handwriting by machine.

Impossible

A back-propagation network reads handwritten zip code digits supplied by the U.S. Postal Service, with a 1% error rate and a 9% reject rate.

1% error, 9% reject, isolated digits

source

25 years, 4 months

Routine

Google ships Handwriting Input for Android, reading printed and cursive writing in 82 languages, with or without an internet connection.

82 languages, 20 scripts

source

Grandmaster chess

Playing chess above the strongest human players.

Impossible

IBM's Deep Blue beats world champion Garry Kasparov 3.5–2.5 in their six-game rematch in New York.

3.5–2.5 over six games

source

12 years, 3 months

Routine

Pocket Fritz 4 scores 9.5/10 at the Mercosur Cup in Buenos Aires running on a handheld Pocket PC, an Elo performance of 2938 that included outplaying a 2522 grandmaster.

9.5/10; Elo performance 2938

source

The same pairs as data: /dataset/deltas.csv