Sorted by the date it became routine, newest first. Nothing on this site is ordered by payment.
Forecasting global weather with a learned model instead of simulated atmospheric physics.
Impossible
2023-11-14 GraphCast, a research model, beats the industry gold-standard physics forecast on more than 90% of tested variables.
10-day forecast in under a minute on one TPU machine
source 2 years, 9 months
Routine
2026-09-02 WeatherNext 3 becomes the new leader on Brightband's Operational WeatherBench, a live third-party comparison of AI and physics global medium-range models, edging out its predecessor WeatherNext 2.
lowest 2m-temperature error on 26 of the last 30 days in August
source Completing multi-step tasks on a real computer — opening applications, clicking through interfaces, finishing the job.
Impossible
2024-04-11 The OSWorld benchmark arrives and the best model completes 12.24% of its 369 real-computer tasks.
12.24% on OSWorld, the original 369-task set
source 2 years, 3 months
Routine
2026-08-03 Qwen3.8-Max posts 86.1% on OSWorld-Verified, the repaired 2025 revision of the benchmark, with competing agents from two other labs already above 83%.
86.1% on OSWorld-Verified, the 2025 revision
source Generating a short video with matching audio from a written prompt.
Impossible
2022-10-05 Google publishes Imagen Video and states it has decided not to release the model or its source code.
source 2 years, 9 months
Routine
2025-07-17 Veo 3 generates video with dialogue, effects and music in one pass, sold by the second on a public API.
$0.75 per second of output
source Writing formal, machine-verifiable proofs for olympiad-level mathematics.
Impossible
2021-08-31 The miniF2F benchmark debuts and its strongest baseline, a GPT-f/PACT prover, closes 29.2% of the olympiad-level test problems in Lean.
29.2% of test problems at Pass@8; 24.6% at Pass@1
source 3 years, 7 months
Routine
2025-04-30 DeepSeek-Prover-V2-671B, released as open weights, proves 88.9% of the miniF2F test set at Pass@8192 — and 61.9% of it at Pass@1.
88.9% at Pass@8192; 61.9% at Pass@1
source Handing a model an open bug report from a real repository and getting back a patch that passes the tests.
Impossible
2023-10-10 SWE-bench reports that the best model it tested resolves a mere 1.96% of real GitHub issues.
1.96% of issues resolved
source 1 year, 4 months
Routine
2025-02-24 A model sold by the million tokens reports 62.3% on the same benchmark's human-verified subset.
62.3% on SWE-bench Verified
source Solving timed competitive programming problems, judged by hidden test cases.
Impossible
2022-02-08 AlphaCode, a purpose-built research system, averages a ranking in the top 54.3% — mid-pack among humans — across Codeforces contests with more than 5,000 entrants.
top 54.3% of entrants
source 2 years, 11 months
Routine
2025-02-03 o3, a general-purpose reasoning model, reaches a 2724 Codeforces rating — the 99.8th percentile — and a gold-medal score at the 2024 International Olympiad in Informatics.
Codeforces 2724, 99.8th percentile
source Solving competition mathematics problems and showing the working.
Impossible
2021-03-05 The MATH paper concludes that scaling to 40% accuracy would need around 10^35 parameters, which it calls impractical.
GPT-3, 175B parameters: 5.2%
source 3 years, 10 months
Routine
2025-01-22 A model published with open weights reports 97.3% on MATH-500 and is distilled into versions from 1.5B parameters up.
97.3% pass@1 on MATH-500
source Answering graduate-level science questions written so that even unrestricted web search does not help.
Impossible
2023-11-20 GPQA debuts with domain PhDs scoring 65%, skilled non-experts with full web access 34%, and the strongest GPT-4 baseline 39%.
best model 39%; PhD experts 65%
source 1 year, 2 months
Routine
2025-01-22 DeepSeek-R1, its weights free to download, reports 71.5% on GPQA Diamond, the benchmark's hardest split.
71.5% pass@1
source Holding a spoken conversation with a model at the pace people actually talk.
Impossible
2023-09-25 ChatGPT gets voice as a relay — Whisper transcribes the speech, the text model composes a reply, and a separate synthesizer reads it out.
source 1 year, 1 month
Routine
2024-10-25 GPT-4o's system card records the model answering speech directly in 320 milliseconds on average — the speed of human conversational turns.
232 ms fastest, 320 ms average
source Downloading the weights of a largest-class language model and running them yourself.
Impossible
2019-02-22 OpenAI withholds the full GPT-2, publishing a much smaller version and keeping back the datasets and training code.
source 5 years, 5 months
Routine
2024-07-23 Meta publishes Llama 3.1's weights for anyone to download, customise and fine-tune.
405 billion parameters
source Putting a long document into a model's input and asking questions about all of it.
Impossible
2020-05-28 GPT-3 is published with a fixed context window of 2,048 tokens for every model size.
2,048 tokens
source 4 years
Routine
2024-06-27 Google opens a two-million-token context window to every developer on its public API.
2,000,000 tokens
source Running a genuinely capable language model on a phone, offline.
Impossible
2020-09-07 GPT-3, served from a datacenter at 175 billion parameters, scores 43.9% on the new 57-subject MMLU exam.
175B parameters, 43.9% MMLU
source 3 years, 7 months
Routine
2024-04-22 Phi-3-mini, quantized to 1.8 GB, runs fully offline on an iPhone 14 at more than 12 tokens per second and scores 69% on MMLU.
3.8B parameters on a phone, 69% MMLU
source Generating a complete song, music and singing, as audio.
Impossible
2020-04-30 OpenAI's Jukebox generates music with singing in the raw audio domain, at about nine hours of sampling per minute of song.
about 9 hours of compute per minute of audio
source 3 years, 10 months
Routine
2024-03-21 Suno's v3 model makes full two-minute songs in seconds and is available to all users.
a two-minute song in seconds
source Fine-tuning a large language model into your own specialized version on a single machine.
Impossible
2021-06-17 Microsoft researchers measure full fine-tuning of GPT-3 at 1.2 TB of GPU memory and call deploying independent fine-tuned copies prohibitively expensive.
1.2 TB of GPU memory
source 1 year, 11 months
Routine
2023-05-23 QLoRA fine-tunes a 65-billion-parameter model on a single 48 GB GPU in 24 hours, reaching 99.3% of ChatGPT's level on the Vicuna benchmark.
one 48 GB GPU, 24 hours
source Transcribing ordinary recorded conversation about as accurately as a person does.
Impossible
2016-10-18 Microsoft researchers report the first human-parity error rate on the standard Switchboard task.
5.9% word error rate
source 6 years, 1 month
Routine
2022-12-06 OpenAI releases Whisper's models and inference code, reporting accuracy and robustness approaching human transcribers.
trained on 680,000 hours of audio; weights public
source Producing an original image from a written description of it.
Impossible
2021-02-24 DALL-E is published as a research paper: a transformer trained to generate images from text captions.
source 1 year, 5 months
Routine
2022-08-22 Stable Diffusion's weights are published for anyone to download and run on one consumer graphics card.
6.9 GB of VRAM
source Getting the three-dimensional shape a protein folds into from its sequence of amino acids.
Impossible
2020-11-30 AlphaFold's CASP14 result answers what the assessors had called a 50-year grand challenge in biology.
median 92.4 GDT across all targets
source 1 year, 7 months
Routine
2022-07-28 Predicted structures for over 200 million proteins are free to search and to bulk-download.
200 million structures, no cost
source Translating news text from one language into another as accurately as a human translator.
Impossible
2018-03-14 Microsoft reports human parity on Chinese-to-English news translation, for one language pair on one test set.
newstest2017, judged by bilingual evaluators
source 4 years, 3 months
Routine
2022-07-11 Meta open-sources NLLB-200, one model translating 200 languages, with its benchmark and training code.
200 languages; 44% better BLEU than the previous state of the art
source Riding in a car on public roads with no human driver aboard.
Impossible
2004-03-13 No vehicle finishes DARPA's Grand Challenge; the top-scoring one travels 7.5 miles of the course.
7.5 miles of a 142-mile course
source 16 years, 6 months
Routine
2020-10-08 Waymo opens its fully driverless service to the general public in metro Phoenix.
service area larger than San Francisco
source Beating professional players at no-limit Texas hold'em poker.
Impossible
2017-01-31 Libratus beats four heads-up specialists by $1,766,250 in chips over 120,000 hands, drawing on about 600 nodes of the Bridges supercomputer.
roughly 600 of Bridges' 846 compute nodes
source 2 years, 5 months
Routine
2019-07-11 Pluribus beats professionals at six-player no-limit hold'em after computing its strategy in eight days and 12,400 core hours, playing live on 28 cores.
12,400 core hours to train; 28 cores in live play
source Training an image classifier to state-of-the-art accuracy on ImageNet.
Impossible
2012-12-03 AlexNet wins ILSVRC-2012 with a 15.3% top-5 test error, against 26.2% for the runner-up, after five to six days of training on two GTX 580 GPUs.
five to six days on two consumer GPUs
source 5 years, 8 months
Routine
2018-08-10 fast.ai trains ImageNet to 93% accuracy in 18 minutes on 16 rented public cloud instances, for about $40.
18 minutes, about $40
source Playing Go at the level of professional players.
Impossible
2016-01-27 AlphaGo beats the European champion 5–0, a feat its Nature paper says was previously thought to be at least a decade away.
5–0; first program to beat a professional at full-size Go
source 2 years, 3 months
Routine
2018-05-02 Facebook releases ELF OpenGo's trained model and code under a BSD license, after a 14–0 record against four top-30 professionals playing on a single GPU.
14–0 versus top-30 professionals; one GPU
source Detecting diabetic eye disease from a photograph of the retina.
Impossible
2016-11-29 A deep network detects diabetic retinopathy in retinal photographs on par with the ophthalmologists it was measured against, in a JAMA paper.
F-score 0.95 versus a median 0.91 for eight ophthalmologists
source 1 year, 4 months
Routine
2018-04-11 The FDA permits marketing of IDx-DR, the first device authorized to use AI autonomously to detect a medical condition — its results do not require review by a specialized clinician.
first autonomous AI diagnostic authorized by the FDA
source Generating human-sounding speech directly as a raw audio waveform.
Impossible
2016-09-08 DeepMind publishes WaveNet and calls building up samples one step at a time computationally expensive.
source 1 year
Routine
2017-10-04 The production version generates the Google Assistant voices for US English and Japanese on all platforms.
50 ms of compute per second of speech, 1,000x faster than the research model
source Naming the species of bird in a photograph.
Impossible
2014-09-24 xkcd uses checking whether a photo is of a bird as its example of a task needing a research team and five years.
source 8 months
Routine
2015-06-05 Cornell's free Merlin Bird Photo ID names the species in an uploaded photo.
400 species, free
source Reading handwriting by machine.
Impossible
1989-11-27 A back-propagation network reads handwritten zip code digits supplied by the U.S. Postal Service, with a 1% error rate and a 9% reject rate.
1% error, 9% reject, isolated digits
source 25 years, 4 months
Routine
2015-04-15 Google ships Handwriting Input for Android, reading printed and cursive writing in 82 languages, with or without an internet connection.
82 languages, 20 scripts
source Playing chess above the strongest human players.
Impossible
1997-05-11 IBM's Deep Blue beats world champion Garry Kasparov 3.5–2.5 in their six-game rematch in New York.
3.5–2.5 over six games
source 12 years, 3 months
Routine
2009-08-27 Pocket Fritz 4 scores 9.5/10 at the Mercosur Cup in Buenos Aires running on a handheld Pocket PC, an Elo performance of 2938 that included outplaying a 2522 grandmaster.
9.5/10; Elo performance 2938
source