impossible → routine

Google-proof questions

Answering graduate-level science questions written so that even unrestricted web search does not help.

Impossible

GPQA debuts with domain PhDs scoring 65%, skilled non-experts with full web access 34%, and the strongest GPT-4 baseline 39%.

best model 39%; PhD experts 65%

source

1 year, 2 months

Routine

DeepSeek-R1, its weights free to download, reports 71.5% on GPQA Diamond, the benchmark's hardest split.

71.5% pass@1

source

All dated pairs