impossible → routine
Google-proof questions
Answering graduate-level science questions written so that even unrestricted web search does not help.
Impossible
GPQA debuts with domain PhDs scoring 65%, skilled non-experts with full web access 34%, and the strongest GPT-4 baseline 39%.
best model 39%; PhD experts 65%
source1 year, 2 months
Routine
DeepSeek-R1, its weights free to download, reports 71.5% on GPQA Diamond, the benchmark's hardest split.
71.5% pass@1
source