impossible → routine

Real bugs in real repositories

Handing a model an open bug report from a real repository and getting back a patch that passes the tests.

Impossible

SWE-bench reports that the best model it tested resolves a mere 1.96% of real GitHub issues.

1.96% of issues resolved

source

1 year, 4 months

Routine

A model sold by the million tokens reports 62.3% on the same benchmark's human-verified subset.

62.3% on SWE-bench Verified

source

All dated pairs