ModelRadar
/
← All benchmarks

SWE-bench Verified

coding · repo-level-bugfix↗ Quelle
71Evidence
active

What it measures

Whether a model fixes a real bug in a real Python repository such that the project's own tests pass. Measures agentic work on existing code.

What it does not measure

Languages other than Python, frontend work, architectural decisions, readability.

Die Aufgaben stammen aus oeffentlichen GitHub-Issues, Kontamination ist also plausibel. Herstellerzahlen schwanken stark mit dem Agent-Geruest; nur Werte mit dokumentiertem Scaffold sind vergleichbar.

How the score is composed

Currency93
Discrimination59
Headroom74
Contamination safety55
Gaming resistance55
Methodological rigor100
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
high
Public test set
yes
Human-verified
yes
LLM-judged
no
Sample size
500
Letzter Messwert
2026-09-01
Skala
0–100 % solved (höher besser)

Best measured models

  1. 1GLM-5Z.ai (Zhipu)vendor number77.8 % solved
  2. 2Muse GlimmerMetavendor number76 % solved
  3. 3Nex-N2-ProNex AGI (Shanghai Innovation Inst.)vendor number74.4 % solved
  4. 4Gemini 3.5 FlashGoogle DeepMind71.8 % solved
  5. 5Nemotron 3 UltraNVIDIAvendor number70.7 % solved
  6. 6upstage--solar-open2-250bvendor number70.4 % solved