71Evidence
activeWhat it measures
Whether a model fixes a real bug in a real Python repository such that the project's own tests pass. Measures agentic work on existing code.
What it does not measure
Languages other than Python, frontend work, architectural decisions, readability.
Die Aufgaben stammen aus oeffentlichen GitHub-Issues, Kontamination ist also plausibel. Herstellerzahlen schwanken stark mit dem Agent-Geruest; nur Werte mit dokumentiertem Scaffold sind vergleichbar.
How the score is composed
Currency93
Discrimination59
Headroom74
Contamination safety55
Gaming resistance55
Methodological rigor100
Reproducibility60
Details
- Benchmaxxing risk
- medium
- Difficulty
- high
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- no
- Sample size
- 500
- Letzter Messwert
- 2026-09-01
- Skala
- 0–100 % solved (höher besser)
Best measured models
- 1GLM-5Z.ai (Zhipu)vendor number77.8 % solved
- 2Muse GlimmerMetavendor number76 % solved
- 3Nex-N2-ProNex AGI (Shanghai Innovation Inst.)vendor number74.4 % solved
- 4Gemini 3.5 FlashGoogle DeepMind71.8 % solved
- 5Nemotron 3 UltraNVIDIAvendor number70.7 % solved
- 6upstage--solar-open2-250bvendor number70.4 % solved