66Evidence
activeWhat it measures
Aufgaben, die den korrekten Einsatz echter Bibliotheken verlangen, nicht nur Algorithmik.
What it does not measure
Mehrdateien-Aenderungen, Projektkontext.
How the score is composed
Currency95
Discrimination0
Headroom100
Contamination safety55
Gaming resistance100
Methodological rigor100
Reproducibility60
Details
- Benchmaxxing risk
- low
- Difficulty
- medium
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- no
- Sample size
- 1140
- Letzter Messwert
- 2026-09-10
- Skala
- 0–100 % pass@1 (höher besser)
Best measured models
- 1DeepSeek V4 ProDeepSeekvendor number56.8 % pass@1
- 2DeepSeek V4.1 FlashDeepSeekvendor number56.8 % pass@1