81Evidence
activeWhat it measures
Kurze Faktenfragen mit eindeutiger Antwort. Misst vor allem, ob ein Modell zugibt, etwas nicht zu wissen, statt zu erfinden.
What it does not measure
Schliessen, Rechnen, laengere Texte.
How the score is composed
Currency95
Discrimination100
Headroom100
Contamination safety55
Gaming resistance55
Methodological rigor75
Reproducibility60
Details
- Benchmaxxing risk
- medium
- Difficulty
- high
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- yes
- Sample size
- 4326
- Letzter Messwert
- 2026-09-10
- Skala
- 0–100 % correct (höher besser)
Best measured models
- 1DeepSeek V4 ProDeepSeekvendor number46.2 % correct
- 2DeepSeek V4.1 FlashDeepSeekvendor number30.1 % correct