90Evidence
activeWhat it measures
Verstehen und Schliessen ueber echte lange Dokumente statt synthetischer Nadeln.
What it does not measure
Sehr kurze Aufgaben, Werkzeugnutzung.
How the score is composed
Currency95
Discrimination100
Headroom100
Contamination safety55
Gaming resistance100
Methodological rigor100
Reproducibility60
Details
- Benchmaxxing risk
- low
- Difficulty
- high
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- no
- Sample size
- 503
- Letzter Messwert
- 2026-09-10
- Skala
- 0–100 % correct (höher besser)
Best measured models
- 1Nemotron 3 UltraNVIDIAvendor number61.9 % correct
- 2DeepSeek V4.1 FlashDeepSeekvendor number44.7 % correct
- 3DeepSeek V4 ProDeepSeekvendor number40.2 % correct