40Evidence
not yet measuredWhat it measures
23 Teilaufgaben, bei denen Modelle frueher schlechter waren als Menschen.
What it does not measure
Aktuelle Grenzfaehigkeiten.
Von aktuellen Modellen weitgehend ausgereizt.
How the score is composed
Currency50
Discrimination50
Headroom50
Contamination safety9
Gaming resistance55
Methodological rigor50
Reproducibility60
Details
- Benchmaxxing risk
- medium
- Difficulty
- low
- Public test set
- yes
- Human-verified
- no
- LLM-judged
- no
- Skala
- 0–100 % correct (höher besser)
Best measured models
No results ingested yet.