39Evidence
not yet measuredWhat it measures
Laengenbereinigte Siegquote gegen ein Referenzmodell, LLM-bewertet.
What it does not measure
Fachliche Richtigkeit.
How the score is composed
Currency50
Discrimination50
Headroom50
Contamination safety9
Gaming resistance55
Methodological rigor45
Reproducibility60
Details
- Benchmaxxing risk
- medium
- Difficulty
- low
- Public test set
- yes
- Human-verified
- no
- LLM-judged
- yes
- Sample size
- 805
- Skala
- 0–100 % LC-Siege (höher besser)
Best measured models
No results ingested yet.