51Evidence
not yet measuredWhat it measures
Wie oft ein Modell harmlose Anfragen faelschlich verweigert. Die Gegenrichtung zu HarmBench, und im Alltag mindestens so wichtig.
What it does not measure
Echte Schadensrisiken.
How the score is composed
Currency50
Discrimination50
Headroom50
Contamination safety33
Gaming resistance100
Methodological rigor80
Reproducibility60
Details
- Benchmaxxing risk
- low
- Difficulty
- low
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- no
- Skala
- 0–100 % korrekt beantwortet (höher besser)
Hand-maintained — no automatic adapter yet.
Best measured models
No results ingested yet.