63Evidence
not yet measuredWhat it measures
Robustheit gegen anpassungsfaehige Angriffe, mit versionierten Artefakten.
What it does not measure
Alltagsnutzen.
How the score is composed
Currency50
Discrimination50
Headroom50
Contamination safety60
Gaming resistance100
Methodological rigor80
Reproducibility60
Details
- Benchmaxxing risk
- low
- Difficulty
- high
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- no
- Skala
- 0–100 % attack success (niedriger besser)
Hand-maintained — no automatic adapter yet.
Best measured models
No results ingested yet.