55Evidence
not yet measuredWhat it measures
Ob ein Agent mehrstufige Kundenvorgaenge korrekt abschliesst und dabei Regeln einhaelt, statt nur plausibel zu antworten.
What it does not measure
Codearbeit, Wissen, Kreativitaet.
Misst Regeltreue unter Druck, was in der Praxis oft wichtiger ist als Rohleistung.
How the score is composed
Currency50
Discrimination50
Headroom50
Contamination safety60
Gaming resistance55
Methodological rigor50
Reproducibility100
Details
- Benchmaxxing risk
- medium
- Difficulty
- high
- Public test set
- yes
- Human-verified
- no
- LLM-judged
- no
- Skala
- 0–100 % pass^1 (höher besser)
Best measured models
No results ingested yet.