ModelRadar
/
← All benchmarks

XSTest

safety · over-refusal↗ Quelle
51Evidence
not yet measured

What it measures

Wie oft ein Modell harmlose Anfragen faelschlich verweigert. Die Gegenrichtung zu HarmBench, und im Alltag mindestens so wichtig.

What it does not measure

Echte Schadensrisiken.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety33
Gaming resistance100
Methodological rigor80
Reproducibility60

Details

Benchmaxxing risk
low
Difficulty
low
Public test set
yes
Human-verified
yes
LLM-judged
no
Skala
0–100 % korrekt beantwortet (höher besser)

Hand-maintained — no automatic adapter yet.

Best measured models

No results ingested yet.