ModelRadar
/
← All benchmarks

IFEval

instruction-following · verifiable-constraints↗ Quelle
45Evidence
not yet measured

What it measures

Ob maschinell pruefbare Vorgaben eingehalten werden — Wortzahl, Format, verbotene Woerter. Kein Geschmacksurteil, sondern eine Pruefung.

What it does not measure

Inhaltliche Qualitaet, Wahrheitsgehalt.

Objektiv pruefbar, deshalb belastbarer als LLM-bewertete Schreibvergleiche.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety33
Gaming resistance55
Methodological rigor70
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
low
Public test set
yes
Human-verified
no
LLM-judged
no
Sample size
541
Skala
0–100 % eingehalten (höher besser)

Best measured models

No results ingested yet.