ModelRadar
/
← All benchmarks

AlpacaEval 2.0

instruction-following · llm-judged-preference↗ Quelle
39Evidence
not yet measured

What it measures

Laengenbereinigte Siegquote gegen ein Referenzmodell, LLM-bewertet.

What it does not measure

Fachliche Richtigkeit.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety9
Gaming resistance55
Methodological rigor45
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
low
Public test set
yes
Human-verified
no
LLM-judged
yes
Sample size
805
Skala
0–100 % LC-Siege (höher besser)

Best measured models

No results ingested yet.