ModelRadar
/
← All benchmarks

CRUXEval

coding · code-reasoning↗ Quelle
48Evidence
not yet measured

What it measures

Ob ein Modell Ein- und Ausgabe von Code vorhersagen kann, ohne ihn auszufuehren.

What it does not measure

Codeerzeugung.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety33
Gaming resistance55
Methodological rigor70
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
medium
Public test set
yes
Human-verified
no
LLM-judged
no
Sample size
800
Skala
0–100 % correct (höher besser)

Best measured models

No results ingested yet.