ModelRadar
/
← All benchmarks

HumanEval

coding · function-synthesis↗ Quelle
57Evidence
saturated

What it measures

164 kleine Python-Funktionen aus einer Docstring-Beschreibung.

What it does not measure

Alles, was ueber eine einzelne kurze Funktion hinausgeht.

Historisch wichtig, heute praktisch ausgereizt und stark kontaminiert. Wird hier gefuehrt, damit alte Herstellerangaben einordbar bleiben, nicht weil er noch etwas unterscheidet.

How the score is composed

Currency95
Discrimination100
Headroom30
Contamination safety15
Gaming resistance55
Methodological rigor80
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
low
Public test set
yes
Human-verified
yes
LLM-judged
no
Sample size
164
Letzter Messwert
2026-09-10
Skala
0–100 % pass@1 (höher besser)

Best measured models

  1. 1edge0-edge0-35b-a3b-previewvendor number90.9 % pass@1
  2. 2Motif-3Motif Technologiesvendor number73.7 % pass@1
  3. 3DeepSeek V4.1 FlashDeepSeekvendor number69.5 % pass@1
  4. 4Nemotron 3.5 LightningNVIDIAvendor number66.46 % pass@1
  5. 5DeepSeek V4 ProDeepSeekvendor number62.8 % pass@1