ModelRadar
/
← All benchmarks

BigCodeBench

coding · library-use↗ Quelle
66Evidence
active

What it measures

Aufgaben, die den korrekten Einsatz echter Bibliotheken verlangen, nicht nur Algorithmik.

What it does not measure

Mehrdateien-Aenderungen, Projektkontext.

How the score is composed

Currency95
Discrimination0
Headroom100
Contamination safety55
Gaming resistance100
Methodological rigor100
Reproducibility60

Details

Benchmaxxing risk
low
Difficulty
medium
Public test set
yes
Human-verified
yes
LLM-judged
no
Sample size
1140
Letzter Messwert
2026-09-10
Skala
0–100 % pass@1 (höher besser)

Best measured models

  1. 1DeepSeek V4 ProDeepSeekvendor number56.8 % pass@1
  2. 2DeepSeek V4.1 FlashDeepSeekvendor number56.8 % pass@1