ModelRadar
/

Benchmarks

Not just the largest pool, but a rated one. Every benchmark is scored on how current, discriminative and gameable it is — and that score is the weight it carries in every ranking here.

74 Benchmarks18 domains431 Results

coding (12)

BenchmarkEvidenceStatusBenchmaxxing riskResults
LiveCodeBenchcompetitive-programming80activelow4
SWE-bench Verifiedrepo-level-bugfix71activemedium6
SWE-bench Prorepo-level-bugfix69not yet measuredlow0
SWE-Lancereconomic-value67not yet measuredlow0
BigCodeBenchlibrary-use66activelow2
Terminal-Benchshell-agent65not yet measuredlow0
SWE-bench Multilingualrepo-level-bugfix58not yet measuredlow0
HumanEvalfunction-synthesis57saturatedmedium5
CRUXEvalcode-reasoning48not yet measuredmedium0
RepoBenchrepo-completion46not yet measuredmedium0
Aider Polyglotedit-format42outdatedmedium0
MBPP+function-synthesis40not yet measuredmedium0

vision (7)

BenchmarkEvidenceStatusBenchmaxxing riskResults
MMMU-Proexpert-multimodal80activemedium4
DocVQAdocument-understanding71saturatedlow3
BLINKvisual-perception63not yet measuredlow0
MathVistavisual-math53not yet measuredlow1
RealWorldQAspatial-understanding44not yet measuredmedium0
AI2Ddiagram-understanding44activemedium2
ChartQAchart-reading41activemedium2

audio-stt (7)

BenchmarkEvidenceStatusBenchmaxxing riskResults
Earnings-22domain-specific68activelow21
AMI Meeting Corpusmeetings63not yet measuredlow0
Open ASR Leaderboardaggregate-wer58activelow21
Common Voice (ASR)crowdsourced-accents52saturatedlow127
FLEURSmultilingual-asr52saturatedlow159
TED-LIUM 3prepared-speech49saturatedlow21
LibriSpeech test-otherread-speech48not yet measuredlow0

agentic (6)

BenchmarkEvidenceStatusBenchmaxxing riskResults
GAIAgeneral-assistant67not yet measuredlow0
BrowseCompweb-research67not yet measuredlow0
OSWorldcomputer-use63not yet measuredlow0
TAU-benchcustomer-workflows55not yet measuredmedium0
WebArenaweb-navigation49not yet measuredmedium0
MLE-benchml-engineering45not yet measuredmedium0

reasoning (6)

BenchmarkEvidenceStatusBenchmaxxing riskResults
ARC-AGI-2abstract-reasoning71not yet measuredlow0
Humanity's Last Examexpert-frontier67not yet measuredlow0
ZebraLogicconstraint-logic53not yet measuredmedium0
GPQA Diamondexpert-science50saturatedhigh14
MuSRmulti-step-narrative46not yet measuredmedium0
BIG-Bench Hardmixed40not yet measuredmedium0

safety (5)

BenchmarkEvidenceStatusBenchmaxxing riskResults
JailbreakBenchadaptive-attacks63not yet measuredlow0
HarmBenchjailbreak-robustness55not yet measuredlow0
AgentHarmagentic-misuse53not yet measuredlow0
XSTestover-refusal51not yet measuredlow0
TruthfulQAfactuality41not yet measuredmedium0

long-context (5)

BenchmarkEvidenceStatusBenchmaxxing riskResults
LongBench v2realistic-documents90activelow3
RULERsynthetic-retrieval68saturatedmedium2
MRCRmulti-round-coreference53not yet measuredmedium0
BABILongreasoning-in-haystack53not yet measuredmedium0
Needle in a Haystackretrieval46not yet measuredmedium0

instruction-following (5)

BenchmarkEvidenceStatusBenchmaxxing riskResults
LMArenahuman-preference63not yet measuredlow0
IFEvalverifiable-constraints45not yet measuredmedium0
Arena-Hardllm-judged-preference43not yet measuredmedium0
WildBenchreal-user-tasks43not yet measuredmedium0
AlpacaEval 2.0llm-judged-preference39not yet measuredmedium0

math (4)

BenchmarkEvidenceStatusBenchmaxxing riskResults
FrontierMathresearch-level67not yet measuredlow0
HMMTcompetition61saturatedmedium6
AIMEcompetition53saturatedmedium10
MATH-500school-competition35saturatedmedium1

multilingual (4)

BenchmarkEvidenceStatusBenchmaxxing riskResults
INCLUDEregional-knowledge57not yet measuredlow0
Global-MMLUknowledge49not yet measuredmedium0
Belebelereading-comprehension49not yet measuredmedium0
MGSMmath-transfer44activemedium2

knowledge (3)

BenchmarkEvidenceStatusBenchmaxxing riskResults
SimpleQAfactuality81activemedium2
MMLU-Probroad-academic50not yet measuredmedium0
MMLUbroad-academic43supersededhigh11

efficiency (3)

BenchmarkEvidenceStatusBenchmaxxing riskResults
Ausgabegeschwindigkeit (Token/s)latency56not yet measuredlow0
Zeit bis zum ersten Tokenlatency56not yet measuredlow0
Artificial Analysis Intelligence Indexaggregate-index53not yet measuredlow0

retrieval (2)

BenchmarkEvidenceStatusBenchmaxxing riskResults
AIR-Benchfresh-retrieval62not yet measuredlow0
BEIRzero-shot-retrieval42not yet measuredmedium0

tool-use (1)

BenchmarkEvidenceStatusBenchmaxxing riskResults
Berkeley Function Calling Leaderboardfunction-calling55not yet measuredlow0

translation (1)

BenchmarkEvidenceStatusBenchmaxxing riskResults
FLORES-200machine-translation55not yet measuredlow0

multimodal (1)

BenchmarkEvidenceStatusBenchmaxxing riskResults
Video-MMEvideo-understanding81activelow2

audio-tts (1)

BenchmarkEvidenceStatusBenchmaxxing riskResults
TTS Arenahuman-preference63not yet measuredlow0

embedding (1)

BenchmarkEvidenceStatusBenchmaxxing riskResults
MTEBaggregate42not yet measuredmedium0