ModelRadar
/
← All benchmarks

MLE-bench

agentic · ml-engineering↗ Quelle
45Evidence
not yet measured

What it measures

Ob ein Agent Kaggle-Wettbewerbe eigenstaendig bearbeitet.

What it does not measure

Interaktion, Erklaerbarkeit.

Kaggle-Loesungen stehen oeffentlich im Netz — Kontamination ist hier strukturell.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety9
Gaming resistance55
Methodological rigor50
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
high
Public test set
yes
Human-verified
no
LLM-judged
no
Skala
0–100 % medals (höher besser)

Hand-maintained — no automatic adapter yet.

Best measured models

No results ingested yet.