ModelRadar
/
← All benchmarks

BIG-Bench Hard

reasoning · mixed↗ Quelle
40Evidence
not yet measured

What it measures

23 Teilaufgaben, bei denen Modelle frueher schlechter waren als Menschen.

What it does not measure

Aktuelle Grenzfaehigkeiten.

Von aktuellen Modellen weitgehend ausgereizt.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety9
Gaming resistance55
Methodological rigor50
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
low
Public test set
yes
Human-verified
no
LLM-judged
no
Skala
0–100 % correct (höher besser)

Best measured models

No results ingested yet.