ModelRadar
/
← All benchmarks

SimpleQA

knowledge · factuality↗ Quelle
81Evidence
active

What it measures

Kurze Faktenfragen mit eindeutiger Antwort. Misst vor allem, ob ein Modell zugibt, etwas nicht zu wissen, statt zu erfinden.

What it does not measure

Schliessen, Rechnen, laengere Texte.

How the score is composed

Currency95
Discrimination100
Headroom100
Contamination safety55
Gaming resistance55
Methodological rigor75
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
high
Public test set
yes
Human-verified
yes
LLM-judged
yes
Sample size
4326
Letzter Messwert
2026-09-10
Skala
0–100 % correct (höher besser)

Best measured models

  1. 1DeepSeek V4 ProDeepSeekvendor number46.2 % correct
  2. 2DeepSeek V4.1 FlashDeepSeekvendor number30.1 % correct