ModelRadar
/
← All benchmarks

LongBench v2

long-context · realistic-documents↗ Quelle
90Evidence
active

What it measures

Verstehen und Schliessen ueber echte lange Dokumente statt synthetischer Nadeln.

What it does not measure

Sehr kurze Aufgaben, Werkzeugnutzung.

How the score is composed

Currency95
Discrimination100
Headroom100
Contamination safety55
Gaming resistance100
Methodological rigor100
Reproducibility60

Details

Benchmaxxing risk
low
Difficulty
high
Public test set
yes
Human-verified
yes
LLM-judged
no
Sample size
503
Letzter Messwert
2026-09-10
Skala
0–100 % correct (höher besser)

Best measured models

  1. 1Nemotron 3 UltraNVIDIAvendor number61.9 % correct
  2. 2DeepSeek V4.1 FlashDeepSeekvendor number44.7 % correct
  3. 3DeepSeek V4 ProDeepSeekvendor number40.2 % correct