ModelRadar
/
← All benchmarks

ChartQA

vision · chart-reading↗ Quelle
41Evidence
active

What it measures

Ablesen und Rechnen an Diagrammen.

What it does not measure

Freie Bildbeschreibung, Video.

How the score is composed

Currency82
Discrimination18
Headroom53
Contamination safety15
Gaming resistance55
Methodological rigor50
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
low
Public test set
yes
Human-verified
no
LLM-judged
no
Letzter Messwert
2026-07-25
Skala
0–100 % correct (höher besser)

Best measured models

  1. 1Mage-VLMicrosoftvendor number83.96 % correct
  2. 2Phi-4-multimodalMicrosoftvendor number81.8 % correct