ModelRadar
/
← All benchmarks

Earnings-22

audio-stt · domain-specific↗ Quelle
68Evidence
active

What it measures

Erkennung in Analystenkonferenzen mit Fachbegriffen, Zahlen und Akzenten.

What it does not measure

Alltagssprache.

How the score is composed

Currency100
Discrimination26
Headroom10
Contamination safety100
Gaming resistance100
Methodological rigor80
Reproducibility60

Details

Benchmaxxing risk
low
Difficulty
high
Public test set
yes
Human-verified
yes
LLM-judged
no
Letzter Messwert
2026-09-28
Skala
0–100 % WER (niedriger besser)

Best measured models

  1. 1elevenlabs-scribe-v29.99 % WER
  2. 2assemblyai-universal-3-pro10.59 % WER
  3. 3speechmatics-enhanced10.75 % WER
  4. 4transcribe-03Cohere12.66 % WER
  5. 5whisper-v3OpenAI13.1 % WER
  6. 6whisper-v3OpenAI13.16 % WER
  7. 7parakeet-tdt-v2NVIDIA13.25 % WER
  8. 8parakeet-rnntNVIDIA14.14 % WER
  9. 9parakeet-ctcNVIDIA14.46 % WER
  10. 10parakeet-tdt-v3NVIDIA14.58 % WER