68Evidence
activeWhat it measures
Erkennung in Analystenkonferenzen mit Fachbegriffen, Zahlen und Akzenten.
What it does not measure
Alltagssprache.
How the score is composed
Currency100
Discrimination26
Headroom10
Contamination safety100
Gaming resistance100
Methodological rigor80
Reproducibility60
Details
- Benchmaxxing risk
- low
- Difficulty
- high
- Public test set
- yes
- Human-verified
- yes
- LLM-judged
- no
- Letzter Messwert
- 2026-09-28
- Skala
- 0–100 % WER (niedriger besser)
Best measured models
- 1elevenlabs-scribe-v29.99 % WER
- 2assemblyai-universal-3-pro10.59 % WER
- 3speechmatics-enhanced10.75 % WER
- 4transcribe-03Cohere12.66 % WER
- 5whisper-v3OpenAI13.1 % WER
- 6whisper-v3OpenAI13.16 % WER
- 7parakeet-tdt-v2NVIDIA13.25 % WER
- 8parakeet-rnntNVIDIA14.14 % WER
- 9parakeet-ctcNVIDIA14.46 % WER
- 10parakeet-tdt-v3NVIDIA14.58 % WER