63Evidence
not yet measuredWhat it measures
Menschlicher Blindvergleich zweier Sprachsynthesen, Elo-Wertung.
What it does not measure
Latenz, Stimmklonen, Sprachabdeckung.
How the score is composed
Currency50
Discrimination50
Headroom50
Contamination safety100
Gaming resistance100
Methodological rigor80
Reproducibility20
Details
- Benchmaxxing risk
- low
- Difficulty
- medium
- Public test set
- no
- Human-verified
- yes
- LLM-judged
- no
- Skala
- 800–1600 Elo (höher besser)
Hand-maintained — no automatic adapter yet.
Best measured models
No results ingested yet.