ModelRadar
/
← All benchmarks

BEIR

retrieval · zero-shot-retrieval↗ Quelle
42Evidence
not yet measured

What it measures

Retrieval ohne aufgabenspezifisches Training ueber 18 Datensaetze.

What it does not measure

Reranking-Ketten, Latenz.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety9
Gaming resistance55
Methodological rigor50
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
medium
Public test set
yes
Human-verified
no
LLM-judged
no
Skala
0–100 nDCG@10 (höher besser)

Best measured models

No results ingested yet.