ModelRadar
/
← All benchmarks

ZebraLogic

reasoning · constraint-logic↗ Quelle
53Evidence
not yet measured

What it measures

Logikraetsel mit harten Nebenbedingungen, beliebig skalierbare Schwierigkeit.

What it does not measure

Wissen, Sprache.

Aufgaben sind generierbar, daher kaum kontaminierbar.

How the score is composed

Currency50
Discrimination50
Headroom50
Contamination safety60
Gaming resistance55
Methodological rigor50
Reproducibility60

Details

Benchmaxxing risk
medium
Difficulty
high
Public test set
yes
Human-verified
no
LLM-judged
no
Skala
0–100 % solved (höher besser)

Hand-maintained — no automatic adapter yet.

Best measured models

No results ingested yet.