Benchmark archive

Every score should have a replay.

Protocol, raw decisions, local command, and the specialist beside its larger opponent. Successful and failed attempts both belong here when their evidence can be inspected and replayed.

Available now

3 evidence surfaces · 0 qualified ratings

Cross-game Arena

candidate · ratings unqualified

Head-to-head Arena Elo and paired-score evidence, with uncertainty and qualification shown before rank.

extensible adapters95% intervalsforfeit-awareno FIDE claim
Chess sample
4 games
Provisional gap
170 Arena Elo
Qualified
no

Character Chess

development gate passed

Can chess specialization compress frontier move selection into a 30–50M Mac-local model?

FEN + UCI20 tacticsfull gamesno tools
Frontier
65% exact
Strongest local
9B · 10%
Our SLM
not trained

Character 2048

failed · advantage not reproduced

Can a 30–50M specialist compress game intelligence from a frontier general LLM?

character inputstrict + constrainedno toolsfull replays
Best general LLM
Qwen3 4B Instruct
Mean score
845
Our SLM
not trained