Benchmark archive
Every score should have a replay.
Protocol, raw decisions, local command, and the specialist beside its larger opponent. Successful and failed attempts both belong here when their evidence can be inspected and replayed.
Available now
3 evidence surfaces · 0 qualified ratingsCross-game Arena
candidate · ratings unqualifiedHead-to-head Arena Elo and paired-score evidence, with uncertainty and qualification shown before rank.
- Chess sample
- 4 games
- Provisional gap
- 170 Arena Elo
- Qualified
- no
Character Chess
development gate passedCan chess specialization compress frontier move selection into a 30–50M Mac-local model?
- Frontier
- 65% exact
- Strongest local
- 9B · 10%
- Our SLM
- not trained
Character 2048
failed · advantage not reproducedCan a 30–50M specialist compress game intelligence from a frontier general LLM?
- Best general LLM
- Qwen3 4B Instruct
- Mean score
- 845
- Our SLM
- not trained