Development Gate 0 · 20 novel positions · no tools
The gradient
is real.
Can chess specialization compress frontier move selection into a 30–50M Mac-local model? Puzzles are the ruler. Complete games show what those scores look like when policies must keep choosing.
- Frontier
- 65%
- 9B local
- 10%
- 4B local
- 0%
- Random calibrated
- 6.74%
Prerecorded evidence
Watch every decision.
Primary ruler
Tactical move selection
Swipe the table to inspect latency and interpretation →
| Policy | Role | Exact | Legal | Mean decision | Interpretation |
|---|---|---|---|---|---|
| Codex gpt-5.5 | frontier | 13 / 20 · 65% | 100% | 23.69 s | Gate anchor passed |
| Qwen3.5 9B | local general | 2 / 20 · 10% | 90% | 577 ms | Above 4B, far below frontier |
| Qwen3 4B | local general | 0 / 20 · 0% | 95% | 294 ms | Legal executor, no tactical hits |
| Random legal | calibrated diagnostic | 6.74% expected | 100% | <1 ms | 2,000-seed mean |
| Our chess specialist30–50M · ≤50M | specialist | Not trained—freeze the benchmark first | |||
Why this passes: frontier exceeded calibrated random by 58.26 points and the strongest local model by 55 points, while returning a legal move on every position. The ruler is promising; it is not frozen benchmark evidence yet.
Expanded candidate audit
More data exposed the ceiling problem.
Different coverage is shown explicitly; do not rank a 10-position row against a 40-position row →
| Policy | Coverage | Exact | Raw legal | Executed legal | Redirect | Mean decision |
|---|---|---|---|---|---|---|
| GPT-5.5codex-cli | 40 / 40 | 28 / 40 · 70% | 100% | 100% | 0% | 18.94 s |
| GPT-5.5high reasoning | 40 / 40 | 24 / 40 · 60% | 97.5% | 100% | 2.5% | 26.92 s |
| GPT-5.4codex-cli | 40 / 40 | 30 / 40 · 75% | 97.5% | 100% | 2.5% | 44.70 s |
| GPT-5.4-minicodex-cli | 10 / 40 | 8 / 10 · 80% | 100% | 100% | 0% | 70.47 s |
| Claude Sonnet 5claude-cli | 20 / 40 | 13 / 20 · 65% | 90% | 100% | 10% | 30.09 s |
| Claude Opus 4.8claude-cli | 10 / 40 | 7 / 10 · 70% | 70% | 100% | 30% | 86.80 s |
| Devin GLM-5.2devin-cli | 20 / 40 | 14 / 20 · 70% | 100% | 100% | 0% | not exposed |
Why this is not frozen: Stockfish's top move was stable at depths 16 and 20 on all 100 candidates, but no frontier lane reached the required near-100% ceiling. GPT-5.5 high reasoning scored lower than medium reasoning. The calibrated random mean was 6.7%. Human audit comes next; small-model grading and training remain blocked.
Secondary demonstration
A match score can lie.
Qwen 4B recorded two wins and two draws against Qwen 9B—but both wins were notation forfeits, and both draws were 14-ply repetitions. That is why the dense puzzle ruler decides admission while complete games remain inspectable behavior.
- Recorded games
- 4
- 4B wins
- 2
- Draws
- 2
- 9B wins
- 0
- Notation forfeits
- 2
Reproduce locally
Same board. Same legal set. No engine.
Same FEN, sorted legal UCI set, greedy decoding, eight output tokens maximum, no engine/tools/search/code/rollouts.
python3.12 scripts/chess_mlx_pilot.py --suite evals/chess/fixtures/development-puzzles-v1.json --model <MODEL_PATH> --model-ref <MODEL_REF> --policy-id <ID> --output runs/chess/<ID>.jsonWhat this does not prove
- This is a 20-position development gate, not the future frozen benchmark.
- The frontier identity is a mutable Codex model alias and cannot support a frozen public ceiling claim yet.
- Engine-generated tactical-gap positions test move selection but do not yet provide broad human-authored theme coverage.
- The paired local games are demonstrative only: notation forfeits and short repetitions make them unsuitable for Elo or capability ranking.
- No 30–50M specialist has been trained or evaluated.