Development Gate 0 · 20 novel positions · no tools

The gradient
is real.

Can chess specialization compress frontier move selection into a 30–50M Mac-local model? Puzzles are the ruler. Complete games show what those scores look like when policies must keep choosing.

Gate 0 passedEligible for a frozen suite—not training yet.
Frontier
65%
9B local
10%
4B local
0%
Random calibrated
6.74%

Watch every decision.

Puzzle 1 / 20
Turn

Tactical move selection

Stockfish 18 labels · depth 12

Swipe the table to inspect latency and interpretation →

PolicyRoleExactLegalMean decisionInterpretation
Codex gpt-5.5frontier13 / 20 · 65%100%23.69 sGate anchor passed
Qwen3.5 9Blocal general2 / 20 · 10%90%577 msAbove 4B, far below frontier
Qwen3 4Blocal general0 / 20 · 0%95%294 msLegal executor, no tactical hits
Random legalcalibrated diagnostic6.74% expected100%<1 ms2,000-seed mean
Our chess specialist30–50M · ≤50MspecialistNot trained—freeze the benchmark first

Why this passes: frontier exceeded calibrated random by 58.26 points and the strongest local model by 55 points, while returning a legal move on every position. The ruler is promising; it is not frozen benchmark evidence yet.

More data exposed the ceiling problem.

100 generated · 86 admitted · 40 reviewed

Different coverage is shown explicitly; do not rank a 10-position row against a 40-position row →

PolicyCoverageExactRaw legalExecuted legalRedirectMean decision
GPT-5.5codex-cli 40 / 40 28 / 40 · 70% 100% 100% 0% 18.94 s
GPT-5.5high reasoning 40 / 40 24 / 40 · 60% 97.5% 100% 2.5% 26.92 s
GPT-5.4codex-cli 40 / 40 30 / 40 · 75% 97.5% 100% 2.5% 44.70 s
GPT-5.4-minicodex-cli 10 / 40 8 / 10 · 80% 100% 100% 0% 70.47 s
Claude Sonnet 5claude-cli 20 / 40 13 / 20 · 65% 90% 100% 10% 30.09 s
Claude Opus 4.8claude-cli 10 / 40 7 / 10 · 70% 70% 100% 30% 86.80 s
Devin GLM-5.2devin-cli 20 / 40 14 / 20 · 70% 100% 100% 0% not exposed

Why this is not frozen: Stockfish's top move was stable at depths 16 and 20 on all 100 candidates, but no frontier lane reached the required near-100% ceiling. GPT-5.5 high reasoning scored lower than medium reasoning. The calibrated random mean was 6.7%. Human audit comes next; small-model grading and training remain blocked.

A match score can lie.

Qwen 4B recorded two wins and two draws against Qwen 9B—but both wins were notation forfeits, and both draws were 14-ply repetitions. That is why the dense puzzle ruler decides admission while complete games remain inspectable behavior.

Recorded games
4
4B wins
2
Draws
2
9B wins
0
Notation forfeits
2

Same board. Same legal set. No engine.

Same FEN, sorted legal UCI set, greedy decoding, eight output tokens maximum, no engine/tools/search/code/rollouts.

python3.12 scripts/chess_mlx_pilot.py --suite evals/chess/fixtures/development-puzzles-v1.json --model <MODEL_PATH> --model-ref <MODEL_REF> --policy-id <ID> --output runs/chess/<ID>.json
ObservationFEN + ply + sorted legal UCI moves
Strict outputOne legal UCI move, nothing else
Rulespython-chess 1.999
LabelsStockfish 18, depth 12

What this does not prove