Arena Elo
Only two interacting policies, win/draw/loss outcomes, one frozen pool, and match-level uncertainty.
Candidate arena · 2 rating families · 0 hidden conversions
Head-to-head games earn Arena Elo. Independent games earn paired score evidence. Neither becomes a universal intelligence number.
Rating console
The 170-point gap is a regularized diagnostic from four games—not a qualified rank. Two 4B wins came from 9B notation forfeits. This pool-specific Arena Elo is not FIDE, human, engine, Chess.com, or Lichess Elo.
Paired delta compares each model with random legal play on the same seed. The interval crossing zero means no demonstrated advantage.
Evidence ledger
Only two interacting policies, win/draw/loss outcomes, one frozen pool, and match-level uncertainty.
Single-player policies share seeds with a baseline. Mean score and uncertainty stay in the game’s own units.
The math can produce a diagnostic before the sample can support a public rank. We show both states.
Extension contract
Choose a family. Head-to-head result or independent paired score.
Normalize evidence. Participant identity, outcomes, failures, trace hashes, and source paths.
Freeze qualification. Sample, balance, connectivity, and failure limits before scoring.
Reuse the report. The validator, uncertainty engine, and dashboard remain unchanged.