Sealed accuracy
57.1% 63 held-out instances, two passesPace Intent Router v8
This is the most important rejection in the published model set. Pace v8 scored 95.5% on 14,995 source-matched synthetic examples at roughly 3ms, but only 57.1% on a new 63-instance sealed suite that frontier aced. The weights remain useful as a latency floor and as proof that the synthetic generator overfit its template families—not as the production winner.
Headline Numbers
Frontier retained
57.1% Codex gpt-5.5 scored 100% on the same rulerWarm latency
3.8ms mean local sealed-eval latencySource-matched holdout
95.5% 14,995 synthetic rows; did not generalizeCompetitive Context
| System | Metric | Score | Size / Class | Comparable? | Readout |
|---|---|---|---|---|---|
| Pace Intent Router v8 | sealed exact intent accuracy | 57.1% | 49.5M | Direct | Fastest local entry, but rejected because fresh phrasing exposed distribution overfit. |
| Qwen3-4B-Instruct-2507 4-bit | same sealed exact intent accuracy | 93.7% | 4B | Direct | Capability winner among local entries at 211ms mean warm latency. |
| Apple FoundationModels | same sealed exact intent accuracy | 92.1% | on-device system model | Direct | 522ms mean warm latency; much slower than Pace v8 but materially more accurate on fresh phrasing. |
| Codex gpt-5.5 | same sealed exact intent accuracy | 100% | frontier anchor | Direct | Passed the frozen 99% frontier-ceiling gate. Its amortized batch latency is not comparable with per-instance local timing. |
Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.
Attempts and decisions
| Evidence stage | Accuracy | Latency | Decision | Lesson |
|---|---|---|---|---|
| Source-matched synthetic holdout | 95.5% | 3.1ms p50 | promising | The pipeline learned its generator distribution |
| Sealed V1 | 57.1% | 3.8ms mean | reject-production-winner | Fresh user-like phrasing broke the apparent win |
Sealed V1 result
| Entry | Accuracy | Unknown recall | Warm latency | Readout |
|---|---|---|---|---|
| Codex gpt-5.5 | 100.0% | 100.0% | 519ms amortized batch | Frontier ruler passed |
| Qwen3 4B 4-bit | 93.7% | 77.8% | 211ms | Local capability winner |
| Apple FoundationModels | 92.1% | 77.8% | 522ms | Accurate, slower |
| Pace v8 | 57.1% | 55.6% | 3.8ms | Latency floor; rejected |
Release Blockers
Fresh-distribution generalization failed
The earlier 95.5% result was source-matched to the synthetic generator. Accuracy fell to 57.1% on leakage-checked sealed V1.
Unblock: Generate public-development examples from failure themes—not sealed prompts—and require a newly generated sealed V2 for any successor claim.
Rejected as the production winner
Qwen3 4B and Apple FoundationModels both exceeded 92% on the same sealed suite.
Unblock: Keep v8 only as a latency floor and training-pipeline artifact until a successor clears the sealed gate.