49.5M on-device classifier Release-ready weights 2026-07-18

Pace Intent Router v8

This is the most important rejection in the published model set. Pace v8 scored 95.5% on 14,995 source-matched synthetic examples at roughly 3ms, but only 57.1% on a new 63-instance sealed suite that frontier aced. The weights remain useful as a latency floor and as proof that the synthetic generator overfit its template families—not as the production winner.

Headline Numbers

Sealed accuracy

57.1% 63 held-out instances, two passes

Frontier retained

57.1% Codex gpt-5.5 scored 100% on the same ruler

Warm latency

3.8ms mean local sealed-eval latency

Source-matched holdout

95.5% 14,995 synthetic rows; did not generalize

Competitive Context

System Metric Score Size / Class Comparable? Readout
Pace Intent Router v8 sealed exact intent accuracy 57.1% 49.5M Direct Fastest local entry, but rejected because fresh phrasing exposed distribution overfit.
Qwen3-4B-Instruct-2507 4-bit same sealed exact intent accuracy 93.7% 4B Direct Capability winner among local entries at 211ms mean warm latency.
Apple FoundationModels same sealed exact intent accuracy 92.1% on-device system model Direct 522ms mean warm latency; much slower than Pace v8 but materially more accurate on fresh phrasing.
Codex gpt-5.5 same sealed exact intent accuracy 100% frontier anchor Direct Passed the frozen 99% frontier-ceiling gate. Its amortized batch latency is not comparable with per-instance local timing.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Attempts and decisions

Evidence stageAccuracyLatencyDecisionLesson
Source-matched synthetic holdout95.5%3.1ms p50promisingThe pipeline learned its generator distribution
Sealed V157.1%3.8ms meanreject-production-winnerFresh user-like phrasing broke the apparent win

Sealed V1 result

EntryAccuracyUnknown recallWarm latencyReadout
Codex gpt-5.5100.0%100.0%519ms amortized batchFrontier ruler passed
Qwen3 4B 4-bit93.7%77.8%211msLocal capability winner
Apple FoundationModels92.1%77.8%522msAccurate, slower
Pace v857.1%55.6%3.8msLatency floor; rejected

Release Blockers

Fresh-distribution generalization failed

The earlier 95.5% result was source-matched to the synthetic generator. Accuracy fell to 57.1% on leakage-checked sealed V1.

Unblock: Generate public-development examples from failure themes—not sealed prompts—and require a newly generated sealed V2 for any successor claim.

Rejected as the production winner

Qwen3 4B and Apple FoundationModels both exceeded 92% on the same sealed suite.

Unblock: Keep v8 only as a latency floor and training-pipeline artifact until a successor clears the sealed gate.

Evidence

Next Release Action

Do not tune on sealed V1. Build a broader public-development corpus from the failure themes, train a successor, and judge it once on a newly generated sealed V2.