Teacher-free breadth recovery Release-ready weights 2026-07-13

Qwen3-4B ReST Fused

This is the factory's first narrow ship decision from an existing measured candidate: the public weights, frozen fixtures, package, and routing boundary are preserved, while missing historical performance and trace evidence stays visible.

Headline Numbers

File-ops hard gate

100% held from the distilled depth anchor

Breadth after ReST

65% up from 59.6% stock 4B

Breadth delta

+5.4pp 52 held-out non-file-ops tasks

Paid API cost

$0 teacher-free local ReST iteration

Competitive Context

System Metric Score Size / Class Comparable? Readout
posttrainllm Qwen3-4B ReST out-of-domain breadth 65% 4B, 8.06GB stored Direct Recorded on the same 52-task breadth fixture and prompt family as stock.
Stock Qwen3-4B same out-of-domain breadth fixture 59.6% 4B Direct The ReST iteration recovers breadth without giving up the file-ops depth gate.
File-ops-only distilled 4B same out-of-domain breadth fixture 42.3% 4B Direct Shows the negative transfer that the ReST iteration was designed to recover.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Recorded result

GateStockReST candidateReadout
File-ops hard gate0.581.00Depth preserved
Out-of-domain breadth0.5960.65+5.4 points over stock
Latency / RAM / tok-snot preservednot preservedNo estimated values

Release Blockers

Historical performance evidence missing

The original run did not preserve latency, RAM, tok-s, elapsed time, or raw predictions.

Unblock: Run a fresh product-specific gate only when a downstream integration justifies the heavy model work.

Not a Pace planner

Pace uses a different intent envelope and its own six-dimension ship gate.

Unblock: Re-distill on Pace's action surface and clear the Pace gate before runtime wiring.

Fine-Tune Report Card

The portable before/after proof for this artifact: baseline, candidate, deltas, regressions, slices, performance, eval validity, and the ship/retry/reject decision — with every value labeled measured, derived, historical, skipped, or not recorded. Compiled offline from the recorded factory evidence; no model was run to produce it.

Shipped — routed only Not fully verified — see the report card for why

Evidence

Next Release Action

Keep this package research-only. Freeze a product-specific target before spending compute on another eval or training run.