Negative-transfer case study Report artifact 2026-07-03

Qwen3-4B Multibackend Distilled

Adding multibackend teacher data did not create a broader agent. It preserved the saturated file-ops score while producing the worst breadth result in the recorded Qwen3-4B lineage. The public weights are useful as a reproducible warning: breadth gates must be frozen before specialist distillation.

Headline Numbers

File-ops depth

100% saturated narrow gate

Breadth

31% down from 59.6% stock

Breadth delta

-28.6pp recorded stock-to-candidate change

Decision

Reject failed negative-transfer artifact

Competitive Context

System Metric Score Size / Class Comparable? Readout
Multibackend-distilled 4B recorded depth / breadth 100% / 31% 4B Direct The target gate stayed saturated while breadth collapsed.
Stock Qwen3-4B same recorded depth / breadth family 58% / 59.6% 4B Direct The candidate gained depth but lost 28.6 breadth points.
File-ops-distilled 4B same recorded depth / breadth family 100% / 42.3% 4B Direct Already showed negative transfer; multibackend distillation made breadth worse.
ReST-fused 4B same recorded depth / breadth family 100% / 65% 4B Direct The later teacher-free ReST attempt is the breadth-preserving research candidate.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Attempts and decisions

AttemptMethodDepthBreadthDecision
A0Stock Qwen3-4B58%59.6%Baseline
A1File-ops distillation100%42.3%Route only
A2Multibackend distillation100%31%Reject
A3Teacher-free ReST100%65%Research-only routed ship

Release Blockers

Breadth collapsed

The recorded 31% breadth result is 28.6 points below stock and 11.3 points below the narrower file-ops distillation.

Unblock: Do not promote this checkpoint. Start from a breadth-preserving recipe with a frozen regression gate.

No committed machine-readable eval package

The exact historical result is preserved in the attempt ledger and session record, but raw predictions and a package-level eval report were not committed.

Unblock: Retain the result as historical evidence; only rerun if a new recipe needs the same checkpoint as a controlled baseline.

Evidence

Next Release Action

Keep the weights public as a failed comparison artifact. Do not spend compute revalidating them unless a new breadth-preserving recipe explicitly needs this checkpoint as its baseline.