posttrainllm fine-tune report card · schema v1 · compiler 1.0.0
Qwen3-4B File-Ops Distilled
— not recordedThe specialist package format does not record the owner goal that framed the run.
- Target
- Qwen3-4B File-Ops Distilled
- Base model
- Qwen/Qwen3-4B-Instruct-2507 (bf16)
- Candidate
- qwen3-4b-file-ops-distilled
- Method
- frontier/gold trajectory distillation on GorillaFileSystem multi-turn tasks
- Compiled from
specialists/qwen3-4b-file-ops-distilled(specialist-package)
Decision
Shipped Shipped for a named route only
ship only as a routed file-ops specialist; do not use as the general planner
Routing constraint
ship only as a routed file-ops specialist; do not use as the general planner. Do not use for: general Pace planner; multi-domain agentic planning without routing.
This candidate is safe only inside that envelope. It is not a general replacement for the base model.
Verification status
Not fully verified. This report does not claim a verified ship. Reasons:
- Primary gate `file_ops_hard_gate` baseline is `historical`, not a current measurement.
- Primary gate `file_ops_hard_gate` candidate is `historical`, not a current measurement.
- Primary gate `file_ops_hard_gate` has no threshold value (state `missing`).
- Primary gate `file_ops_hard_gate` has no passed value (state `missing`).
- Frozen-eval identity is not recorded as a current measurement.
- Train/eval overlap (leakage) was not checked with current evidence.
- Failure reason
- Not applicable
- Failure-reason confidence
- not-applicable
- Lesson
- Not recorded
Next action
Before and after
This candidate does not present an unconditional win: out_of_domain_breadth failed. Target and regression gates are reported independently below.
| Gate | Role | Metric | Baseline | Candidate | Delta | Threshold | Result | n | Frontier ceiling |
|---|---|---|---|---|---|---|---|---|---|
| file_ops_hard_gate | primary | file_ops_hard_gate | 0.5800 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 1.0000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.4200 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | — not recordedThe specialist package records no ship threshold for the primary gate, so a pass/fail result cannot be derived. | 12 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 1.0000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). |
| file_ops_hardgen_heldout | regression | file_ops_hardgen_heldout | 0.6000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.9500 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.3500 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | yes derivedNo threshold was recorded. Derived as passing because the candidate did not score below the baseline. | 40 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | — not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability. |
| out_of_domain_breadth | breadth | out_of_domain_breadth | 0.5960 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.4230 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | -0.1730 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | no derivedNo threshold was recorded. Derived as failing because the candidate scored below the baseline on a non-primary gate. | 52 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | — not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability. |
Eval identity
| Gate | Suite | Command | Date | Frozen |
|---|---|---|---|---|
| file_ops_hard_gate | file_ops_hard_gate | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-06-19 | — not recordedThe package does not record whether this suite is frozen. |
| file_ops_hardgen_heldout | file_ops_hardgen_heldout | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-06-19 | — not recordedThe package does not record whether this suite is frozen. |
| out_of_domain_breadth | out_of_domain_breadth | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-06-19 | — not recordedThe package does not record whether this suite is frozen. |
Per-slice evidence
No slice evidence was recorded for this candidate.
Cost and performance
| Metric | Value | Source |
|---|---|---|
| Latency | — not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. | specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.latency_ms |
| RAM / peak RSS | — not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. | specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.peak_rss_mb |
| Decode throughput | — not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. | specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.tokens_per_second |
| Training time | — not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. | specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.training_time_seconds |
| Training cost | — not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. | specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.training_cost_usd |
| Eval time | — not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. | specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.eval_time_seconds |
Eval validity and leakage
| Check | Result | Source |
|---|---|---|
| Frontier ceiling | 1.0000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-file-ops-distilled/eval_report.json#scores[0].frontier |
| Frozen eval | — not recordedThe package records no frozen held-out split identity. | specialists/qwen3-4b-file-ops-distilled/eval_report.json |
| Train/eval overlap | — not recordedNo train/eval overlap check is recorded for this package. | specialists/qwen3-4b-file-ops-distilled/eval_report.json |
Known eval limitations
- Gate `file_ops_hardgen_heldout` has no frontier-ceiling evidence, so its absolute score is not calibrated against frontier capability.
- Gate `out_of_domain_breadth` has no frontier-ceiling evidence, so its absolute score is not calibrated against frontier capability.
- Values come from a committed specialist package, not a canonical factory-run folder: eval commands, dataset hashes, and raw predictions are unavailable for independent replay.
Caveats
- The out-of-domain breadth ceiling is not frontier-validated yet.
- The negative-transfer result is relative and directly comparable: same 52 tasks, same prompt, stock vs distilled.
- The model is a fused full HF/MLX directory, not a small adapter package.
- Do not use for: general Pace planner
- Do not use for: multi-domain agentic planning without routing
Source evidence
Every number above traces to one of these artifacts. Content hashes are recorded where the source file is committed.
- model card
specialists/qwen3-4b-file-ops-distilled/model_card.md· committed-package-file · sha256d26f8cff8f6b5dbf… - eval report
specialists/qwen3-4b-file-ops-distilled/eval_report.json· committed-package-file · sha25656e73eeb14a0f4c3… - reproducibility lock
specialists/qwen3-4b-file-ops-distilled/tinygpt.lock.json· committed-package-file · sha256e28380d3b483cc88… - prompt contract
specialists/qwen3-4b-file-ops-distilled/prompt.md· committed-package-file · sha2561507e58e65e50604… - specialist registry entry
specialists/registry.json· committed-registry · sha25697a78c4c1783fad0… - recorded result source
docs/learn/tool-calling-frontier-parity.md#81-climbing-the-cliff--frontier-trajectory-distillation-2026-06-16· historical-record · sha256b3e2ac1d4d4d53dd…· The document the legacy score was recorded in. - recorded result source
docs/learn/tool-calling-frontier-parity.md#84-breadth--narrow-distillation-causes-negative-transfer-2026-06-16· historical-record · sha256b3e2ac1d4d4d53dd…· The document the legacy score was recorded in. - published weights
hf://models/posttrainllm/qwen3-4b-file-ops-distilled· external-artifact · Public weight location; not hashed by this compiler.
How to read the evidence states
- measured
- Read directly from a source artifact for this candidate.
- derived
- Computed from other recorded values (for example a delta).
- historical
- Imported from a legacy record without current canonical provenance. Treat as weaker than a measurement.
- skipped
- Deliberately not run for this candidate.
- not recorded
- Evidence should exist but does not. No number is implied.
- not applicable
- The check does not apply to this candidate.