posttrainllm fine-tune report card · schema v1 · compiler 1.0.0

Qwen3-4B File-Ops Distilled

not recordedThe specialist package format does not record the owner goal that framed the run.

Target
Qwen3-4B File-Ops Distilled
Base model
Qwen/Qwen3-4B-Instruct-2507 (bf16)
Candidate
qwen3-4b-file-ops-distilled
Method
frontier/gold trajectory distillation on GorillaFileSystem multi-turn tasks
Compiled from
specialists/qwen3-4b-file-ops-distilled (specialist-package)

Decision

Shipped Shipped for a named route only

ship only as a routed file-ops specialist; do not use as the general planner

Routing constraint

ship only as a routed file-ops specialist; do not use as the general planner. Do not use for: general Pace planner; multi-domain agentic planning without routing.

This candidate is safe only inside that envelope. It is not a general replacement for the base model.

Verification status

Not fully verified. This report does not claim a verified ship. Reasons:

Failure reason
Not applicable
Failure-reason confidence
not-applicable
Lesson
Not recorded

Next action

not recordedThe specialist package format records no machine-readable next action. The release action lives in the public artifact registry (docs/factory/public-artifacts.md) as prose.

Before and after

This candidate does not present an unconditional win: out_of_domain_breadth failed. Target and regression gates are reported independently below.

Every gate with baseline, candidate, derived delta, threshold, result, sample size, and frontier-ceiling evidence.
Gate Role Metric Baseline Candidate Delta Threshold Result n Frontier ceiling
file_ops_hard_gate primary file_ops_hard_gate 0.5800 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 1.0000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.4200 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. not recordedThe specialist package records no ship threshold for the primary gate, so a pass/fail result cannot be derived. 12 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 1.0000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions).
file_ops_hardgen_heldout regression file_ops_hardgen_heldout 0.6000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.9500 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.3500 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. yes derivedNo threshold was recorded. Derived as passing because the candidate did not score below the baseline. 40 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability.
out_of_domain_breadth breadth out_of_domain_breadth 0.5960 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.4230 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). -0.1730 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. no derivedNo threshold was recorded. Derived as failing because the candidate scored below the baseline on a non-primary gate. 52 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability.

Eval identity

Which suite produced each gate, how it was invoked, and whether it is frozen.
Gate Suite Command Date Frozen
file_ops_hard_gate file_ops_hard_gate not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-06-19 not recordedThe package does not record whether this suite is frozen.
file_ops_hardgen_heldout file_ops_hardgen_heldout not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-06-19 not recordedThe package does not record whether this suite is frozen.
out_of_domain_breadth out_of_domain_breadth not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-06-19 not recordedThe package does not record whether this suite is frozen.

Per-slice evidence

No slice evidence was recorded for this candidate.

Cost and performance

Latency, memory, throughput, and training cost/time. Absent evidence is reported as not measured, never as zero.
MetricValueSource
Latency not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.latency_ms
RAM / peak RSS not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.peak_rss_mb
Decode throughput not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.tokens_per_second
Training time not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.training_time_seconds
Training cost not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.training_cost_usd
Eval time not recordedThe specialist package records no value for this metric. It is reported as not measured rather than as zero. specialists/qwen3-4b-file-ops-distilled/eval_report.json#performance.eval_time_seconds

Eval validity and leakage

Whether the benchmark is a trustworthy ruler: frontier ceiling, frozen-eval identity, and train/eval overlap.
CheckResultSource
Frontier ceiling 1.0000 historicalImported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-file-ops-distilled/eval_report.json#scores[0].frontier
Frozen eval not recordedThe package records no frozen held-out split identity. specialists/qwen3-4b-file-ops-distilled/eval_report.json
Train/eval overlap not recordedNo train/eval overlap check is recorded for this package. specialists/qwen3-4b-file-ops-distilled/eval_report.json

Known eval limitations

Caveats

Source evidence

Every number above traces to one of these artifacts. Content hashes are recorded where the source file is committed.

How to read the evidence states

measured
Read directly from a source artifact for this candidate.
derived
Computed from other recorded values (for example a delta).
historical
Imported from a legacy record without current canonical provenance. Treat as weaker than a measurement.
skipped
Deliberately not run for this candidate.
not recorded
Evidence should exist but does not. No number is implied.
not applicable
The check does not apply to this candidate.