posttrainllm fine-tune report card · schema v1 · compiler 1.0.0

Qwen3-4B ReST Fused

not recordedThe specialist package format does not record the owner goal that framed the run.

Target
Qwen3-4B ReST Fused
Base model
Qwen/Qwen3-4B-Instruct-2507 (bf16)
Candidate
qwen3-4b-rest-fused
Method
teacher-free ReST over checker-passing interleaved trajectories with a file-ops gold depth anchor
Compiled from
specialists/qwen3-4b-rest-fused (specialist-package)

Decision

Shipped Shipped for a named route only

ship as a research specialist package; do not use as the Pace default planner

Routing constraint

ship as a research specialist package; do not use as the Pace default planner. Do not use for: Pace default planner without re-distillation and its product ship gate; unqualified broad general-agent claims.

This candidate is safe only inside that envelope. It is not a general replacement for the base model.

Verification status

Not fully verified. This report does not claim a verified ship. Reasons:

Failure reason
Not applicable
Failure-reason confidence
not-applicable
Lesson
Not recorded

Next action

not recordedThe specialist package format records no machine-readable next action. The release action lives in the public artifact registry (docs/factory/public-artifacts.md) as prose.

Before and after

Every gate with baseline, candidate, derived delta, threshold, result, sample size, and frontier-ceiling evidence.
Gate Role Metric Baseline Candidate Delta Threshold Result n Frontier ceiling
file_ops_hard_gate primary file_ops_hard_gate 0.5800 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 1.0000 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.4200 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. not recordedThe specialist package records no ship threshold for the primary gate, so a pass/fail result cannot be derived. 12 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability.
out_of_domain_breadth breadth out_of_domain_breadth 0.5960 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.6500 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). 0.0540 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. not recordedThe specialist package format records no per-gate threshold. yes derivedNo threshold was recorded. Derived as passing because the candidate did not score below the baseline. 52 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability.

Eval identity

Which suite produced each gate, how it was invoked, and whether it is frozen.
Gate Suite Command Date Frozen
file_ops_hard_gate file_ops_hard_gate not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-06-17 not recordedThe package does not record whether this suite is frozen.
out_of_domain_breadth out_of_domain_breadth not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. 2026-06-17 not recordedThe package does not record whether this suite is frozen.

Per-slice evidence

No slice evidence was recorded for this candidate.

Cost and performance

Latency, memory, throughput, and training cost/time. Absent evidence is reported as not measured, never as zero.
MetricValueSource
Latency not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. specialists/qwen3-4b-rest-fused/eval_report.json#performance.latency_ms
RAM / peak RSS not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. specialists/qwen3-4b-rest-fused/eval_report.json#performance.peak_rss_mb
Decode throughput not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. specialists/qwen3-4b-rest-fused/eval_report.json#performance.tokens_per_second
Training time not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_time_seconds
Training cost 0 USD historicalTeacher-free local ReST iteration; no paid model API was used. Recorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_cost_usd
Eval time not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. specialists/qwen3-4b-rest-fused/eval_report.json#performance.eval_time_seconds

Eval validity and leakage

Whether the benchmark is a trustworthy ruler: frontier ceiling, frozen-eval identity, and train/eval overlap.
CheckResultSource
Frontier ceiling not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability. specialists/qwen3-4b-rest-fused/eval_report.json#scores[0]
Frozen eval not recordedThe package records no frozen held-out split identity. specialists/qwen3-4b-rest-fused/eval_report.json
Train/eval overlap not recordedNo train/eval overlap check is recorded for this package. specialists/qwen3-4b-rest-fused/eval_report.json

Known eval limitations

Caveats

Source evidence

Every number above traces to one of these artifacts. Content hashes are recorded where the source file is committed.

How to read the evidence states

measured
Read directly from a source artifact for this candidate.
derived
Computed from other recorded values (for example a delta).
historical
Imported from a legacy record without current canonical provenance. Treat as weaker than a measurement.
skipped
Deliberately not run for this candidate.
not recorded
Evidence should exist but does not. No number is implied.
not applicable
The check does not apply to this candidate.