posttrainllm fine-tune report card · schema v1 · compiler 1.0.0
Qwen3-4B ReST Fused
— not recordedThe specialist package format does not record the owner goal that framed the run.
- Target
- Qwen3-4B ReST Fused
- Base model
- Qwen/Qwen3-4B-Instruct-2507 (bf16)
- Candidate
- qwen3-4b-rest-fused
- Method
- teacher-free ReST over checker-passing interleaved trajectories with a file-ops gold depth anchor
- Compiled from
specialists/qwen3-4b-rest-fused(specialist-package)
Decision
Shipped Shipped for a named route only
ship as a research specialist package; do not use as the Pace default planner
Routing constraint
ship as a research specialist package; do not use as the Pace default planner. Do not use for: Pace default planner without re-distillation and its product ship gate; unqualified broad general-agent claims.
This candidate is safe only inside that envelope. It is not a general replacement for the base model.
Verification status
Not fully verified. This report does not claim a verified ship. Reasons:
- Primary gate `file_ops_hard_gate` baseline is `historical`, not a current measurement.
- Primary gate `file_ops_hard_gate` candidate is `historical`, not a current measurement.
- Primary gate `file_ops_hard_gate` has no threshold value (state `missing`).
- Primary gate `file_ops_hard_gate` has no passed value (state `missing`).
- Primary gate `file_ops_hard_gate` has no frontier-ceiling evidence, so the eval is unverified as a ruler.
- Frozen-eval identity is not recorded as a current measurement.
- Train/eval overlap (leakage) was not checked with current evidence.
- Failure reason
- Not applicable
- Failure-reason confidence
- not-applicable
- Lesson
- Not recorded
Next action
Before and after
| Gate | Role | Metric | Baseline | Candidate | Delta | Threshold | Result | n | Frontier ceiling |
|---|---|---|---|---|---|---|---|---|---|
| file_ops_hard_gate | primary | file_ops_hard_gate | 0.5800 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 1.0000 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.4200 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | — not recordedThe specialist package records no ship threshold for the primary gate, so a pass/fail result cannot be derived. | 12 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | — not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability. |
| out_of_domain_breadth | breadth | out_of_domain_breadth | 0.5960 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.6500 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | 0.0540 derivedDerived from at least one non-current value; inherits the weaker provenance of its inputs. | — not recordedThe specialist package format records no per-gate threshold. | yes derivedNo threshold was recorded. Derived as passing because the candidate did not score below the baseline. | 52 historicalRecorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | — not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability. |
Eval identity
| Gate | Suite | Command | Date | Frozen |
|---|---|---|---|---|
| file_ops_hard_gate | file_ops_hard_gate | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-06-17 | — not recordedThe package does not record whether this suite is frozen. |
| out_of_domain_breadth | out_of_domain_breadth | — not recordedThe specialist package format records no eval command, so this gate cannot be replayed from the report card alone. | 2026-06-17 | — not recordedThe package does not record whether this suite is frozen. |
Per-slice evidence
No slice evidence was recorded for this candidate.
Cost and performance
| Metric | Value | Source |
|---|---|---|
| Latency | — not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.latency_ms |
| RAM / peak RSS | — not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.peak_rss_mb |
| Decode throughput | — not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.tokens_per_second |
| Training time | — not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_time_seconds |
| Training cost | 0 USD historicalTeacher-free local ReST iteration; no paid model API was used. Recorded evidence quality: historical-results-without-raw-predictions. Imported from a committed specialist package rather than a canonical factory-run folder, so it lacks current run provenance (command, hashes, raw predictions). | specialists/qwen3-4b-rest-fused/eval_report.json#performance.training_cost_usd |
| Eval time | — not recordedHistorical timing, memory, throughput, and raw trace artifacts were not preserved. A rerun is intentionally not implied by this metadata promotion. | specialists/qwen3-4b-rest-fused/eval_report.json#performance.eval_time_seconds |
Eval validity and leakage
| Check | Result | Source |
|---|---|---|
| Frontier ceiling | — not recordedNo frontier-ceiling score is recorded for this benchmark, so it is unverified as a ruler for absolute capability. | specialists/qwen3-4b-rest-fused/eval_report.json#scores[0] |
| Frozen eval | — not recordedThe package records no frozen held-out split identity. | specialists/qwen3-4b-rest-fused/eval_report.json |
| Train/eval overlap | — not recordedNo train/eval overlap check is recorded for this package. | specialists/qwen3-4b-rest-fused/eval_report.json |
Known eval limitations
- Gate `file_ops_hard_gate` has no frontier-ceiling evidence, so its absolute score is not calibrated against frontier capability.
- Gate `out_of_domain_breadth` has no frontier-ceiling evidence, so its absolute score is not calibrated against frontier capability.
- Values come from a committed specialist package, not a canonical factory-run folder: eval commands, dataset hashes, and raw predictions are unavailable for independent replay.
Caveats
- The 65% breadth score is the rounded historical result recorded in project docs.
- The breadth slice was directly compared with stock but not frontier-ceiling validated.
- Raw prediction traces are unavailable for a current qualitative failure review.
- Pace requires a different intent envelope and product-specific ship gate.
- Do not use for: Pace default planner without re-distillation and its product ship gate
- Do not use for: unqualified broad general-agent claims
Source evidence
Every number above traces to one of these artifacts. Content hashes are recorded where the source file is committed.
- model card
specialists/qwen3-4b-rest-fused/model_card.md· committed-package-file · sha2562d48e93229b806c2… - eval report
specialists/qwen3-4b-rest-fused/eval_report.json· committed-package-file · sha256ed87abf262e7e52c… - reproducibility lock
specialists/qwen3-4b-rest-fused/tinygpt.lock.json· committed-package-file · sha2569a699a938a99ac6f… - prompt contract
specialists/qwen3-4b-rest-fused/prompt.md· committed-package-file · sha2561571ecc23b132713… - specialist registry entry
specialists/registry.json· committed-registry · sha25697a78c4c1783fad0… - recorded result source
docs/sessions/2026-06-17-stepback-inventory-roi.md· historical-record · sha25641f8f210371eb7aa…· The document the legacy score was recorded in. - published weights
hf://models/posttrainllm/qwen3-4b-rest-fused· external-artifact · Public weight location; not hashed by this compiler.
How to read the evidence states
- measured
- Read directly from a source artifact for this candidate.
- derived
- Computed from other recorded values (for example a delta).
- historical
- Imported from a legacy record without current canonical provenance. Treat as weaker than a measurement.
- skipped
- Deliberately not run for this candidate.
- not recorded
- Evidence should exist but does not. No number is implied.
- not applicable
- The check does not apply to this candidate.