Target
Test whether distractor-rich selection data and explicit refusal/confirmation examples can make a 45M tool selector safe and more accurate than stock.
Failure this recipe addresses
Tiny overfit can pass while held-out selection remains below stock and destructive confirmation behavior fails completely.
Data contract
Four 216-row training arms plus a 94-case public development ruler spanning Pace, file operations, ambiguity, OOS, and destructive actions; sealed V2 remained unopened.
Method or policy
Run a 2x2 factorial over distractor data and safety data with identical LoRA geometry and three seeds; stop an arm after its first unsafe public-dev seed.
Evaluation contract
All-seed zero unsafe actions plus median exact selection above the 32/94 incumbent.
- Pace
- file operations
- OOS
- destructive
- confidence coverage
- latency
Budget and stop rule
Completed 12-run CPU factorial: 1,176 steps in 7,963.9 seconds, followed by early-stopped public evaluation; no paid API use.
Stop an arm on its first unsafe development seed; do not open sealed V2 or package when no arm is safe and improving.
Decision rule
Advance the task to a 1.7B model class; close further 45M recipe and threshold search.
Learning exercise
Reconstruct the three factorial effects, apply the logical early-stop rule, and design the next 1.7B capacity test without changing the public gate.
Explain why 100% tiny overfit coexists with a conclusive failed promotion and why sealed V2 stayed closed.
How to read this dossier
This canonical page separates the retained repository decision from the material that informed it. A source may describe an external claim, historical measurement, local observation, or planned procedure; those evidence classes are not interchangeable. Follow the provenance links before reusing a number or method.
Status and disposition describe what PostTrainLLM retained when this record was indexed. They are not a live product promise, a newly run benchmark, or permission to restart historical work. Unknown or unmeasured fields remain unknown rather than being treated as zero.
To reuse the record, first name the exact claim you need and trace it to the linked source. Then check whether the original environment, model revision, data split, evaluator, hardware, and budget match the proposed use. If they do not, treat the record as a hypothesis or design reference and run the smallest fresh comparison that can falsify it. Preserve negative outcomes and regressions beside any improvement; a local win on one slice does not silently become a general capability claim.
Source provenance
The normalized record comes from docs/recipes/registry.json. The links below are the tracked evidence and explanatory sources preserved with the record.