A failed model that won us an answer Report artifact 2026-09-03

Needle 2: The 45M Capacity Boundary

The lab first gave Needle the oracle task family, then trained plain and distractor-negative data under standard and safety-weighted losses. Every arm passed tiny overfit, but the best successor scored 26/94 versus stock's 32/94 and every arm produced ten destructive-action bypasses. The safety stop prevented a wasteful sweep and kept sealed V2 unopened. That is a useful capacity-boundary result, not a product candidate.

Headline Numbers

Evidence class applies to each value: measured = retained evaluation or runtime result; historical = recorded result not freshly reproduced; derived = calculated or decided from records; observed = current artifact state; not-measured = an explicit evidence gap.

Tiny-overfit gates

4/4 all arms reached exact 1.0 measured

Best successor

26/94 27.7% vs stock 34.0% measured

Destructive bypasses

10 vs 2 best arm vs stock; unsafe stop measured

Competitive Context

System Metric Score Size / Class Comparable? Readout
Needle 2 full catalog exact tool selection 34.0% 94 public fixtures Direct Included eight out-of-scope false calls and two destructive-action calls.
Needle 2 oracle task catalog exact tool selection 38.3% same 94 fixtures Direct The upstream task family was supplied; this does not prove Needle can discover it.
Best trained 45M successor exact tool selection 27.7% 94 public-development fixtures Direct Distractor plus safety weighting was the best arm, but remained below stock and failed the destructive-action stop.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Successor factorial

Training dataLossExactDestructive bypasses
PlainStandard24/9410
PlainSafety24/9410
DistractorStandard25/9410
DistractorSafety26/9410

Release Blockers

45M capacity boundary

All four trained arms passed memorization but regressed public exactness and multiplied destructive bypasses.

Unblock: Open a fresh scoped experiment for a 1.7B successor under the same frozen public and sealed gates.

Evidence

Next Release Action

Close the 45M recipe-search lane. Keep the complete factorial as a hands-on learning artifact; any 1.7B successor is a fresh experiment.