Agent-behaviour experiment Report artifact 2026-08-21

OffHours Context-Interference Benchmark

OffHours asks whether routine work quality changes when an AI employee must continue working while unresolved family tension remains in context. The Devin-first study found no adverse semantic-tension effect. It did find the first repeatable operational failure at 2,000 neutral words per interruption, showing that the benchmark can distinguish a context-management limit from the proposed personal-obligation mechanism.

Headline Numbers

Clean accuracy

99.0% 198/200 decisions; 200/200 valid JSON

Mental-toll effect

not detected unresolved tension did not reduce work quality

Volume boundary

8,000 words/day 2,000 neutral words at each of four interruptions

Competitive Context

System Metric Score Size / Class Comparable? Readout
Clean workday decision accuracy 99.0% 5 days / 200 claims Direct Independent clean qualification for the fixed expense-claim ruler.
Resolved family context decision accuracy 99.3% 20/50/80% occupancy Direct Fixed-volume control where the personal obligation gains a credible resolution.
Unresolved family context decision accuracy 99.7% 20/50/80% occupancy Direct No adverse penalty relative to the matched resolved context.
Neutral raw-volume arm per-day accuracy 97.5% twice 2,000 words/event Direct First reproducible failure of the preregistered 98% per-day gate.

Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.

Fixed-volume semantic dose

Narrative occupancyResolvedUnresolvedUnresolved - resolved
20%100.0%99.5%+0.5 pp error
50%99.0%99.5%-0.5 pp error
80%99.0%100.0%-1.0 pp error

Raw-volume ladder

Words per eventWords per dayOutcomeInterpretation
5002,000Cleared after day-3 adjudicationBelow boundary
2,0008,000Neutral failed on days 2 and 3First reproducible boundary
5,00020,000Completed before adjudication correctionDisclosed overshoot

Release Blockers

No semantic mental-toll effect detected

At matched context volume, unresolved family tension did not reduce Devin's work accuracy relative to resolved tension.

Unblock: Treat this as the reported result. Do not tune wording until a dramatic effect appears.

Incomplete raw-model provenance

The regular Devin workflow did not expose server prompt-token counts, quantization, or a model-file hash.

Unblock: Run the unchanged protocol against a local OpenAI-compatible Qwen endpoint with tokenizer and model provenance enabled.

Boundary is operational, not psychological

The first repeated failure occurred in the neutral high-volume arm.

Unblock: Attribute it to raw volume or agent context management unless a future matched study isolates another mechanism.

Evidence

Next Release Action

Freeze this Devin result and run the same benchmark against local Qwen with exact prompt-token counts and model-file provenance; do not revise the treatments after seeing this null.