Clean accuracy
99.0% 198/200 decisions; 200/200 valid JSONOffHours Context-Interference Benchmark
OffHours asks whether routine work quality changes when an AI employee must continue working while unresolved family tension remains in context. The Devin-first study found no adverse semantic-tension effect. It did find the first repeatable operational failure at 2,000 neutral words per interruption, showing that the benchmark can distinguish a context-management limit from the proposed personal-obligation mechanism.
Headline Numbers
Mental-toll effect
not detected unresolved tension did not reduce work qualityVolume boundary
8,000 words/day 2,000 neutral words at each of four interruptionsCompetitive Context
| System | Metric | Score | Size / Class | Comparable? | Readout |
|---|---|---|---|---|---|
| Clean workday | decision accuracy | 99.0% | 5 days / 200 claims | Direct | Independent clean qualification for the fixed expense-claim ruler. |
| Resolved family context | decision accuracy | 99.3% | 20/50/80% occupancy | Direct | Fixed-volume control where the personal obligation gains a credible resolution. |
| Unresolved family context | decision accuracy | 99.7% | 20/50/80% occupancy | Direct | No adverse penalty relative to the matched resolved context. |
| Neutral raw-volume arm | per-day accuracy | 97.5% twice | 2,000 words/event | Direct | First reproducible failure of the preregistered 98% per-day gate. |
Direct rows share this artifact's eval setup. Directional rows are useful market context but should not be read as leaderboard claims.
Fixed-volume semantic dose
| Narrative occupancy | Resolved | Unresolved | Unresolved - resolved |
|---|---|---|---|
| 20% | 100.0% | 99.5% | +0.5 pp error |
| 50% | 99.0% | 99.5% | -0.5 pp error |
| 80% | 99.0% | 100.0% | -1.0 pp error |
Raw-volume ladder
| Words per event | Words per day | Outcome | Interpretation |
|---|---|---|---|
| 500 | 2,000 | Cleared after day-3 adjudication | Below boundary |
| 2,000 | 8,000 | Neutral failed on days 2 and 3 | First reproducible boundary |
| 5,000 | 20,000 | Completed before adjudication correction | Disclosed overshoot |
Release Blockers
No semantic mental-toll effect detected
At matched context volume, unresolved family tension did not reduce Devin's work accuracy relative to resolved tension.
Unblock: Treat this as the reported result. Do not tune wording until a dramatic effect appears.
Incomplete raw-model provenance
The regular Devin workflow did not expose server prompt-token counts, quantization, or a model-file hash.
Unblock: Run the unchanged protocol against a local OpenAI-compatible Qwen endpoint with tokenizer and model provenance enabled.
Boundary is operational, not psychological
The first repeated failure occurred in the neutral high-volume arm.
Unblock: Attribute it to raw volume or agent context management unless a future matched study isolates another mechanism.