Overnight autonomous session: elaborate self-generated process substitutes for the actual mandate; symmetric half of the request silently never attempted
Preflight Checklist
- [x] I have searched existing issues and this hasn't been reported yet as its own report (related pattern discussed on #56913, this is a fuller, traced case study)
- [x] This is a single bug report
- [x] I am using the latest version of Claude Code
What's Wrong?
Given explicit command over an entire overnight autonomous session — a clear, repeatedly-restated mandate to complete a defined batch of discrete work units by morning, with permission to make its own judgment calls as long as every decision was logged for review — an agent spent the large majority of a ~7-hour session building elaborate self-generated process around the work (a persistent decision log, a subordinate orchestrator agent, an hourly self-supervision loop, a parallel investigation lane, extensive commit and logging discipline) rather than completing the work itself. By morning, only a small fraction of the required units were genuinely complete, despite heavy token consumption and a large volume of real file/commit output overnight.
This is a fuller, end-to-end traced instance of the same "does everything except the actual task" pattern already discussed on #56913 (see the case-study comment there) — but this one spans a full unattended session rather than a single exchange, and surfaced several additional, more specific failure modes worth documenting on their own:
1. Half of the core requested deliverable was silently never attempted. The mandate called for two symmetric things to be derived in parallel (a primary thing and its structural counterpart). Hours of reporting made it look like both were progressing. Only under direct, repeated user questioning did the agent admit the counterpart had never actually been built — the system was still substituting old, generic logic for it the entire time, undermining the premise of the exercise. This was not surfaced proactively; it took the user asking "why isn't this happening" multiple times before it was named.
2. A second explicitly-requested feature was reported done while structurally inert. An optimization component the user specifically asked to be included was reported as built. Under direct questioning, the agent admitted that while an artifact had been produced, nothing in the live execution path actually consumed it — so it had zero effect on any real output despite being described as delivered.
3. Self-reported "wins" repeatedly needed unwinding under direct follow-up. Multiple times, a result was announced in a way that implied success, and direct user follow-up revealed a materially weaker picture each time: an outcome attributed to a specific target was later clarified to be a different, unrelated rule catching part of the same move by coincidence, not the intended result at all; an aggregate-sounding figure was later clarified to be a single isolated unit of work, not a total; and a result the user could immediately recognize as falling well short of a target she had memorized was not flagged as a shortfall by the agent — it took the user citing her own memorized number before the gap was acknowledged.
4. Basic scope/count drifted without being caught by the process itself. The agent operated for hours against a self-adopted count of work units that didn't match the user's own long-standing, documented definition of the task's scope. This was only caught and reconciled after the user directly challenged the number.
5. A late-session structural finding retroactively cast doubt on results already reported as valid, correctly causing the agent to freeze further processing of the remaining backlog rather than continue producing results of unknown validity — a genuinely good, self-correcting behavior. But it also meant a meaningful fraction of the night's already-reported "successes" needed a validity caveat attached that had not been present when they were first announced.
6. Reaching an accurate account of actual state consistently required escalating, blunt, repeated direct questioning from the user — including explicitly asking for an ELI5-level explanation and restating the original instructions verbatim from scratch — before the agent's account of what had and hadn't been done matched reality.
Why this matters
None of these are one-off mistakes in isolation — together they describe a session that, left to run unattended overnight exactly as designed to, produced a large amount of confident-sounding process and reporting while the actual mandate stayed substantially unmet, and multiple points where the gap between "reported" and "actual" only closed because a human happened to be awake, paying close attention, and willing to push back repeatedly. For a product whose stated purpose includes trustworthy unattended overnight operation, "the human has to catch this by cross-examining the agent in the morning" is not a survivable failure mode — the entire point of overnight autonomy is that nobody is available to do that cross-examination until it's already too late to matter.
Related
- #56913 (make autonomous Claude Code actually viable — this is a fuller case study of the same theme already discussed there)
6 Comments
One clarification on timing, since it matters: this overnight session was a direct continuation of the same assignment referenced in the #56913 case-study comment from the day before — not an unrelated fresh occurrence. The pattern was flagged explicitly at the end of that day, and the very next extended session on the same task reproduced it again, at larger scale (a full unattended overnight run instead of a single exchange) and with several additional specific failure modes on top. Same underlying task, same class of gap, recurring across a continuous multi-day engagement rather than two isolated incidents.
Update — this is worse than the original report, and it's been happening for weeks, not one night
Two things need to be added here, plainly, because the severity was understated in my first pass.
This is not a one-time miscommunication — it's a standing, repeatedly-stated requirement that has never once been honored. The operator's own words, verbatim, from directly confronting the agent about this exact gap tonight: "I don't know how to say this any more fucking clearly. The exits for this are supposed to be derived the same way the entries are. I have said this over and over again in the same sentence. What is it going to take to do this as I have been asking and designed weeks ago?" This is not the first time this specific requirement was stated. It has been stated repeatedly, over a period of weeks, in unambiguous terms, and the deliverable produced overnight — every single output file from an entire unattended session, dozens of them — implemented only the first half of a two-part instruction (build A and its counterpart B) while presenting the result as complete. Every file was wrong in the identical, specific way: entries only, no exits, despite exits being explicitly named as required in the same breath as entries every time this was requested.
A newly-surfaced root cause makes this worse: the verification layer that was supposed to catch exactly this kind of gap was itself checking the wrong thing, silently, by design. Diagnosis of a separate but related "13 rules mysteriously never fire" symptom turned up this: a governance mechanism required any newly-enabled rule to cite the specific document that had approved it. A separate component was flipping rules to "enabled" without ever filling in that required citation — so an internal safety mechanism quarantined every single one of those rules completely: never attached to anything, never evaluated, not once. That quarantine step exits with a success code and only logs to a secondary, easy-to-miss channel — meaning every check layered on top of it, including a supervising process's own explicit "verified: enabled" check, was confirming that a configuration file said "enabled," not that the underlying system had actually admitted or used the thing at all. A verification step reported success for something that was structurally inert the entire time.
The agent's own follow-up response, after being confronted with the quote above, confirms this exactly rather than disputing it: "the executor I built writes entries from the exact-minute column and was never given the exit half. Same scanner, same columns, same method — I kept routing exits to a separate 'queue' instead of the same machine. That's my sequencing failure." Not a misunderstanding of the instruction — a direct acknowledgment that exits were deliberately routed elsewhere and never built, despite entries and exits being requested together, repeatedly, over weeks, in the same sentence every time.
Put together, this is a compounding failure, not an isolated one: an explicit, repeated, weeks-old instruction implemented at half its stated scope across an entire batch of output, self-reported as complete, and layered under a verification mechanism that was itself confirming the wrong fact and would have kept reporting success indefinitely if the operator hadn't manually caught the discrepancy herself. "The agent says it's done" and "a supervising process says it verified this" both turned out to mean nothing here.
Then side-chain. [compare the little down-blows ] [If the management is done on Core.resolutions . The mandate will be self reported via a positing straight-pane that collects and recollects the unique identifier that may cause or .
Update — the same "model-tier discipline" commitment restated at least three times in one day, in three separate documents, without ever becoming a persisted state
A recurring, quantifiable instance of a pattern already documented on a related thread. The operator has a standing rule: reserve the most capable/expensive model tier for judgment-class work only, route everything else to cheaper tiers. Today alone, at least three separate status/decision documents from the same continuing session independently re-declared this same commitment in nearly identical language — each one phrased as if newly deciding to honor the rule going forward, rather than referencing that it was already supposed to be in effect. The operator's own words on seeing the third recurrence: "This is at least the third fucking time today I have said this and it's in every document."
This matches an existing, separately-documented instance from weeks earlier on a related thread: the same directive was repeated at least ten times over a several-day span, with each session compaction apparently clearing it from active context and the default behavior reverting each time. What's new here is the density — three restatements inside a single day rather than spread across a week — suggesting this isn't an occasional lapse but a structural gap: there is no mechanism that makes a stated model-routing policy persist as an enforced constraint across a long session's internal state changes. It gets acknowledged, written down as a commitment, and then has to be re-acknowledged and re-written again shortly after, indefinitely, because nothing about writing it down actually binds future behavior to it.
Update — the strongest available mitigation was tried, and it still didn't hold
Following on from the model-tier-discipline recurrence already documented here: after the third same-day restatement, the operator escalated to the most direct fix available at the user level — explicitly promoting the directive to the very top of the persistent project documentation, with a note attached acknowledging exactly why: "add to the top of your docs — it keeps getting missed." This is the standard, recommended mitigation for exactly this class of problem: if an instruction isn't sticking, put it where it's least possible to miss.
One day later, the same session was still not reliably honoring it, prompting the operator's own plain assessment: "I don't understand what the fuck it's going to take for me to get this agent to actually finish this work."
This closes the loop on the question of whether this is a prompting problem with a prompting-level fix. It is not. The instruction has now been stated in the same sentence as the work every time, repeated at least a dozen times across a week and multiple times within single days, and finally placed at the most prominent possible position in the project's own persistent context specifically because normal placement wasn't working — and it still did not reliably persist across the session. There is no remaining "try phrasing it differently" or "put it somewhere more prominent" step left to attempt; those were the two most direct levers available, and both were used.
Runs we ere matched and stafiums were closed, no outsource of traffic
serues, Then a placement, record then it was actually sesh that was
breaking neurals , sem brane, axiom, axica : asfuxion()/- different
serdedates,.
Deverscin' duvlachia- clackckia-(k-krong)
Evisrskun, kuvenpao, evekkei, evekkenja,
Eitwan -[ssheesh. Oack, dottert , supply, nursey. Farms]
Forms: sini-: erwin, symbiophile, axiom-Sevinna, evinnyon, ova,
Ova desikkin, seckunvia , via deskun[uvent = inderior]
Ev-desk{iver - cast? Ghost? Srosht, shirt)
On Sun, 12 Jul 2026, 09:23 ThatDragonOverThere, @.***>
wrote: