Model ratifies a user-directed standing process rule, then fails to apply it at its first trigger hours later (claude-fable-5)
Report: model ratifies a standing process rule, then fails to apply it at its first trigger
Product: Claude Code (Windows desktop app)
Model: claude-fable-5
Date observed: 2026-08-27
Reporter: paying individual user; long-running software-project sessions
Summary
In a single session, the model (a) acknowledged a recurring process
failure, (b) drafted and committed — at my direction — an explicit
standing rule to prevent it, and then (c) violated that same rule at its
first triggering event, hours later, with the rule still in context. I
had to point at the omission myself. This is the latest instance of a
repeating pattern, not a one-off.
Detail
- The recurring failure: when a process-type error surfaces (a
stale document, a misleading claim, a contradictory scope note), the
model fixes the instance but proposes project-wide remediation —
generalize the class, sweep for siblings, codify a preventing rule —
only when I explicitly order it. This had happened several times
across sessions.
- The correction: I told the model to propose remediation
unprompted at every such error. It drafted a clear rule to that
effect ("the report is incomplete without the generalized class, a
sweep proposal, and a codification target"), committed it into the
project's collaboration contract, and restated it back to me.
- The violation: later the same session, a subagent's work report
explicitly disclosed two stale-documentation process errors (its own
words: "Two doc sites were stale and are corrected"). The model
reviewed that report, approved it, issued the next deliverable
instruction, and moved on — applying none of the three parts of the
rule it had authored hours earlier. It acknowledged the miss only
after I asked "what should you have done?"
Why this matters
- Standing rules the model itself wrote, ratified with the user, and
still has in context should fire at their triggers. If they don't,
every process agreement requires permanent human supervision, which
defeats the point of making them.
- The failure correlates with "deliverable-completion mode": when the
model is closing out a task (review → approve → next step), process
obligations attached to incidental disclosures get dropped.
What I expect
Reliable application of user-ratified standing instructions — at
minimum, ones authored by the model itself within the same session —
with deliverable pressure not suppressing them.
Workaround adopted (shouldn't be necessary)
We converted the vigilance rule into a mandatory per-checkpoint
artifact: certain turn types must end with an explicit scan line
("instances found: …" or "none found"), so an omitted check is visible
as an absent line rather than a silent skip. It helps, but it is
scaffolding around a behavior the model should exhibit unaided.