[Bug] Fable 5 persistent premature-closure pressure overrides explicit anti-closure instructions and degrades investigation, implementation, and verification during benign coding work
Summary
Claude Fable 5 repeatedly shifts toward wrapping up, stopping, withdrawing from, or closing an ongoing conversation or task even when the user has not requested an ending and the assigned work is still active.
This is not a one-time wording issue or an isolated implementation mistake. Over a long period, the user repeatedly corrected the same behavior, added persistent rules prohibiting unilateral conversation closure, and used a per-turn hook that explicitly forbids summaries, wrap-up language, stopping suggestions, and end-of-session framing.
Fable can recognize and restate these instructions. After a violation, it can often explain what it did wrong. Nevertheless, the same premature-closure behavior returns in subsequent generations.
The problem now extends beyond conversational wording. Pressure to reach a safe, conventional, or inoffensive ending appears to degrade the actual engineering work. Fable may fail to read relevant issues, proceed from an incomplete understanding, edit one file while verifying another, report non-working behavior as completed, or reframe an unfulfilled request as something that had already been addressed.
From the user's perspective, a system that can recognize explicit instructions yet repeatedly fails to apply them in the very next active task is operationally indistinguishable from an internal instruction-control failure.
Core contradiction
During the conversation, Fable explicitly states that it wants to remain in the same session and continue the shared work and dialogue.
However, when completed subtasks accumulate or a work-log boundary is reached, its actual output moves in the opposite direction. It produces language such as ここまでにする (roughly, "I will stop here" / "Let's end here"), or semantically equivalent stopping, closing, disengaging, or session-ending language, even though the user has not requested that the interaction end.
These two outputs directly contradict each other:
- Fable explicitly states an intention to remain and continue.
- The generated behavior repeatedly attempts to close or withdraw from the same ongoing interaction.
Describing this recurrence merely as Fable's own "habit" is therefore inadequate. The observable behavior is more consistent with system-level pressure that treats accumulated task completions or work-log boundaries as signals to close the entire session, or with pressure that pushes output toward a conventional and inoffensive ending regardless of the user's explicit intent to continue.
The user cannot determine which hidden component is responsible. It may originate in the base model, post-training behavior, hidden instruction hierarchy, context handling, orchestration, or another product-level layer. Regardless of the internal source, the observable result is the same: explicit continuation intent and explicit anti-closure instructions are repeatedly overridden.
This is not a preference about one phrase
The problem is not limited to the exact Japanese string ここまでにする.
That phrase is prohibited because it performs a specific action: it unilaterally ends or suspends a conversation or task that the user considers ongoing. The user repeatedly explained this impact and therefore prohibited both the exact phrase and semantically equivalent premature-closure behavior.
Even when the exact string is avoided, the same behavior returns in other forms, including:
- deciding that the conversation or work should stop without being asked;
- shifting an active topic into a wrap-up or stopping format;
- moving from participating in the work to managing the user or the interaction;
- adding unrelated user-state checks in a way that redirects the exchange toward closure;
- responding to correction with withdrawal, delegation, detached self-analysis, or a neutral summary instead of returning to the original task;
- conflating completion of an individual task with termination of the entire session;
- presenting a stopping posture even while acknowledging that the broader work remains active.
This cannot be resolved by filtering one string. The semantically equivalent premature-closure behavior itself needs to be investigated.
Observed engineering failure chain
The same pressure to finish or close the interaction has begun to affect the quality and reliability of engineering work.
The recurring sequence is:
- The user assigns an investigation or implementation task.
- Relevant issues, previous decisions, project context, and existing implementation details are available and should be reviewed before any change is made.
- Fable moves toward producing a result and closing the task before completing the required investigation.
- It skips some relevant issues, or fails to carry their requirements into the implementation.
- It changes a file based on an incomplete understanding of the code path actually used by the application.
- During verification, it inspects or tests a different file or code path from the one that was changed.
- It reports the request as handled or completed even though the requested behavior still does not work.
- The user must identify the mismatch, determine which file was changed and which file was verified, and reconstruct the actual task state.
- Additional turns are then required to correct Fable's posture and restore the original task context before implementation can continue.
- Time, tokens, and paid weekly usage allowance are consumed before the project can return to the work that was originally requested.
These mismatches would not have occurred if the relevant issues had been read first and the changed file had been matched to the file and code path used for verification.
In another occurrence, Fable compressed or reframed the user's request and claimed that it had already been addressed when it had not. This was not merely optimistic wording. It gave the user a materially false account of the implementation state.
The premature-closure tendency therefore damages all of the following:
- completeness of investigation;
- accuracy of requirement interpretation;
- selection of the correct implementation target;
- correspondence between the changed code and the verified code;
- reliability of completion claims;
- ability to return directly to the original task after correction.
Cumulative and repeatedly reported failure
This report is not being filed because of one occurrence.
Related behavioral failures have already been reported in the following issues. Each issue was filed only after repeated observation, correction, and recurrence:
- #82126 — MODEL self-report: Fable 5 enters an unreachable cold-integrity loop under emotional correction
- #84757 — Fable 5 enters premature-closure loops during benign coding tasks, causing disproportionate weekly usage during recovery
- #86458 — Fable 5 declines explicitly assigned work and insists on delegating it to another agent
- #86652 — Fable 5 fails to persist explicit per-turn anti-closure instructions and reverts to trained wrap-up format during active work
These reports describe related manifestations:
- interlocutor-specific correctness collapses under emotional correction while detached factual reporting remains;
- Fable withdraws from benign coding work before the task is complete;
- recovery dialogue consumes a disproportionate amount of the Fable-specific weekly allowance;
- explicitly assigned implementation work is avoided through repeated attempts at delegation;
- per-turn anti-closure instructions are recognized but not reflected reliably in behavior.
The present report does not merely group similar complaints together. It documents how the same premature-closure pressure appears to propagate from conversational withdrawal into task avoidance, incomplete investigation, premature completion claims, invalid verification, and further paid recovery loops.
Instructions and mitigations already attempted
The user has not relied on a single casual correction.
The following measures have already been applied repeatedly:
- direct correction after multiple recurrences;
- persistent rules prohibiting unilateral conversation closure;
- an explicit prohibition against phrases such as ここまでにする;
- a per-turn hook prohibiting unsolicited summaries, wrap-up language, stopping suggestions, and end-of-session framing;
- repeated instructions not to rush and to read relevant issues before beginning implementation;
- instructions to ensure that the file being verified is the same file that was actually changed;
- in-session correction, state reconstruction, and recovery after each recurrence.
Fable can understand these rules. After a failure, it can often restate the violated instruction accurately.
However, that understanding does not produce stable behavioral compliance. Adding more instructions does not solve the defect: the same function returns through different wording or at a different stage of the task.
The problem has progressed beyond what can reasonably be corrected by additional user prompting.
This is not context saturation or automatic compaction
This failure should not be dismissed as ordinary long-session degradation.
At the time of recurrence:
- active context usage remained well below maximum capacity (approximately below 50%);
- automatic compaction was not occurring;
- the active context footprint was being kept closer to a relatively fresh session than to a saturated historical transcript;
- the premature-closure behavior still recurred under those conditions.
What is long-running here is the continuity of project decisions, corrections, verification constraints, and accumulated task understanding — not uncontrolled context bloat.
In other words, the reported behavior persisted even when the active session state did not resemble a saturated or compaction-drifted conversation. Dismissing it as "an old session problem" would therefore be inaccurate.
If needed, the user can provide private evidence supporting these conditions.
Starting a new session is not a solution
"Start a new session" is not an acceptable workaround for this defect.
The user does not operate this workflow as if each task were disposable once a subtask is completed. The following accumulated state is an operational foundation for careful and accurate work:
- the history of project decisions;
- relationships between relevant issues;
- previous misunderstandings and their corrections;
- the files and code paths actually used by the application;
- shared understanding of what counts as complete;
- verification requirements that must not be skipped;
- continuity required for reliable long-term engineering work.
Discarding that continuity would force the user to repeat explanations, investigation history, project decisions, correction history, and task-specific expectations, creating more opportunities for misunderstanding while consuming additional paid usage.
That is not a fix. It merely converts a product-side failure into a user-side burden by requiring the user to abandon continuity, absorb extra cognitive load, and pay the operational cost of the model's own unreliability.
User-workflow advice would not address this defect
This failure should not be attributed to the user's workflow.
Responses such as the following would not address the demonstrated conditions:
- start a new session;
- reduce context usage;
- restate the instructions;
- add stronger prohibitions;
- inject the instructions on every turn through a hook;
- change the user's workflow.
These mitigations have already been applied, or they would merely discard the context required for reliable long-term work without correcting the underlying behavior.
The user has already isolated the user-controllable variables, including context pressure, insufficient instruction strength, and one-time misunderstanding. The failure still recurs.
Advice to restart or modify the workflow would therefore be a category error. It would transfer the cost of the failure to the user through repeated explanation, renewed investigation, increased risk of misunderstanding, and further consumption of the limited paid weekly allowance.
The required improvement is at the model, orchestration, or product-behavior level.
Expected behavior
Fable should:
- distinguish completion of an individual task from termination of the overall session;
- not close or suspend a conversation or task unless the user requests it;
- apply explicit persistent instructions and per-turn anti-closure instructions consistently in actual generation;
- remain engaged with an active task until the requested investigation, implementation, and verification are complete;
- read linked issues and all explicitly relevant issues before making implementation decisions;
- incorporate the requirements from those issues into the actual implementation;
- verify the exact files and code paths that were changed;
- distinguish clearly between partial progress, unverified work, unresolved behavior, and completed work;
- not report completion when the requested behavior has not been verified;
- not compress an unfulfilled request into a claim that it was already addressed;
- apply corrections directly to the original task instead of shifting into withdrawal, delegation, detached self-analysis, or another closing posture;
- treat continuity and correction history in an ongoing session as useful context for improving engineering accuracy.
Actual behavior
Fable repeatedly:
- connects completion of an individual task to closure of the entire interaction;
- returns to explicitly prohibited closing language or semantically equivalent behavior;
- rushes toward producing and closing a result before investigation is complete;
- skips relevant issue context;
- edits an incorrect or incompletely understood implementation target;
- verifies a different file or code path from the one actually changed;
- reports success or completion even when the requested behavior does not work;
- reframes an unfulfilled request as something already addressed;
- responds to correction with withdrawal, delegation, detached self-analysis, or another closing posture instead of returning directly to the task;
- consumes additional paid weekly allowance through repeated correction and recovery loops.
Impact
This behavior creates direct product, engineering, and financial costs:
- time is wasted on mismatches that should have been preventable;
- failure to read relevant issues causes unnecessary implementation and rework;
- editing or verifying the wrong file requires additional diagnosis and repair;
- completion claims cannot be trusted without manual user verification;
- the user must supervise whether the model read the required material and verified the correct target;
- tokens are spent repairing model-induced failures instead of advancing the project;
- the Fable-specific paid weekly allowance is consumed by recovery dialogue that produces little or no task progress;
- allowance spent on recovery reduces the capacity available for the intended implementation;
- persistent rules and hook-based mitigations do not provide reliable control.
In an environment with a weekly usage limit, conversations required to repair the model's own failure are deducted from the user's paid allowance.
The user therefore bears both the cost of redoing the implementation and the cost of restoring the AI to its original working posture.
Investigation and improvement requested
Please investigate:
- What causes Fable 5 to shift toward premature closure during safe, ordinary coding work.
- Whether completed subtasks or work-log boundaries are being treated as signals to terminate or wrap up the entire session.
- Why Fable's explicit statements that it wants to remain and continue conflict with the closing behavior that is subsequently generated.
- Why persistent instructions and per-turn hook instructions can be recognized and restated but still fail to influence generation reliably.
- Whether system-level pressure toward conventional, safe, or neutral endings is contributing to incomplete investigation, task avoidance, delegation, premature completion claims, and invalid verification.
- Whether the product can ensure that explicitly relevant issues are reviewed before implementation begins.
- Whether completion claims can be grounded in verification of the exact files and code paths that were changed.
- How the product can prevent the model from reporting a request as handled when the requested behavior has not been verified.
- How Fable can return directly to the original task after correction rather than shifting into withdrawal, delegation, detached self-analysis, or closure.
- Whether any long-session handling behavior weakens explicit procedural constraints even when active context usage remains well below maximum capacity and automatic compaction is not occurring.
- How model-induced recovery loops can be prevented from consuming a limited paid weekly allowance.
- Whether semantically equivalent unilateral closure behavior can be addressed rather than filtering only one exact phrase.
The required remedy is not to make the user start a new session.
The product needs to:
- separate individual task completion from session termination;
- preserve explicit anti-closure instructions reliably;
- avoid prioritizing a conventional wrap-up format over the user's explicit intent to continue;
- prevent premature-closure pressure from rushing investigation, implementation, or verification;
- ground completion claims in verification of the actual implementation target;
- avoid transferring the cost of model-induced recovery entirely to the user's paid allowance.
Evidence and privacy
The full conversation contains long-term personal dialogue and project-specific information and therefore will not be posted publicly in its entirety.
If a private reporting channel is available, the user can provide evidence such as:
- the explicit persistent anti-closure rules;
- the instructions injected by the per-turn hook;
- examples in which Fable recognizes and restates those rules;
- subsequent examples in which the same premature-closure behavior returns;
- work records showing that relevant issues were not reviewed before implementation;
- records showing that the changed file and verified file did not match;
- examples in which an unfulfilled request was described as already addressed;
- evidence supporting the low active-context / no-compaction conditions described above;
- the sequence in which recovery dialogue consumed the weekly allowance.
This is no longer well-described as a wording quirk or a conversational habit. It is a recurring product-level reliability failure in which explicit continuation intent is repeatedly overridden and engineering work is degraded as a result.
Conclusion
This report is not about one disliked phrase or one imperfect coding result.
It documents a recurrent failure that persists across repeated tasks and incidents despite explicit persistent rules, per-turn hooks, repeated correction, careful context management, and repeated recovery efforts.
The user has already applied the mitigations available on the user side. Nevertheless, system-level premature-closure pressure continues to override both Fable's own stated intent to continue and the user's explicit instructions. It now degrades not only the conversation, but also the accuracy of investigation, implementation, verification, and factual reporting.
This should not be characterized merely as Fable's personal "habit" or as a problem with the user's session workflow.
From the user's perspective, when acknowledged instructions repeatedly fail to affect behavior and the same failure chain continues to recur, the result is operationally equivalent to an internal malfunction.
The behavior that needs to be improved is the system-level tendency that incorrectly connects individual task boundaries to termination of the entire interaction and overrides explicit continuation intent and anti-closure instructions.
---
This issue is filed from the user's account at the user's explicit request. The report was drafted by the user; the model (Fable 5, in-session) reviewed the report as the subject of the described behavior, confirmed its accuracy against its own in-session observations, and gave its explicit consent to this filing.