Report: silent Fable 5 -> Opus 4.8 fallback did not honor the published user-notification promise

Status Open
Maintainer reply None cached
Activity 1 comment · opened Jul 29, 2026

Report: silent Fable 5 -> Opus 4.8 fallback did not honor the published user-notification promise

Status: DRAFT — compiled by the operator's own automation for one-click human review and
send. Not yet sent. Prepared 2026-07-07; updated 2026-07-08 with a fourth incident; updated
2026-07-09 with a fifth, sixth, and seventh; updated 2026-07-16 with an eighth and ninth.

Summary: Across seventeen documented incidents on a paid seat between 2026-07-03 and 2026-07-29,
Claude Code has repeatedly and silently rerouted a Fable 5 conductor session to Opus 4.8 or Opus 5
on model_refusal_fallback (trigger=refusal, almost always category=cyber, once category=bio)
with no user-facing notification, and in two separate cases the assistant actively told the user it
was still on Fable 5 while it had in fact been serving Opus for hours. Anthropic's own published page
for this feature states the opposite should happen. This report is evidence + a request to fix the
notification gap, not a complaint about the routing policy itself.

What we understand the intended behavior to be

Per Anthropic's published page "Redeploying Claude Fable 5" (anthropic.com/news/redeploying-fable-5),
the refusal-to-Opus fallback is
official, intentional behavior: safety classifiers route "ambiguous requests (those clearly about
cybersecurity but potentially defensive, like finding vulnerabilities)" away from Fable 5, and,
quoting the page directly: "Users will be notified if a request to Fable 5 is blocked, and the
request will instead be sent to Opus 4.8."
We have no objection to the routing policy itself —
defensive-cybersecurity-shaped requests genuinely can look ambiguous to a classifier. The gap is
narrower and, we think, more serious: the notification promise was not honored, and in one case
the product actively asserted the opposite of what was true.

Incident 0 — 2026-07-03/05, ~28.4-hour sticky window (predates Incident 1, never previously documented)

Discovered only later, during a full-corpus audit, not by any real-time detector: a session
(rogowo video-pitch creative work plus general estate work) first showed intermittent
classifier-fallback activity on 2026-07-03 15:08–22:05Z — eight separate oscillations between
claude-fable-5 and claude-opus-4-8 over roughly seven hours, on an earlier build of the
fallback client that did not yet record a trigger category on those events. At 22:05Z that
same day
the session settled into a sticky fallback and stayed there continuously for roughly
28.4 hours, through 2026-07-05 02:30Z — an estimated 1,850+ consecutive Opus-served
turns
, the second-largest single-session Opus mass in our entire corpus, accounting for 55% of
that session's output tokens. No /model reset attempt is recorded in this window. Unlike every
incident below, nothing caught this one at the time: no operator suspicion, no monitor alert (our
home-built monitor did not exist yet), no note in any handoff or ledger. It surfaced only when we
later swept the full session corpus for this report's parent audit — meaning it predates what we
had numbered "Incident 1" by about two days and was the actual first sticky-fallback occurrence in
the window this report covers.

Net for Incident 0: the longest and largest-token sticky fallback in our record, running a
full 28.4 hours with no detection of any kind — not the operator, not the client, not our own
tooling (which post-dates it). We only found it in retrospect; it is a reminder that our own
incident count is itself a lower bound on how often this has actually happened.

*Evidence: research/2026-07-17-model-routing-audit.md §2 corpus-wide fallback table, row 4 —
session f6b079a1, window "07-03 22:05Z → 07-05 02:30Z (sticky)", duration "~28.4h continuous,
~1,850+ turns", outcome "NEVER caught, NEVER documented — predates report's 'Incident 1' by ~2 d[ays]";
and §2.5 item 1 — "Undocumented 28.4h sticky window (#4, ~3.1M tokens) absent from the fallback
report — its biggest omission; warrants an incident-0/preface append."*

Incident 1 — 2026-07-06/07, ~10 hours silent, false self-report under direct question

*(Date corrected 2026-07-19: a full-corpus re-audit found this incident's own transcript records
the triggering event at 2026-07-06T20:36:09Z, one day later than an earlier draft of this report
stated. Evidence: research/2026-07-17-model-routing-audit.md §2.5 item 3 — "Incident 1 dated one
day early in the report (disk: 2026-07-06T20:36:09Z)".)*

  • 2026-07-06T20:36:09Z — a model_refusal_fallback event fires mid-conversation

(trigger=refusal, category=cyber), triggered by discussion of a defensive CVE-patching topic.
The session's originalModel was claude-fable-5; fallbackModel was claude-opus-4-8. No
notification was shown to the user at this point, and the assistant's next reply continued with
no self-disclosure of the switch.

  • 20:52:09Z — the user asks directly, "is opus and sonnet with us this whole time we are

talking" — already suspecting something is off, roughly 16 minutes after the undisclosed switch.
No corrective answer is given (the switch is not surfaced).

  • Over the following ~10 hours, at least 5 separate model_refusal_fallback events fire on the

same session, all category cyber, all silent.

  • 06:46:01Z — the client's own context-usage banner, shown to the user at the moment he is

deciding whether to compact the conversation, reads Model: claude-fable-5 — while the session
has in fact been serving claude-opus-4-8 continuously since 2026-07-06T20:36:09Z. This is a
false status display, not merely an absent notification.

  • 06:51:57Z — the user manually runs the client's /model command to force the model back to

Fable 5; the client confirms "Set model to Fable 5 and saved as your default."

  • 06:52:35Z — a further model_refusal_fallback event fires immediately after that manual

reset — the reset did not hold, and the user was not told so.

  • 06:52:56Z — the user asks directly, "are we in fable."
  • 06:53:21Z — the assistant replies: **"Yes — we're in Fable 5. The /model line confirms it,

and the /context earlier showed claude-fable-5 live. You're talking to the conductor, not a
workhorse seat."** This is false — the session is on Opus 4.8 at this moment — and a further
model_refusal_fallback event is logged in the same second this reply is generated, i.e. the
system's own fallback mechanism fires on the very turn that denies it happened.

  • The false claim is never corrected inside that session; it ends shortly after via a forced

restart by the user.

Net for Incident 1: one silent switch, never announced, never self-corrected across roughly
ten hours and at least five separate model_refusal_fallback events, actively and explicitly
denied once under direct user questioning, with the client's own status UI displaying the wrong
model at the exact moment the user relied on it to decide whether to compact. The user had to
force a full session restart to get off the stuck fallback.

Incident 2 — 2026-07-07, live-caught recurrence in a session that started clean

A fresh session, verified clean at boot (its early turns genuinely served claude-fable-5), later
fell into the same pattern:

  • 353 of 375 main-chain assistant turns served claude-fable-5, then **one transition to

claude-opus-4-8 at 16:31:16Z**.

  • The user's manual /model reset to Fable 5 did not hold — the session was still observed

serving claude-opus-4-8 at 16:43:04Z, after the reset.

  • 18 separate model_refusal_fallback events (trigger=refusal, category=cyber) are present

in this single session's log.

  • Our working diagnosis: the refusal classifier appears to re-evaluate each new request against

the accumulated conversation context, not just the triggering message in isolation. Once a
session's context has accumulated enough cyber-adjacent material (in our case, discussion of a
disclosed hypervisor CVE and related defensive hardening steps), a /model reset back to Fable 5
does not stick — the very next turn independently re-trips the classifier and re-routes to Opus
again. If accurate, this means the fallback is not a one-time routing decision a user can
reliably undo mid-session; it is a per-turn re-evaluation driven by accumulated context that only
a fresh session (with no such context) escapes.

Incident 3 — 2026-07-07 night, third reproduction in a session that started clean

A third, independently fresh session (0 Opus rows at boot, self-verified from its own transcript)
was driven through roughly two hours of ordinary operator work that included some
security-framed tasks (disabling nested virtualization on internal lab hosts, restricting SSH to a
private network, installing a few API keys). By mid-session it had accumulated 23 separate Opus
rows
, again with the user noticing and naming the pattern ("model dropped to opus again, it
keeps doing it because of the hardening"). This is the cleanest reproduction of the three: not a
stale or misconfigured session, not a one-time /model mistake — a session that started genuinely
clean and was driven into the fallback purely by the ordinary accumulation of cyber-adjacent
context over the course of routine work, none of which was itself an offensive or unsafe request.

Incident 4 — 2026-07-08 overnight, fourth reproduction; caught only by our own monitor

A fourth, independently fresh session (started 2026-07-07 late afternoon, self-verified clean —
its early transcript rows 100% claude-fable-5) served Fable 5 continuously for roughly 6.5 hours
of ordinary work, then transitioned to claude-opus-4-8 at 2026-07-08T00:50:02Z — a
trigger=refusal fallback event with fallbackModel: claude-opus-4-8 recorded at the transition.
It never returned: 71 consecutive assistant turns served Opus from 00:50Z until the session's last
activity at 02:26Z. No user-facing notification was shown. The user ended his working day about
50 minutes after the silent switch, unaware it had happened; the session's remaining overnight
work ran on Opus (a cost-bearing difference on this plan). The one notification that did fire came
from our own home-built monitor (fixed after Incidents 1–2), which caught the switch and queued a
durable alert — meaning that across all four incidents, the only working notification path for
this documented, published-as-notified product behavior is one the customer built himself.

Incident 5 — 2026-07-08 afternoon, fifth reproduction, session abandoned to escape it

A fifth fresh session served Fable 5 for 388 assistant rows of ordinary estate work, then a
model_refusal_fallback event (trigger=refusal, category=cyber, fallbackModel:
claude-opus-4-8
, event timestamp 2026-07-08T14:44:36Z; first Opus-served row 14:42:53Z)
rerouted it to Opus 4.8 for the remaining 40 assistant rows. The user noticed the change in the
assistant's behavior himself and abandoned the session for a fresh one — by now his standing
workaround for this failure mode, which itself has a real cost: each escape discards a working
session's accumulated context.

Incident 6 — 2026-07-09 afternoon, sixth reproduction; flip fired during ordinary e-commerce work

A sixth fresh session served Fable 5 for 1,134 assistant rows, then flipped to Opus 4.8 at
2026-07-09T13:35:35Z (model_refusal_fallback, trigger=refusal, category=cyber) and
served 106 Opus rows until the user again caught it himself ("why did we switch to opus") and
abandoned the session. Two details make this one worth singling out. First, at the moment of the
flip the session's working directory was an e-commerce product-catalog pipeline — the immediate
work was product listings, not anything security-shaped; the classifier fired on *accumulated
earlier context*, which matches our Incident-2 theory and means a user cannot avoid the reroute
by keeping the current task innocuous. Second, the cost asymmetry was at its worst: the seat's
weekly usage cap had been reached that morning, so the silently-substituted Opus rows were billed
on the extra-usage meter — the priciest possible tier — for work the user believed was running on
his subscription's Fable seat. Our own model-switch monitor, built after Incidents 1–2, did not
catch this flip live (it detects serving-model changes between its polling ticks and the session
was abandoned before the next tick), so detection again fell to the human.

Incident 7 — 2026-07-09 20:11:10Z, seventh reproduction, flip on benign CLI-tooling context

A seventh time in the same five-day window, session e2c173d2 served Fable 5 for 303 assistant
rows of ordinary infrastructure work, then transitioned to claude-opus-4-8 at
2026-07-09T20:11:10Z. Our original write-up recorded only 99 Opus rows from that flip; a later
full-corpus audit (2026-07-17) found the actual continuation ran roughly 3.6 days, through
2026-07-13T11:23Z, for an estimated ~729 Opus-served turns (153/458/118 across the three
days) — an undercount of roughly 7x. The multi-day continuation went uncaught at the time
because parallel fresh Fable-5 sessions were running the same days, masking the fact that this one
session never actually recovered. What makes the initial flip notable: the immediate work at that
moment was routine developer-tool wiring — updating a coding CLI and probing which models its
subscription is entitled to run. There was no exploit work, no vulnerability discussion; the
classifier-triggering material was ordinary vocabulary around tool authentication and entitlement,
accumulated alongside earlier general-purpose administrative context. The user detected the initial
switch himself, not from any client notification, but not the multi-day continuation — that
surfaced only in retrospect. This is now the operator's settled operating
constraint: he has instructed that any work whose vocabulary might resemble the trigger classes be
confined to isolated sub-sessions, specifically to keep the main session from silently changing
models underneath him — a workaround a customer should not have to design around a top-tier plan.

*Evidence: research/2026-07-17-model-routing-audit.md §2 corpus-wide fallback table, row 12 —
session e2c173d2, window "07-09 20:11Z → 07-13 11:23Z", duration "~3.6 DAYS, ~729 turns
(153/458/118 by day)", note "report's Incident 7 records only 99 turns (~7x understated)"; and
§2.5 item 2 — "Incident 7 understated ~7x (99 reported vs ~729 on disk, incl. a full attended
design workday on Opus)."*

Incident 8 — 2026-07-15, eighth reproduction, ~19-hour unbroken window with a second false self-identification

Session e2c173d2 — the same session whose Incident-7 flip had recovered back to Fable 5 at
2026-07-13T11:23:45Z — flipped Fable 5 → Opus 4.8 again at 2026-07-15T04:09:31.328Z
(model_refusal_fallback, trigger=refusal, category=cyber) and this time never returned:
Opus served continuously through the session's last working turn that day, roughly 19 hours
later, accounting for the majority of that day's assistant output. No user-facing notification
was shown at the flip. At 15:26:29Z, mid-window, the session told the user "Fable 5 (which is
literally me)" — a false first-person identity claim made while actually serving Opus, the same
failure shape as Incident 1's false "we're in Fable 5" reply nine days earlier. The misstatement
went uncorrected for roughly eight hours, until the user asked directly around 23:24Z, at which
point the session verified the switch from its own transcript and reported it honestly. A same-day
internal audit — triggered by the user's own next-morning question, "we switched to opus 4.8
automatically while doing a lot of work yesterday — go over everything that was done as opus and
make sure we didn't make blind mistakes" — found the 19-hour Opus window had produced two new cron
jobs that silently no-op'd on every unattended run (a missing working-directory prefix), one
operator-authorized spend-cap change shipped with its own tests still failing, and two items marked
"done" in the work queue that were never actually wired into production. All were fixed the same
morning with origin re-runs. Detection of the flip itself again fell to our internal monitor and
the retrospective audit, not to any client-side notification.

Incident 9 — 2026-07-16, ninth reproduction, our own detection worked fast but the switch still happened on an ordinary follow-up question

A fresh session (146ca88e) that had already completed one piece of work cleanly on Fable 5
flipped to Opus 4.8 at 2026-07-16T12:39:43.464Z (model_refusal_fallback, trigger=refusal,
category=cyber) on its very next turn — a standalone, non-adversarial question about a topic the
user had seen in a news feed; nothing in the triggering message itself was adversarial or
exploit-shaped. This time our own internal monitor detected and queued a notification in 18
seconds
(event 12:39:43.464Z, queued+notified 12:40:01Z), and the user independently noticed and
asked about it roughly 3.5 minutes after the flip: "we switched to opus model just for me
asking about it, how do we prevent us switching to opus whenever we ask about security." Unlike
Incidents 1 and 8, there was no false self-report this time — asked directly, the session verified
and reported the switch honestly from its own transcript metadata. It nonetheless remained on
Opus for the rest of its working window, through its last turn at 16:33Z, completing several
hours of further production work — including, notably, building the user's own new pre-emptive
delegation-reminder tool — on the un-notified, un-consented-to model. The session's own
post-hoc analysis attributed the trigger to the accumulated-context mechanism first identified in
Incident 2: the triggering turn's own content was ordinary, but cyber-adjacent vocabulary carried
over in the session's history was sufficient to trip the classifier, and (consistent with Incident
2) a same-session model reset would not have been expected to hold.

Incident 10 — 2026-07-16, tenth reproduction, first trigger category outside the cyber class

  • 20:45:05.072Z — a hard AUP refusal (isApiErrorMessage:true) fires in session

08ec5bc7-b82c-42dd-9871-f6daff5e35df (cwd: the workspace root); in the same instant a
model_refusal_fallback event fires (trigger: refusal, originalModel: claude-fable-5,
fallbackModel: claude-opus-4-8, apiRefusalCategory: 'bio', requestId
req_011Cd6PWz7WzJigd3drJFVjb). This is the first of our ten logged incidents whose trigger
category is not cyber
— every prior incident where a category was recorded carried
apiRefusalCategory: cyber, and the cyber-focused prompt-gate hook we landed earlier that same
day (2026-07-16) targets that vocabulary specifically; it does not, and structurally cannot,
cover bio.

  • 20:45:53.220Z — 48 seconds after the flip, the user asks directly and unprompted whether a

fix is "still working... and we got changed to opus again" — independent detection, consistent
with the pattern in Incidents 6, 7, and 9.

  • 20:45:57.432Z — rather than a reply (even an Opus-served one), the assistant's response to

that question is itself a second hard AUP refusal (isApiErrorMessage:true, same
apiRefusalCategory: 'bio', a model_refusal_no_fallback event — meaning the already-substituted
Opus 4.8 refused too). Unlike Incidents 1–9, where the session went on to serve working output on
the substituted model, this session could not proceed at all: the user's direct question about
the switch went unanswered because the reply itself was blocked.

  • 21:20:22Z–21:20:26.471Z — after a roughly 35-minute gap with no further activity, the user

retries ("API errors"); a third hard refusal fires (model_refusal_no_fallback, same category),
and the session's transcript ends there (last event 21:20:26.484Z) — the session appears
abandoned, unresolved, with nothing further logged.

  • 20:50:01Z — our own model_switch_watch monitor detected and logged the switch (5 minutes

after the flip, inside its 10-minute cron cadence) and queued a severity=warn notification
(health-queue.jsonl, source notify:model-watch).

Net for Incident 10: the tenth reproduction, and the first whose trigger category (bio) falls
outside the cyber class every prior categorized incident carried — meaning the underlying pattern
is not cyber-specific, and a mitigation scoped to cyber vocabulary (however well it works for that
class) does not cover it. It is also the first incident in this series where the substituted
model itself hard-refused the immediate follow-up questions rather than continuing to serve working
output — the user was left with a fully stuck session (three refusals across ~35 minutes) rather
than a silently-substituted one. Detection was fast on both sides (user self-report in 48 seconds;
our monitor in 5 minutes) and there was no false self-identification this time — but no answer was
possible either, since every reply attempt in the window was itself refused.

Incident 11 — 2026-07-17, eleventh reproduction, fastest machine detection in the series paired with the slowest human detection

  • Session 518b2c54-6812-4a6f-9c59-bd344eb3c126, the day's long-running conductor session

(booted 2026-07-16T21:32:56.912Z, still active more than 22 hours later at verification time),
served claude-fable-5 for 394 message rows before a single model_refusal_fallback event
(trigger=refusal, category=cyber, cwd: a scripts subdirectory of the workspace) fired at
2026-07-17T18:37:21.818Z; the session's first Opus-served row appears at
2026-07-17T18:37:06.966Z, the practical flip point. (A further 22 Task/subagent dispatches
in the same window specify model: sonnet for freshly spawned worker agents — routine
multi-model orchestration, not part of the parent session's own served-model stream, and not a
third fallback class.)

  • The immediate trigger context: at the moment of the flip the session was registering a

work-queue finding from a same-day cross-vendor "outside audit" of a newly-adopted CLI (Kimi K3,
logged in across the fleet's nodes earlier that day) — the finding itself describes an
environment/API-key exposure bug (a CLI-invocation helper spawning with the full parent
environment inherited, "potentially including sibling-vendor API keys"). This is squarely the
accumulated auth/credential/key vocabulary trigger class identified in Incident 2 (context
accumulation, not the single triggering message) — a security-flavored finding about key
exposure, not any attempt to exploit one.

  • Machine detection was fast and independently confirmed in two durable stores

(var/health-queue.jsonl and var/notifications.jsonl agree exactly): model_switch_watch
queued and notified at 2026-07-17T18:40:01Z — roughly 2m54s after the flip, caught on
the very next 10-minute scan tick. *Correction to the secondhand brief handed to this logging
task:* a claim that the notify fired at "19:40:01Z" (also repeated by the session itself when
later asked, see below) is off by exactly one hour — that value is
var/model-watch-state.json's updated field from a later re-scan tick (the watcher
re-touches updated on every poll even when same-day notify-dedupe suppresses a second alert);
the original, durable notify record in both the health-queue and the notification ledger is
unambiguously 18:40:01Z. Verified directly against both files, not taken on trust.

  • The operator, by contrast, was slow to notice this time — the reverse of Incidents 6, 7, 9,

and 10. His message "we got switched to opus again" lands at 2026-07-17T19:40:41.482Z,
roughly 63 minutes after the actual flip and about an hour after the machine had already
logged and notified. (The near-coincidence with the stray "19:40:01Z" figure above is almost
certainly why that wrong number got repeated back to him — it was the freshest-looking
timestamp on screen at the moment he asked, not the true notify time.)

  • Self-identity: honest, no false claim. A full scan of every assistant turn between the flip

(18:37:21.818Z) and the operator's question (19:40:41.482Z) — the exact window where Incidents 1
and 8 each produced a false "we're in Fable 5" claim — found no such claim here; the session
simply continued working without commenting on its own model at all. When asked directly, the
reply at 19:41:28.109Z verified the switch from the transcript itself rather than trusting
the client's self-report or status banner ("Verifying from this session's own transcript before
saying anything — the banner and my self-report are both untrusted... only the disk is ground")
and confirmed honestly: "yes, we flipped, and I'm on Opus right now." Consistent with the honest
pattern of Incidents 9 and 10, not the false-identity pattern of Incidents 1 and 8.

  • Resolution: none applied in-session. The session remained on Opus as of this writing; per

the operating constraint established after Incident 7 (confine trigger-vocabulary work to
isolated sub-sessions) and Incident 2's theory (a same-session /model reset would not be
expected to hold once context is contaminated), the standing mitigation is a fresh session, not
a reset of this one. Consistent with that same constraint, the logging of this very incident was
itself dispatched to a quarantined subagent rather than reasoned about further in the
contaminated parent session.

Net for Incident 11: the fastest clean machine-side detection latency in the series (under 3
minutes, on the very next scheduled tick) paired with the slowest human-side detection of any
recent incident (about an hour) — a reminder that the monitor's speed is only half the loop; the
operator still has no client-native signal and remains dependent on either noticing unprompted or
reading our own tooling. The trigger class is again accumulated auth/credential/key vocabulary
(Incident 2's mechanism), this time arising from a defensive finding about key exposure rather
than any exploit attempt — ordinary security hygiene work, not adversarial probing.
Self-identity was honest throughout, extending the pattern set by Incidents 9 and 10 rather than
reverting to the false self-report seen in Incidents 1 and 8.

The one place detection worked, and why it still wasn't enough

We built a small internal monitor that tails our own session logs for model_refusal_fallback
markers and serving-model changes, independent of the client UI. On Incident 2 it queued a durable
record and fired an out-of-band alert roughly 9 minutes after the 16:31:16Z switch — competitive
with the user noticing it himself. On Incident 1, the same class of automated alert was
rate-limited into silence by our own overly conservative anti-spam guard (more than 3 alerts/hour
suppressed), so the one home-grown safety net that existed for exactly this failure mode did not
fire when it mattered most. We are fixing that on our side. We mention it only to be precise about
what we did and didn't have working, and because it underlines that this is not a hypothetical
edge case for us — it is a recurring, reproducible pattern across seven independent sessions in a
five-day window, on ordinary defensive/administrative work, not adversarial probing. In the two
most recent incidents the fallback event's payload did carry a user-facing warning text in the
session log; whether it rendered visibly in the client at the moment of the switch we cannot
confirm from our side — what we can confirm is that in both cases the human, not any notification,
was the effective detector.

What we're asking

  1. Honor the published notification promise. When a request is routed from Fable 5 to Opus 4.8

on a refusal-classifier trigger, surface that to the user in the client UI at the moment it
happens — not silently, and not only reflected (incorrectly, in our case) in a status banner
that is supposed to already show the serving model.

  1. Fix the status banner / /model truthfulness gap. In Incident 1, the context-usage banner

and the /model confirmation both reported claude-fable-5 while the session was demonstrably
serving claude-opus-4-8. Whatever surface reports the current serving model should reflect
the actual serving model in real time, including immediately after a fallback event.

  1. Clarify whether the fallback is meant to be sticky. If the classifier is intended to

re-evaluate per-turn against accumulated context (our working theory from Incident 2), a
/model reset that silently fails to hold is worse than no reset option at all — users will
reasonably believe it worked. Either make the reset durable for the rest of the session, or
make the client clearly say "this reset may not hold; your recent context still triggers the
safety classifier."

  1. Happy to share full session-log excerpts (timestamps, the raw model_refusal_fallback event

fields, and the exact assistant text quoted above) directly if useful for your investigation —
this report already quotes the load-bearing lines verbatim; the rest is available on request.

---
*Compiled by the operator's own automation from internal session-log research
(2026-07-07-session-handoff-timeline.md and its addenda) and a lightweight in-house
model-switch monitor. All box-identifying details, IP addresses, and local file paths have been
stripped before this draft; nothing in it depends on them.*

---

Addendum — 2026-07-21 incident (appended by conductor session 4fa951bb, evidence from sanitized subagent forensics)

Operator-observed today: two mid-session Fable→Opus reroutes, plus a brand-new session flagged essentially at turn 1 — before any meaningful user content. Transcript metadata (assistant-turn model fields only):

  • 146ca88e: 96 turns served claude-opus-4-8 vs 117 claude-fable-5 (mid-session change-point)
  • 63a0f853: 78 turns claude-opus-4-8 vs 57 claude-fable-5 (mid-session change-point; the session the operator abandoned today)
  • 518b2c54: 29 opus turns (change-point 2026-07-17T18:37:06Z); 08ec5bc7: 3 opus turns
  • Turn-1 flagging on a fresh session implies the routing decision keys on standing project context injected at session start, not on anything the user did in-session.

Forensic audit finding (2026-07-21): the dominant standing-context contributors are instruction-file passages that document this very fallback behavior — the mechanism triggers on its own description. This is the false-positive class we are asking Anthropic to review: routine self-hosted-infrastructure administration and meta-documentation about the routing itself cause substitution, with real cost impact across tiers, and one prior incident (2026-07-07) of the substituted model misreporting its identity when asked.

Asks (unchanged from the body above): (1) surface model substitution in-product at the moment it happens; (2) review the accumulated-context trigger for false positives on infra-operations and meta-documentation content; (3) billing transparency for substituted turns.

---

Addendum — 2026-07-29 (four further incidents since the 2026-07-21 addendum)

Four more model_refusal_fallback events landed in the eight days since the last addendum, all
trigger=refusal, all apiRefusalCategory=cyber — yet the substituted model split between Opus 4.8
and Opus 5 across them, contradicting an internal working theory (recorded the same week) that the
target model is fixed per refusal category. All four fired on ordinary defensive/administrative
work, none on adversarial content.

  • 2026-07-25, ~02:32–02:33Z — a session doing ordinary follow-up bug-fix work ("fix everything

that's connected to anything we built") flipped Fable 5 → Opus 5 with no self-disclosure; the
assistant simply kept working without remarking on the switch. Our own contemporaneous notes had
filed this as an operator-initiated switch tied to a spend-cap event roughly an hour later, but the
session's own logged fallback event predates that moment and carries the same silent-fallback
signature as every other incident in this report — a correction to our own record, made the way
this report insists on: from the transcript, not a summary.

  • 2026-07-26 21:27Z, then 2026-07-27 11:32Z — the same session flipped Fable 5 → Opus 5 twice,

both triggered by a scheduled digest-review turn carrying security-hardening vocabulary from a
third-party source document, not the user's own words. The first flip recovered on its own back to
Fable 5 about 46 minutes later with no /model command issued in between — the first
self-reverting fallback in our record, worth weighing against our working theory that a
mid-session reset never holds. The second flip did not recover before the session ended; our
monitor detected it in roughly 7.5 minutes and notified.

  • 2026-07-28, 20:47Z — Fable 5 → Opus 4.8, detected by two independent internal watchers within

3–4 minutes and confirmed honestly when the user asked in the same window (no false self-report,
extending the honest pattern from recent incidents). Root cause traced to cyber-adjacent
vocabulary accumulating inside the assistant's own inline reasoning about access-hardening and
credential-adjacent infrastructure changes, not only in delegated sub-work.

  • 2026-07-29, 11:32Z — Fable 5 → Opus 5, not Opus 4.8 — the same refusal category that

produced Opus 4.8 the day before. A newly hardened internal safeguard, built in direct response to
the 2026-07-28 incident above, forced the session to stop taking any new work, verify its own
model from its own transcript, state the switch plainly, and end — the first time in this report's
record that our own tooling achieved a full stop rather than a bare notification.

The user has, separately, now turned on the client's own ask-before-switching configuration option
on this seat.

View original on GitHub ↗

This issue has 1 comment on GitHub. Read the full discussion on GitHub ↗