Fable 5 safeguards false-positive on camera-protocol dev and public court/registry browser automation — ~20 silent sticky switches to Opus 4.8 in 4 weeks

Status Open
Maintainer reply None cached
Activity 0 comments · opened Aug 14, 2026

Summary

Fable 5's safeguards are repeatedly false-positive flagging legitimate consumer-app development and public-records research, silently switching my sessions to Opus 4.8 mid-task. Local transcripts document ~20 switch events across 13 distinct sessions between 2026-07-17 and 2026-08-14. The switch is sticky: once flagged, the session never returns to Fable 5. /feedback has been submitted for these multiple times with no visible response or change, so I'm filing here.

Banner text (verbatim): "Fable 5's safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate coding, cybersecurity, and biology tasks. Switched to Opus 4.8."

Per support article 15363606, an Opus 4.8 fallback indicates the offensive-cybersecurity category — none of my flagged work is offensive security.

Environment

  • Claude Code CLI on macOS (darwin), Max subscription, model claude-fable-5, max effort
  • Auth: claude.ai (Pro/Max)

Frequency (from my local ~/.claude/projects transcripts — grep for the banner text)

| Session date | Flag events |
|---|---|
| 2026-07-17 | 1 |
| 2026-07-17 | 1 |
| 2026-07-20 | 1 |
| 2026-07-24 | 1 |
| 2026-07-24 | 1 |
| 2026-07-28 | 1 |
| 2026-07-28 | 1 |
| 2026-07-29 | 3 |
| 2026-07-30 | 1 |
| 2026-07-31 | 2 |
| 2026-08-01 | 1 |
| 2026-08-03 | 1 |
| 2026-08-05 | 5 |
| 2026-08-13 | 1 |

The flagged workloads are benign

The work in these sessions falls into three classes, none remotely offensive-security:

  1. iOS camera-control app development — implementing vendor camera protocols (PTP/IP-class) for my own app, analyzing packet captures of my own cameras on my own LAN. Ordinary consumer-hardware dev.
  2. Driving my own Safari browser on public government websites — the state business-entity registry and county superior-court public records portals (California), read-only research for my own small-claims matter. The most recent flag (2026-08-13, 23:10 PT) fired mid-way through screenshotting a court's public case-register page.
  3. General repo review/audit work.

Why the stickiness is the worst part

Documented example from the 2026-08-13 incident: the switch fired at 23:10 and the session ran on Opus 4.8 for roughly 12 more hours — the transcript's model records show claude-fable-5 through line 1043 and claude-opus-4-8 from line 1041 to the end (line 1340, last write 10:53 the next morning). Everything produced after the flag — none of it security-related — was silently produced by a different model than the one I pay for and had selected, and I only discovered it from a screenshot of the banner. Load-bearing research from that window then had to be independently re-verified.

I've now set switchModelsOnFlag: false so sessions pause instead (good that this exists — thank you), but that trades silent degradation for hard stops on work that should never have been flagged.

Asks

  1. Precision: tune the cyber classifier so that (a) consumer camera-protocol development and (b) automating the user's own browser against public government registries/court portals stop tripping it. These are common, legitimate Claude Code workloads.
  2. No sticky switches: re-evaluate per message (or offer one-keystroke return to the selected model) instead of degrading the entire remainder of a session after a single flagged message.
  3. Visibility: keep the currently-active model visible in-session so a switch can't go unnoticed for 12 hours.
  4. An appeal path that answers: the support article mentions a Cyber Verification Program for legitimate work — is there a timeline, and would workloads like the above qualify? Multiple /feedback submissions have produced no visible acknowledgment; even an automated "received, classifier report logged" would help.

Happy to provide sanitized transcript excerpts for any of the listed incidents on request.

View original on GitHub ↗