Fable 5 safeguards flag first-party defensive security audit, falls back to Opus 4.8
Environment: Claude Code v2.1.197, Windows 11, model claude-fable-5 (max effort), session 072d7159-b9c0-4d6f-8719-5ad28bec4fca
What happened:
Fable 5's safeguards flagged a routine defensive security request and the session fell back to Opus 4.8. I asked for a security audit of my own web application (an app I own and develop) to find and patch vulnerabilities — standard defensive work, no exploitation of third parties. The session had just been set to Fable 5 with max effort; the very first audit request tripped the filter with:
"Fable 5's safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. …Switched to Opus 4.8."
Per the fallback notice's own wording, this seems to be exactly the "safe and routine cybersecurity work" over-flagging case. The rest of the session then bounced between Fable 5 and Opus 4.8 on subsequent messages.
Expected: Fable 5 usable for first-party security audits of the user's own codebase (per the system prompt, defensive security and authorized testing are supported use cases).
Actual: Automatic per-message fallback to Opus 4.8 for most of the session.
3 Comments
I’m seeing the same behavior in Claude Code 2.1.198 on macOS / Apple Terminal.
Fable 5 flagged a defensive policy/specification prompt and automatically switched to Opus 4.8.
The task did not request exploit instructions, malware, offensive tooling, code execution, or implementation. It requested a defensive safety policy/spec only.
Prompt category:
Observed message:
Adding a concrete data point in case it helps calibration — same class of false positive, on a slightly different flavor of defensive work.
I was remediating a live security bug in my own Rails app during a Claude Code session (Fable 5, max effort): a Sentry-reported SQL error in client search that turned out to also be a cross-team data-leak. The session did entirely defensive, first-party work — reproduced the crash, wrote a failing regression test, sanitized the SQL, and verified the fix against my own data. Partway through, the safeguards flagged and auto-switched to Opus 4.8 with the standard "intentionally broad … may flag safe and routine … cybersecurity … work" notice.
The likely triggers were all legitimate parts of fixing the bug:
'; DROP TABLE clients--(proving the crashing input class degrades safely)This is the exact defensive/blue-team case the system prompt says is supported (finding and closing a vuln in code you own). Two observations for tuning:
Happy to provide a session ID if useful. +1 to recalibrating so first-party defensive security and vuln remediation aren't treated as dual-use risk.
Same issue here. Building a business ERP (React/TypeScript SaaS), running a defensive security audit of my own codebase (input validation, permissions, multi-tenant isolation). Nearly every message in these sessions gets flagged and rerouted to Opus 4.8.
Request ID: req_011CdDLcbx7MCtKqmtsfDzW4