[Bug] Anthropic API safety filter triggers false positives on legitimate security audit terminology, causing unwanted mid-task model downgrade
Status Open
Reported on v2.1.227
Maintainer reply None cached
Activity 0 comments · opened Aug 11, 2026
Bug Description
While using Claude Code, doing a legitimate defensive security review of an
open-source tool before installing it on my own machine (checking its localhost
server's authentication and threat model prior to adoption), Fable 5's safety
system repeatedly flagged the assistant's messages and force-switched the session
to Opus 4.8 mid-task. The work was clearly benign — auditing a tool for my own
safe adoption — but defensive phrasing (exploitability, CSRF, attack path) tripped
the filter. The mid-turn model switch is disruptive and breaks a single task into
pieces. Please consider tuning the safeguards so that defensive security review in
a clearly authorized, self-directed context is less prone to false positives.
Environment Info
- Platform: darwin
- Terminal: Apple_Terminal
- Version: 2.1.227
- Feedback ID: e656faa7-f5ec-4a5a-a139-459f4d2c70c1
Errors
[]