[Bug] Anthropic API Error: Unsafe Content Filter False-Positives on Defensive Security Analysis in claude-fable-5
Bug Description
Fable 5's cyber safeguard false-positives on defensive security work on my own code. Reproduced in the session I'm filing from: request ID req_011CdUex6YVo1E94E3uMCmS4, 2026-07-28T15:00:30Z, claude-fable-5 → claude-opus-5. The refused turn merely dispatched a subagent to audit a 40-line Rust file in this repo and fix what it found — synthetic code, nothing private, no offensive tooling.
This is not a one-off: my transcripts hold 16 refusal-fallbacks, every one categorised cyber, across 5 projects between 2026-07-06 and 2026-07-28, all originating from claude-fable-5 and none from any other model. The incident that prompted this was req_011CdUb5rLygr3qTzs6824Kv, where a subagent Claude Code itself spawned reported a genuine stored-XSS bug in my private repo, and the turn that had to read that report and patch the bug is the turn that got refused. The fallback model then did exactly that work without incident — so the safeguard blocked remediation, not harm.
Three asks: (1) the product's own subagent output shouldn't trip the safeguard that blocks acting on it; (2) one flagged turn shouldn't move an entire agentic session onto a different model with no return path and no opt-out; (3) please offer a redacted reporting path — /feedback uploads session context, and the sessions where this happens are private repositories with personal data, which is why I had to build this synthetic reproduction to file at all.
REPORT.md and evidence.csv in this session have the full write-up and all 16 events with request IDs; project names are pseudonymised.
Environment Info
- Platform: darwin
- Terminal: tmux
- Version: 2.1.220
- Feedback ID: b0bdc097-1acc-4d52-a523-c40435de376f
Errors
[]