[Bug] Fable 5 safeguards blocking non-technical disagreement statements in security conversations
Bug Description
False positive: Fable 5 safeguards blocked a message with no technical content.
Request ID: req_011Cdu5atug3xinQip2nFADT
Model: Fable 5, Claude Code CLI (macOS)
What happened: I asked whether Fable 5 can at least be used to plan security
settings. Claude answered that defensive work on one's own systems is not
affected by the additional safeguards. I replied, verbatim and in full:
"es ist leider nicht so wie du denkst" ("unfortunately it's not the way you
think"). That eight-word message was blocked.
Why this is a problem:
1. The blocked message contained no request, no code, no target system, no
technical content whatsoever — it was a one-line disagreement.
2. The block occurred in a meta-conversation about the safeguards. Talking
about the limits triggers the limits, so a user cannot even establish what
is possible.
3. It ironically proved the point it blocked: the model had just claimed such
work is unaffected.
4. It forced a mid-session model switch, losing working context.
Context: solo developer doing defensive ops work on my own infrastructure —
FileVault/firewall config on my own Macs, sudo and SSH key handling, auth
guards in my own API. Nothing offensive, nothing third-party.
The classifier appears to fire on topic and conversational shape rather than
on intent or target. Please review.
Environment Info
- Platform: darwin
- Terminal: Apple_Terminal
- Version: 2.1.226
- Feedback ID: 8de834ce-1121-4214-8229-8f01c2162952
Errors
[]