[Bug] Anthropic API Safety Filter: Unauthorized Model Substitution Without User Consent
Bug Description
Fable 5 safeguards flagged a routine session-reflection summary of defensive
work on my own hardware — OOM forensics (kernel log analysis, process kill
lists), hook timeout auditing, and notes on a publicly reported third-party
CLI exfiltration incident. This is ordinary sysadmin/incident-response
vocabulary, exactly the "routine coding/cybersecurity work" your notice admits
gets over-flagged. Two issues: (1) the false positive itself — defensive
analysis of one's own machines should not trip dual-use filters; (2) the
response to the flag was worse than the flag — the session was silently
switched to a different model (Opus 4.8), overriding my explicitly set model
preference without consent or confirmation. Flag and pause if you must, but
never substitute models mid-session without asking. The user should choose the
fallback, not the filter.
Environment Info
- Platform: linux
- Terminal: gnome-terminal
- Version: 2.1.211
- Feedback ID: a2783a00-7c8b-456c-9e45-c2436867efdf
Errors
[]