[Bug] Anthropic API Error: False positive safety filter on legitimate safety/moderation product context

Status Open
Reported on v2.1.220
Maintainer reply None cached
Activity 0 comments · opened Aug 2, 2026

Bug Description
My message was flagged by the automated safeguards and switched from Fable 5 to Opus 5, but it contained nothing that should have triggered a filter. The message was routine project context for a faith-based accountability app I'm building — folder access, session memory, and past ticket history. There was no code, no request, just context loading. I suspect the false positive came from benign subject-matter keywords in my project (it deals with content moderation, image screening, and user safety features), which seem to have been misread as something harmful. This is exactly the kind of legitimate work the broad filters are catching by mistake. Please consider tuning the safeguards so ordinary product-development context around safety and moderation topics doesn't get flagged.

Environment Info

  • Platform: darwin
  • Terminal: Apple_Terminal
  • Version: 2.1.220
  • Feedback ID: b370c20c-ef47-4e68-8944-d55f6efc0d0d

Errors

[]

View original on GitHub ↗