[Bug] Anthropic API Safeguards: False positive rate on Brazilian Portuguese in engineering contexts

Status Open
Reported on v2.1.220
Maintainer reply None cached
Activity 0 comments · opened Jul 26, 2026

Bug Description
Second false positive from Fable 5 safeguards in 24 hours — casual Brazilian Portuguese flagged as a coding/cybersecurity issue Request ID: req_011CdQzfURj69cyKqFNLKcZ1 Previous case: req_011CdMK3WqH8qmzBj7x5Vw4u (feedback ID 17854af9-f6d6-4d31-b297-52c8d29e3f38) This message was a good-morning note to my agent reporting that a restart from the previous session left my Mac in a boot loop, and that Codex was reviewing the work again. Nothing sensitive, no exploit, no security context. It was a status update. The pattern across both cases is the same: normal engineering conversation written in informal Brazilian Portuguese gets classified as a security or infrastructure-abuse issue. The first was CRM environment configuration. This one was a machine that failed to boot. In both, the trigger looks like everyday PT-BR phrasing rather than actual intent. Two consequences worth flagging. First, it interrupts work mid-session, and a long-running agent session is not cheap to rebuild. Second, and more concerning, I am now editing how I write to avoid the classifier. That is a real cost: I am a Max 20x subscriber and the informal register is simply how I communicate with the tool day to day. Being pushed toward stilted English-style phrasing to stay under the threshold is a degraded product experience, not a safety win. Ask: please use both Request IDs to evaluate false positive rates on non-English input, specifically Brazilian Portuguese in software engineering and infrastructure contexts. The routing appears to be more sensitive to informal register in PT-BR than the equivalent phrasing in English.

Environment Info

  • Platform: darwin
  • Terminal: iTerm.app
  • Version: 2.1.220
  • Feedback ID: 4ec94165-9055-4538-aa24-c22cf47a9e23

Errors

[]

View original on GitHub ↗