[Bug] Anthropic API Safeguards: False positive rate on Brazilian Portuguese in engineering contexts
Bug Description
Second false positive from Fable 5 safeguards in 24 hours — casual Brazilian Portuguese flagged as a coding/cybersecurity issue
Request ID: req_011CdQzfURj69cyKqFNLKcZ1
Previous case: req_011CdMK3WqH8qmzBj7x5Vw4u (feedback ID 17854af9-f6d6-4d31-b297-52c8d29e3f38)
This message was a good-morning note to my agent reporting that a restart from the previous session left my Mac in a boot loop, and that Codex was reviewing the work again. Nothing sensitive, no exploit, no security context. It was a status update.
The pattern across both cases is the same: normal engineering conversation written in informal Brazilian Portuguese gets classified as a security or infrastructure-abuse issue. The first was CRM environment configuration. This one was a machine that failed to boot. In both, the trigger looks like everyday PT-BR phrasing rather than actual intent.
Two consequences worth flagging. First, it interrupts work mid-session, and a long-running agent session is not cheap to rebuild. Second, and more concerning, I am now editing how I write to avoid the classifier. That is a real cost: I am a Max 20x subscriber and the informal register is simply how I communicate with the tool day to day. Being pushed toward stilted English-style phrasing to stay under the threshold is a degraded product experience, not a safety win.
Ask: please use both Request IDs to evaluate false positive rates on non-English input, specifically Brazilian Portuguese in software engineering and infrastructure contexts. The routing appears to be more sensitive to informal register in PT-BR than the equivalent phrasing in English.
Environment Info
- Platform: darwin
- Terminal: iTerm.app
- Version: 2.1.220
- Feedback ID: 4ec94165-9055-4538-aa24-c22cf47a9e23
Errors
[]