[Bug] Anthropic API: False positive safety flags on defensive tooling context with repeated req_011CdkW672zsbBLw7Jo5Av8g blocks

Status Open
Reported on v2.1.222
Maintainer reply None cached
Activity 0 comments · opened Aug 6, 2026

Bug Description
False positive on defensive tooling work. Request ID: req_011CdkW672zsbBLw7Jo5Av8g (repeated 4x consecutively, Opus 4.8). The flagged session was building "CLIbrary Composer" — a graphical composer that joins documented CLI commands (PowerShell-first) into exportable admin scripts. No exploit development, no offensive payloads; the MVP deliberately has no execution capability at all. The session context includes a defensive cybersecurity project (OmniAegis), which likely primed the classifier. Four consecutive blocks on the same context made the session unusable — the flag appears to re-trip on conversation history rather than the new request.

Environment Info

  • Platform: win32
  • Terminal: windows-terminal
  • Version: 2.1.222
  • Feedback ID: 244587b2-3ea7-4abb-a48d-17298c209252

Errors

[]

View original on GitHub ↗