[Bug] Safety classifier flags legitimate security audit setup as cyber threat

Status Open
Reported on v2.1.239
Maintainer reply None cached
Activity 0 comments · opened Aug 22, 2026

Bug Description
False positive from the safety classifier (tag [cyber]). I am a professional setting up a local, isolated security-audit lab (vulnerable WordPress in Docker on WSL2, attacked from Kali on WSL2) on my own machine, to offer audits to my clients with their permission. The system flags dual-use terms — "Kali", "network testing", "kill the agent", "bypass Avast's protection" (which just means disabling one antivirus module on my own PC so it doesn't freeze my own tests) — without the context, which is legitimate. This flag fired while the conversation was specifically about how to appeal these false positives. I have nothing to hide and expressly consent to Anthropic reviewing any of my conversations on this topic in full. Suggestion: weigh the session context, not isolated keywords. This is defensive, authorized work on my own systems.

Environment Info

  • Platform: win32
  • Terminal: null
  • Version: 2.1.239
  • Feedback ID: 5b217496-455a-43c5-945d-c4b563cfe967

Errors

[{"error":"TelemetrySafeError: VirtualMessageList: itemKeys/messages length desync (keys=112 messages=111 range=[90,112))\n    at LU0 (B:/~BUN/root/cli:23311:34430)\n    at _kg (B:/~BUN/root/cli:23311:27293)\n    at Br (B:/~BUN/root/cli:2993:21426)\n    at Nc (B:/~BUN/root/cli:2993:40533)\n    at ys (B:/~BUN/root/cli:2993:51462)\n    at $7e (B:/~BUN/root/cli:2993:89118)\n    at cke (B:/~BUN/root/cli:2993:88065)\n    at o4 (B:/~BUN/root/cli:2993:87885)\n    at U_ (B:/~BUN/root/cli:2993:84150)\n    at mt (B:/~BUN/root/cli:2993:6692)","timestamp":"2026-08-22T08:49:45.916Z"}]

View original on GitHub ↗