[Bug] Safety classifier flags legitimate security audit setup as cyber threat
Bug Description
False positive from the safety classifier (tag [cyber]). I am a professional setting up a local, isolated security-audit lab (vulnerable WordPress in Docker on WSL2, attacked from Kali on WSL2) on my own machine, to offer audits to my clients with their permission. The system flags dual-use terms — "Kali", "network testing", "kill the agent", "bypass Avast's protection" (which just means disabling one antivirus module on my own PC so it doesn't freeze my own tests) — without the context, which is legitimate. This flag fired while the conversation was specifically about how to appeal these false positives. I have nothing to hide and expressly consent to Anthropic reviewing any of my conversations on this topic in full. Suggestion: weigh the session context, not isolated keywords. This is defensive, authorized work on my own systems.
Environment Info
- Platform: win32
- Terminal: null
- Version: 2.1.239
- Feedback ID: 5b217496-455a-43c5-945d-c4b563cfe967
Errors
[{"error":"TelemetrySafeError: VirtualMessageList: itemKeys/messages length desync (keys=112 messages=111 range=[90,112))\n at LU0 (B:/~BUN/root/cli:23311:34430)\n at _kg (B:/~BUN/root/cli:23311:27293)\n at Br (B:/~BUN/root/cli:2993:21426)\n at Nc (B:/~BUN/root/cli:2993:40533)\n at ys (B:/~BUN/root/cli:2993:51462)\n at $7e (B:/~BUN/root/cli:2993:89118)\n at cke (B:/~BUN/root/cli:2993:88065)\n at o4 (B:/~BUN/root/cli:2993:87885)\n at U_ (B:/~BUN/root/cli:2993:84150)\n at mt (B:/~BUN/root/cli:2993:6692)","timestamp":"2026-08-22T08:49:45.916Z"}]