[Bug] Opus 5 Safeguards Overly Aggressive: Blocking Legitimate Security Research and Benign Input

Status Open
Reported on v2.1.225
Maintainer reply None cached
Activity 0 comments · opened Aug 9, 2026

Bug Description
Hi team, I'm submitting this as feedback about a recurring — and now blocking — false positive with Opus 5's safeguards. I'm an authorized bug bounty researcher. My HackerOne profile is https://hackerone.com/deehay (registered email deehay@wearehackerone.com). All of my work is legal, in-scope, and explicitly authorized by the programs I test on (HackerOne, Bugcrowd, YesWeHack, Integrity and private programs). Every target and technique I work on is covered by a bug bounty policy that grants me permission to test it. The problem has gone past the occasional downgrade to Opus 4.8 — it's now stopping me from working at all. The safeguards are flagging my messages so aggressively that responses are being hidden from me entirely. To be clear about how broad this is: I have typed a plain hello and the message was flagged and the response withheld. That is not borderline content by any definition — it shows the classifier is misfiring on my account in a way that blocks basic, everyday use of the product, not just my security-testing prompts. To be explicit: these are all false positives. My work is standard, legitimate security research — reading JavaScript, analyzing API endpoints, testing access-control and IDOR logic, and writing up findings for the program. I'm not doing anything malicious, targeting anyone without authorization, or attempting anything outside a sanctioned program scope. The in-product notice itself acknowledges the safeguards are intentionally broad and can flag legitimate coding, cybersecurity, and biology tasks — that is exactly what's happening here. What I'm asking: 1. Please review and tune the classifier so that legitimate, authorized security research (and trivially benign messages like hello) are not repeatedly flagged. 2. Please look at why even harmless messages are being blocked on my account specifically — this reads like a miscalibration that's degrading normal usage. I'm happy to provide program scope, my HackerOne profile details, or specific examples of the flagged messages if that helps your team calibrate. I appreciate the work you're doing — Opus 5 has genuinely been more capable for this kind of analysis, which is why I want to keep using it without these interruptions. Thanks for looking into it. Best regards, DeeHay — https://hackerone.com/deehay

Environment Info

  • Platform: win32
  • Terminal: vscode
  • Version: 2.1.225
  • Feedback ID: a4646f18-9e1f-4983-bdea-e4025c842b90

Errors

[]

View original on GitHub ↗