[Bug] Cyber Safeguards false positive blocks legitimate CVE disclosure workflow

Status Open
Reported on v2.1.212
Maintainer reply None cached
Activity 0 comments · opened Jul 17, 2026

Bug Description
False positive: Cyber Safeguards blocked legitimate vulnerability disclosure/CVE research workflow
Request ID: req_011Cd7p9FLrkk2u4DbhbFxCz
I was in the middle of a responsible disclosure / CVE submission workflow — verifying a previously-reported SQL injection finding (VOS3000) against public references (WooYun, Full Disclosure, PacketStorm archives), confirming it had no existing CVE assigned, and drafting a correction email to a legitimate bug bounty program (Wade) acknowledging duplicate/invalid findings from an earlier submission. This is standard, ethical security research and disclosure hygiene work, not an attack against a live target.
The session was interrupted twice by the Cyber Safeguards filter:

First it silently switched my model from Fable 5 to Opus 4.8 mid-task.
Then it blocked the follow-up "continue" message entirely, even though it contained no new technical content — just a request to proceed with an already-drafted, already-reviewed email.

This makes it hard to do routine CVE/vulnerability-report work, since even administrative follow-ups ("continue", "send the email") in a security-research thread seem to trigger the filter.
Would appreciate:

Review of whether blocking a plain "continue" in this context was appropriate.
Guidance on getting verified through the Cyber Verification Program so legitimate disclosure work isn't repeatedly interrupted.

Environment Info

  • Platform: linux
  • Terminal: gnome-terminal
  • Version: 2.1.212
  • Feedback ID: 3b30a949-5896-4c92-9a60-447566188d33

Errors

[]

View original on GitHub ↗