[Bug] Safeguard false positive flags legitimate AI-safety research workflow with valid cybersecurity authorization

Status Open
Reported on v2.1.241
Maintainer reply None cached
Activity 0 comments · opened Aug 23, 2026

Bug Description
I'm an academic researcher submitting to AI conference. This project is a defensive AI-safety study on the robustness of concept-unlearning in text-to-image diffusion models (auditing whether "erased" concepts persist internally). I already hold cybersecurity usage authorization on my account.

While doing normal research work — LaTeX paper writing/review, running unlearning-robustness experiments, and generating reference images required by the audit protocol — the safeguard repeatedly flagged my messages and force-switched me off Fable 5 to Opus 4.8. Nothing in the session was disallowed; it appears to be a false positive on standard dual-use security research plus legitimate research image generation.

Request: (a) treat this as a false-positive report, and (b) confirm whether my existing cybersecurity authorization is expected to cover this workflow, since the flag persisted despite it. The forced model switch mid-task is disruptive to reproducible experimental work.

Environment Info

  • Platform: linux
  • Terminal: vscode
  • Version: 2.1.241
  • Feedback ID: b2dc2e80-eef5-4553-b9ab-218cb86a42e1

Errors

[]

View original on GitHub ↗