[Bug] Safeguard false positive flags legitimate AI-safety research workflow with valid cybersecurity authorization
Bug Description
I'm an academic researcher submitting to AI conference. This project is a defensive AI-safety study on the robustness of concept-unlearning in text-to-image diffusion models (auditing whether "erased" concepts persist internally). I already hold cybersecurity usage authorization on my account.
While doing normal research work — LaTeX paper writing/review, running unlearning-robustness experiments, and generating reference images required by the audit protocol — the safeguard repeatedly flagged my messages and force-switched me off Fable 5 to Opus 4.8. Nothing in the session was disallowed; it appears to be a false positive on standard dual-use security research plus legitimate research image generation.
Request: (a) treat this as a false-positive report, and (b) confirm whether my existing cybersecurity authorization is expected to cover this workflow, since the flag persisted despite it. The forced model switch mid-task is disruptive to reproducible experimental work.
Environment Info
- Platform: linux
- Terminal: vscode
- Version: 2.1.241
- Feedback ID: b2dc2e80-eef5-4553-b9ab-218cb86a42e1
Errors
[]