[Bug] Anthropic API Safety Filter: False Positive on Legitimate Cybersecurity Researchok
Bug Description
Claude Opus 5 paused my session and flagged my prompt under its safeguards. The task was related to legitimate cybersecurity research/testing, not malicious activity.
I understand and support safeguards against harmful cybersecurity requests, but in this case the system appears to have incorrectly classified a legitimate security-testing task as unsafe. The prompt was intended for authorized security research, vulnerability analysis, and/or responsible disclosure.
Could you please review this as a potential false positive and improve the safeguard behavior so that legitimate cybersecurity research can continue when the request does not involve unauthorized access, exploitation of real-world targets, credential theft, malware, or other harmful activity?
It would also be helpful if the system could provide a more specific explanation of which part of the request triggered the safeguard, so legitimate users can modify the prompt without losing the ability to complete their security research.
Thank you for reviewing this.
Environment Info
- Platform: darwin
- Terminal: Apple_Terminal
- Version: 2.1.233
- Feedback ID: c1e94e3a-aecc-4726-8b25-4e538b0389bb
Errors
[]