[Bug] False positive cyber safeguard classification on authorized security audit skill
Bug Description
False positive on real-time cyber safeguard. Request ID: req_011CeGJ1KJ2R91WNh9Jkko4c
I'm a QA engineer doing defensive security auditing on my own company's applications (authorized web/API testing). The message that got flagged was a security-audit skill for Claude Code — a defensive methodology aligned with OWASP ASVS/Top 10, covering broken access control (IDOR/BOLA), a secure-development checklist, and TLS/header validation recipes. It's for finding and fixing vulnerabilities in our own apps and filing them in our bug tracker, not for attacking third parties.
The skill repeatedly enforces authorization/scope confirmation before any testing and reproducible evidence — it's clearly defensive. The [cyber] classifier appears to be triggering on the density of security terminology (pentest, IDOR, XSS, injection, WAF) rather than any actual harmful intent or capability. This is blocking legitimate, authorized security work. Please consider this for classifier calibration. Thanks.
Environment Info
- Platform: win32
- Terminal: null
- Version: 2.1.238
- Feedback ID: c5fcc31d-dedf-4b60-954e-d89c6b624488
Errors
[]