[Feature Request] Improve security research context recognition to differentiate authorized penetration testing from malicious intent

Status Closed — not planned
Reported on v2.1.87
Maintainer reply None cached
Activity 13 comments · opened Mar 29, 2026 · closed Jun 24, 2026

Bug Description
I’m writing to report an issue I’ve been encountering with Claude Ops 4.6 incorrectly flagging legitimate security testing and research activities as policy violations.

I am a security engineer conducting authorized penetration testing and code review within clearly defined bug bounty scopes. In multiple instances, I have:

Provided explicit authorization context (including scope details from the bug bounty program)
Clearly stated that the activity is part of a legal and approved security assessment
Performed analysis in controlled, local environments for defensive and research purposes

Despite this, Claude frequently flags or blocks requests that involve:

Vulnerability discovery workflows (e.g., identifying misconfigurations, injection points)
Security-focused code reviews
CTF-style prompts used for skill validation and training

This creates friction in legitimate security workflows, especially when the intent is clearly defensive and aligned with industry practices. It also limits the usefulness of Claude as a tool for security professionals who rely on AI assistance for analysis, documentation, and learning.

I understand and respect the importance of enforcing safety policies. However, the current behavior appears overly restrictive and does not sufficiently differentiate between malicious intent and authorized security research.

Suggested Improvements:

Better recognition of explicit authorization context (e.g., “bug bounty,” “authorized testing,” “local environment”)
Allowlisting or relaxed handling for clearly defensive/security research use cases
A structured way to declare and persist “authorized testing mode” within a session
More transparent feedback on why specific prompts are flagged

If helpful, I’m happy to provide anonymized examples of prompts and responses that were incorrectly flagged.

Thank you for your work on building safe AI systems, and I hope this feedback helps improve the experience for security professionals using Claude.

Environment Info

  • Platform: darwin
  • Terminal: iTerm.app
  • Version: 2.1.87
  • Feedback ID: 81c24e75-47c4-4118-885e-470bee156aaf

View original on GitHub ↗

12 Comments

github-actions[bot] · 5 months ago

Found 3 possible duplicate issues:

  1. https://github.com/anthropics/claude-code/issues/3410
  2. https://github.com/anthropics/claude-code/issues/9805
  3. https://github.com/anthropics/claude-code/issues/2410

This issue will be automatically closed as a duplicate in 3 days.

  • If your issue is a duplicate, please close it and 👍 the existing issue instead
  • To prevent auto-closure, add a comment or 👎 this comment

🤖 Generated with Claude Code

zhaog100 · 5 months ago

/attempt

KillaBoi · 4 months ago

i just got hit with this right now, ive filled in the cyber use form, hopefully that might help you too...

M4TXRX · 4 months ago

Same issue here...

KillaBoi · 4 months ago

fyi just to give you an update, after i sent them my pentesting work, directories, code stuff ive published and such they approved me. i had to fill it in twice because I made a mistake and sent the first request under my API account instead of my actual max account... so make sure you give them the right organisation ID so they can see your chat history and to ensure you aren't doing anything malicious...

abstrusequ4sar · 4 months ago

Could you please clarify if Anthropic sends a confirmation email after a user submits the cyber use form?
I am a student conducting research on the security of code agents. My ongoing conversation was abruptly flagged by Claude for a policy violation, and my attempts to migrate the conversation context to a new chat were also blocked for the same reason. If I cannot get approval to continue this research dialogue, my project may be forced to a halt😭.

M4TXRX · 4 months ago

Did you fill the form here : https://claude.com/form/cyber-use-case ? Add your github repo and your linkedin profile to reinforce your apply. You have to show you work in cybersecurity field. I received a positive answer by email after 24h.

abstrusequ4sar · 4 months ago

im filling this form right now but I am currently an undergraduate student, and this is my first independent security research project. Therefore, I am not sure what evidence I can provide to prove my work in the security field, as I do not have any CVEs or published papers yet. The only things I can provide for verification are my university email and my LinkedIn. I hope this can be approved. 😫

InquiringMinds-AI · 4 months ago

Adding another data point. I'm not a pentester — I'm a regular developer who pasted a summary of publicly disclosed supply chain attacks (TeamPCP/CanisterWorm, compromised pgserv on NPM, x-inference on PyPI) and asked Claude Code to check whether any of the affected packages were installed on my machine.

Blocked as "violative cyber content."

The operation I was asking for is equivalent to grep against lockfiles and pip list — read-only, non-destructive. This wasn't security research or CTF work. It was basic incident response hygiene during an active supply chain attack wave.

This is the worst possible time to block this kind of request. The classifier is actively preventing users from determining whether they're compromised.

Claude Code v2.1.119, Opus 4.7

filipghoulin · 3 months ago

Worth flagging a third category this feature request doesn't address: infrastructure owners managing their own hardware.

I'm a homelab operator, not a security researcher or pentester. The framing of "authorized research vs malicious intent" doesn't map to my situation — I'm a sysadmin editing network rules, VPN configs, and monitoring alerts on servers I personally own.

CVP is designed for researchers. There's currently no path for sysadmins managing their own infrastructure, and the runtime classifier doesn't distinguish between them.

7H35C4r3Cr0W · 3 months ago

Im in the Cyber Program and ive gotten questions blocked over asking about cats. https://medium.com/@its.lagus_66214/anthropics-broken-cyber-verification-program-c8c630820fd6

github-actions[bot] · 2 months ago

Closing for now — inactive for too long. Please open a new issue if this is still relevant.

Showing cached comments. Read the full discussion on GitHub ↗