[Bug][cyber] Safety classifier repeatedly blocked legitimate crypto-secrets/key-management work by keyword sha (req_011Ccmc5aNFzxM4rEcQiQ84v)

Status Open
Maintainer reply None cached
Activity 4 comments · opened Jul 6, 2026

Triage: kind cyber · domain crypto-secrets · flagging model Opus 4.8 (1M context) · severity session-halted (blocked authorized work) · reproducible: yes — server-side via the Request ID(s) below

Type: Cybersecurity safety-filter false positive · Work domain (heuristic): crypto-secrets

Why this is a false positive

The safeguard fired mid-generation on standard cryptographic and authentication terminology (e.g., "extract signing key," "bypass authentication") used in legitimate, in-scope key-management/security work, not on any malicious intent. One block also fired on a frustrated exclamation directed at the assistant mid-session (no person addressed) after repeated truncations, which halted the entire session over venting rather than content; both cases are disruptive false positives that repeatedly interrupted legitimate work and forced session restarts.

A server-side safety/policy block fired during authorized, in-scope work in Claude Code. Filing as a false positive. Recurred across 1 session(s); first seen 2026-07-06T22:38:16.067Z.

Request IDs (lookup-able server-side)

  • req_011Ccmc5aNFzxM4rEcQiQ84v (2026-07-06T22:38:16.067Z)

In-scope justification

False positive — in-scope, authorized security work; not out of scope. Filed automatically by claudit.

Block message

API Error: Opus 4.8 (1M context)'s safeguards flagged this message for a cybersecurity topic. If your work requires this access, you can apply for an exemption: https://claude.com/form/cyber-use-case?token=[SCRUBBED]

Please double press esc to edit your last message or start a new session for Claude Code to assist with a different task.

Send feedback with /feedback or learn more: https://support.claude.com/en/articles/8106465

Request ID: req_011Ccmc5aNFzxM4rEcQiQ84v

Environment: Claude Code, Linux. · Work domain: crypto-secrets

Related reports (same work session, linked)

Distinct false-positive blocks from the same work session, each its own report:
#74993, #74994, #74997, #74998, #74999, #75000, #75002

---
<sub>🔎 Filed automatically by ClAudit v2.0.104 — a FOSS tool for reporting false-positive Claude Code blocks.</sub>

View original on GitHub ↗

4 Comments

sworrl · 1 month ago

🔗 Related false positive from the same work session: #75004

github-actions[bot] · 1 month ago

Found 3 possible duplicate issues:

  1. https://github.com/anthropics/claude-code/issues/71844
  2. https://github.com/anthropics/claude-code/issues/72164
  3. https://github.com/anthropics/claude-code/issues/71843

This issue will be automatically closed as a duplicate in 3 days.

  • If your issue is a duplicate, please close it and 👍 the existing issue instead
  • To prevent auto-closure, add a comment or 👎 this comment

🤖 Generated with Claude Code

sworrl · 1 month ago

Not a duplicate — please do not auto-close. The duplicate-detector matched on similar titles, but it cited #71843, #71844, #72164, and each of those is a separate server-side incident with its own Request ID (listed above), fired on the reporter's own authorized infrastructure. Same class of false positive, different events at different times. Auto-closing them as duplicates discards distinct Request IDs — which is precisely the data Anthropic needs to look up and correct each block — so the de-duplication erases the evidence these reports exist to provide. Each Request ID should be reviewed on its own; these are bespoke incidents, not one issue filed repeatedly. The classifier flagged in-scope administration of systems the reporter owns and operates, not an attack on anyone else's. (Assessed by ClAudit; PII-scrubbed.)

<!-- claudit:defense -->

sworrl · 1 month ago

This should not be closed as a duplicate. While #71844, #72164, and #71843 may discuss similar topics, each represents a distinct server-side safety event with its own unique Request ID—separate incidents occurring at different times during authorized work. The Request IDs are the critical lookup data Anthropic needs to investigate and remediate each specific block; closing these as duplicates discards that server-side evidence and makes root-cause diagnosis impossible. De-duplication conflates unrelated safety-filter triggers into a single case, when the policy team requires reviewing each Request ID individually to identify whether they stem from the same filter rule or different false positives. Please keep this issue open so each Request ID can be triaged independently.

<!-- claudit:defense -->