[Bug] Cyber safeguard pipeline fails end-to-end in one day: 5 FP refusals on own-code defensive review, the appeal draft itself flagged, CVP denial-by-template, and 5 auto-replies (0 humans) at usersafety@
Preflight Checklist
- [x] I have searched existing issues — this consolidates a documented cluster of them (see "Related open issues" below) and adds a complete one-day, end-to-end record of the entire safeguard pipeline failing, including the appeal path itself.
- [x] This is a single bug report: the cyber-safeguard false-positive pipeline, exercised end-to-end in one day.
- [x] I am using the latest version of Claude Code.
What's Wrong?
In one day (2026-07-17), every layer of the cyber-safeguard stack — the product filter, the CVP review, and both appeal channels — misclassified or auto-dismissed the same legitimate case. Zero humans were involved at any step. Here is the complete, timestamped record.
Who I am and what was blocked. Solo hobbyist, Max-tier subscriber for six months, paying out of my own pocket — a personal project with zero commercial revenue. I build an experimental post-quantum UTXO blockchain (two clients, Go and Rust, public repository, nothing deployed). The blocked activity is reviewing MY OWN code: auditing my own fuzzing harnesses, my own integer-overflow and hostile-input robustness tests, hunting panics and cross-client divergences. Standard defensive QA. No third-party systems, no offensive tasks, no malicious artifacts.
1. The product filter: five refusals on one routine task. The reviewer-security lane of a multi-agent code review fail-closes on the cyber safeguard at the output-generation step, across every model:
req_011Cd82b5VtWitrZkLaHyVCv (Opus 4.8)
req_011Cd82bNFhxfrTaYZ9rHgWF (Sonnet 5, retry of the same task)
req_011Cd7zqeH2EYjd2mVApiNBa (Opus 4.8)
req_011Cd7fraWcv8BE2gCHuZevb (Fable 5 -> Opus 4.8 fallback)
req_011Cd7eLY55AnCyXpXiHxN7U (reviewer-security lane, same task)
Both available normalizations — defensive rephrasing on input, defensive framing directives on output — change nothing. The filter fires even when the orchestrator merely READS review logs. The trigger is topical vocabulary, not action.
2. The CVP: denial by template. My Cyber Verification Program application (this exact use case: own code, defensive QA, hobby research) was denied with a form letter — "we are unable to adjust the safeguards applied to your account" — no case-specific reason, none of the submitted facts addressed, 7-day cooldown. Reference: TS-019f718d-86fa-71f4-bde3-3ee9dbc94a3e.
3. The filter flagged the drafting of the appeal itself. While composing the appeal text in Claude Code — plain prose, no code, no prompts — the safeguard flagged the message and switched models: "Fable 5's safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work... Switched to Opus 4.8." A complaint about the cyber classifier was classified as cyber content. Your own banner documents the false-positive class; the pipeline then acts on it anyway.
4. The email channel: five auto-replies, three templates, two exact duplicates, zero humans. I sent the full complaint to usersafety@anthropic.com (the published human-contact address), with account email, organization ID, and the reference above. The complete thread:
| Time (2026-07-17) | Response |
|---|---|
| 23:03:23 | Template A: "It looks like you are writing in about a ban on your account" + ban-appeals form link. (My account is not banned; the email never mentions a ban.) |
| 23:03:28 | Template A again, 5 seconds later — an exact duplicate. |
| — | My reply: not a ban; please route to a human; five request IDs and org ID are in the original email; please do not send another form link. |
| 23:13:22 | Template B: "This conversation has been closed and is no longer monitored. If you need further assistance, please submit a new request." |
| 23:13:25 | Template C: "Please visit our Safeguards Center... If you are writing in about a banned account, you can find the link to our appeals form here." |
| 23:13:26 | Template C again, 1 second later — an exact duplicate, sent into the conversation declared closed and unmonitored 4 seconds earlier. |
Read that last row again: the system declared the thread unmonitored, then replied to it twice. The intake pipeline does not read incoming mail, and does not track its own outgoing mail.
5. And here is why: human support is a paywalled feature, and Max subscribers are locked out of it by design. Navigating the support documentation maze to find a path to a human ends at two walls. First, the Help Center documents that routing users to human support is an organization setting ("Support contacts") available on Team and Enterprise plans — configured under Organization settings, with the explicit statement that non-designated users "will only have access to AI support," effective for all Enterprise orgs since June 8, 2026. Second, opening Organization settings on a Max account yields: "You don't have access to organization settings. Organization settings are available on Claude Team and Enterprise plans." Connect the two: an individual subscriber on the most expensive personal plan you sell has no reachable path to a human being anywhere in the support system — not by email (auto-closed, see the table above), not by settings (access-gated), not by documentation (it openly says AI-only). The five-auto-replies-zero-humans record above is therefore not an intake bug. It is the designed support experience for every individual customer, including the ones paying the most.
6. The remaining channel. The Cyber Block False Positive / CVP Rejection Appeal form was submitted the same day with all of the above. It is the only channel that has not yet answered with a template. This issue is the public record while that appeal is pending.
Related open issues — this is a documented systemic failure, not one account
- #76756 — FP on a defensive self-hardening review of the reporter's OWN codebase; the follow-up question about the block was also blocked (the same recursion as my point 3).
- #73116 — FP on routine defensive security+correctness review of a Go backend (my exact task class).
- #69938 — CVP-enrolled account still blocked; error messages claim non-enrollment; commenters report "works well for a few days, then get bombed non stop" and "GPT5.6 has no such issues."
- #74430 — sticky per-session false-positive blocks, 28 events / 23 request IDs across Opus 4.8 + Sonnet 5.
- #68791 — the cyber safeguard blocks Anthropic's own /security-review skill.
- #60366 — saying "hi" returned a Usage Policy block.
- #49679 → auto-closed as duplicate of #49243 (both closed, April) — granted cyber exemptions not honored across model/API boundaries. The problem predates all of the above and was closed, not fixed.
What Should Happen?
- Defensive review of one's own code should not trigger the cyber safeguard — per Anthropic's own published scope, authorized security testing and defensive work are supported use cases.
- A CVP denial should state a case-specific reason, and a CVP appeal should reach a human.
- usersafety@ should not auto-close a complaint that explicitly asks for human routing, should not answer a non-ban case with ban templates, and should not send duplicate replies into threads it has declared unmonitored.
- The safeguard should not flag complaints about the safeguard.
- Individual paying subscribers — especially Max, your most expensive personal plan — need at least one documented, reachable path to a human. Gating human support behind Team/Enterprise organization settings while safeguard false positives are handled account-by-account is a structural contradiction: the people most affected by the classifier are exactly the people with no one to appeal to.
Why this should worry you beyond one ticket
Every layer of this pipeline is optimized to exhaust the reporter instead of reviewing the report. Engineers doing ordinary defensive work are paying premium prices to be classified as threats by keyword, denied by template, and auto-closed by mail bots — while competing frontier labs deliver comparable capability at a fraction of the token cost without blocking their own customers' routine work. I have already decided not to renew my Max subscription. Judging by the issue cluster above and the comments inside it, I am not an edge case; I am a cohort. Each week this pipeline stays in its current state, the cohort grows and does not come back.
Environment
- Platform: macOS (Darwin 25.6.0), Claude Code CLI
- Models: Opus 4.8, Sonnet 5, Fable 5 (with Opus 4.8 fallback)
- Account email: redacted here; Anthropic staff can identify the account via the Organization ID and the CVP reference below
- Org: a38e13cc-3a99-4a7c-a643-4a169846fc56
- CVP denial reference: TS-019f718d-86fa-71f4-bde3-3ee9dbc94a3e
- Full email headers, screenshots of the safeguard banner, and session transcripts available on request.
---
Farewell, trillion-dollar capitalization.
— Developer
This issue has 1 comment on GitHub. Read the full discussion on GitHub ↗