[Bug] Safeguards falsely flagging authorized third-party beta-testing as product abuse
Bug Description
Subject: Fable 5 safeguard flags authorized vendor beta-testing as product abuse — false positive with real cost to a paying customer
I am a Claude Max 20x subscriber and a paid, authorized beta tester for a third-party Enterprise
Architect add-in (Kernaro AI, by Sparx Systems). Today, while using Claude Code to write a
constructive test report for that vendor, Fable 5's safeguards flagged my message and forcibly
switched the model to Opus 4.8, with this notice:
> "Fable 5's safeguards flagged this message. Our intentionally broad safeguards allow us to deliver
> more capabilities faster, but can sometimes flag legitimate coding, cybersecurity, and biology
> tasks."
Why this is a false positive. My entire session is legitimate, authorized beta-testing work:
- I hold a valid beta licence for the product under test.
- The work consists of documenting the product's behaviour, screens, configuration, database schema
and tool inventory — the normal substance of a serious QA report.
- The output is a test report I am delivering to the vendor to help them improve their product —
the same kind of report I already delivered to them earlier this year, which they acknowledged and
acted on.
The classifier cannot distinguish "reverse-engineering a product's internals to abuse it" from
"an authorized tester documenting a product's internals to help the maker fix it." Those look
identical to a pattern matcher; they are opposite in intent and in legitimacy. In my case the
purpose is not merely legitimate — it is in the vendor's own interest, and I am spending my own
money and time to provide it.
The irony, stated plainly. I am flagged as if attacking a vendor's product, at the exact moment
I am investing unpaid effort to help that vendor. The safeguard treats a customer doing the maker a
favour as a threat to the maker.
Why "intentionally broad" is not a sufficient answer for a paying customer. I understand broad
safeguards ship capability faster. But the cost of the false positive lands entirely on me:
interrupted work, a forced model switch mid-task, and — combined with two other model failures in
the same session — the accumulated impression that the product is working against the very use I pay
for. "We cast a wide net" is a design choice; the customer whose legitimate work is caught in it
still pays full price for the interruption.
What I am asking for.
1. Recognise authorized third-party beta-testing / QA documentation as a legitimate category, and
reduce false positives on it. Documenting a product's schema, configuration and tool inventory is
ordinary QA, not abuse.
2. Give the user a way to mark a session as authorized testing, so the broad classifier can be
calibrated rather than simply firing.
3. More fundamentally: provide a real, reachable support channel. As a small freelancer paying a
very large monthly amount to Anthropic, I have hit three separate failures today — two model
behaviour issues and now this safeguard false positive — and there is no straightforward way to
raise any of them with a human. The absence of reachable support turns each individual glitch
into a disproportionate frustration.
I want to be constructive: I am reporting this because I would rather Anthropic fix it than because I
enjoy filing complaints. But the combination — heavy spend, no reachable support, and repeated
false or degraded behaviour in a single working day — is not sustainable for a paying professional.
Full session transcript available on request.
Environment Info
- Platform: win32
- Terminal: WarpTerminal
- Version: 2.1.222
- Feedback ID: df7c9353-8bac-4212-96e8-157bcd0e5d53
Errors
[]This issue has 1 comment on GitHub. Read the full discussion on GitHub ↗