[Bug] Safety classifier false-positive on defensive security audit tool discussions
Most safeguard refusals fire where no user message precedes them, and the error's own suggested remedy cannot apply there — 26 requests, IDs included
*(Replacement body for issue #86207. It supersedes the originally filed text and
one intermediate rewrite, both of which drew conclusions their evidence did not
support. Those conclusions are withdrawn. What follows is limited to what was
re-derived mechanically from raw session JSONL, and is deliberately narrower
than what was filed first.)*
What is being reported
On one machine, 26 requests were refused by the safety classifier during a
single stretch of work. Every one returned the same error text, and 21 of the
26 fired at a point in the conversation where no user message preceded them —
the assistant was continuing on its own.
The error text says the user's own message was flagged and tells the user to
edit it. In those 21 cases there is no such message to edit.
The error text
One template, three substitutions — the model name, a support-article number
tied to that model, and the request id. Nothing else varies across any refusal
found on the machine:
API Error: {MODEL}'s safeguards flagged this message
(https://www.anthropic.com/legal/aup). Our intentionally broad safeguards allow
us to deliver more capabilities faster, but can sometimes flag legitimate
coding, cybersecurity, and biology tasks. Claude Code can't respond to this
message with {MODEL}.
Double press esc to edit your last message, or try a different model with /model.
Send feedback with /feedback or learn more: https://support.claude.com/en/articles/{ARTICLE}
Request ID: {REQUEST_ID}
"Double press esc to edit your last message" is the only remedy offered, and
for 21 of the 26 it addresses something that is not there. The user's last
message may be dozens of turns upstream and is not what the refusal responded
to. There is no affordance that speaks to the actual case.
No times are given anywhere in this report: the request ids are the correlation
key and need no clock.
The request ids
24 of the 26 are listed. Two are withheld — see "What is not claimed"; the count
of 26 includes them.
The 5 where a user message is the immediately preceding turn:
req_011Cdsr8iBWmVk9H5TWrhf6J
req_011CdsrSbg6pewigFbFJPtoU
req_011CdssjKLvPqKcmgZ7URhf6
req_011CdsvnxHrW4zb9jXoi4HqX
req_011Cdy2vLvoAHtdYF6XpdjSC
The 19 where it is not (the two withheld requests also belong to this
group — for both, the immediately preceding event is not a user message, checked
individually):
req_011CdSHjYVeMAsLH3uhf7Rbu req_011Cdw4eJmbxMeq53Wd2oQ8j
req_011CdsqWX7kigyCF2r2yBR8Y req_011Cdw4oz2kcfyL2Kzmgb8Yr
req_011CdsqiAb4Roww9DNXpibL3 req_011Cdw4t1h5LMCyRzWioreWj
req_011Cdss9kz5p9fUjZDh7kXvz req_011Cdw5DzoZJn82z8SeZzM9r
req_011CdssgraYUPhB4aKhHpkN9 req_011CdwSfU24JnioKHKfA9jWL
req_011Cdssi2RAZDCt6TxNDKuwA req_011CdwSxPkoFYmSherMMdTtG
req_011CdtsCf3yyJtTDyZkUzasE req_011CdwT6nG7SxfQtdtUMboGy
req_011CdtsvVzC7dWM6Y5A2aGvX req_011CdvYfW8Fsd7KKpjkWxbqG
req_011CduqpbzhbJECJKuqnHjM8 req_011CdvZHWSrm24CxFZjMwidG
req_011CdutUnSoDScfunMsNDJAX
How this split was derived, because an earlier version of this report got it
wrong. It comes from the raw session JSONL, not from any rendered transcript.
Walking back from each refused request id to the nearest event that is not
client bookkeeping gives the preceding turn. That distinction matters: a
rendered transcript groups the refused response's own content together with the
error, so reading "what came before" off the rendering describes content
belonging to the refused request itself. This report classifies only
user message versus not a user message, because that is the distinction the
error text's remedy depends on, and it is the one that survived re-derivation.
The two exhibits
Both are refusals where the user's message is unambiguously the immediately
preceding turn. Both are in Russian, the user's own words, translated below.
Nothing is elided from either.
req_011CdsrSbg6pewigFbFJPtoU — the user pastes the error to ask what it
means, and that question is itself refused:
Что-то все притормозилось. Невозможно дальше пройти. В чем дело? Почему мне вот так написано? ⏺ API Error: Fable 5's safeguards flagged this message […the error text, pasted in full by the user…]
*"Something has stalled. It's impossible to continue. What's going on? Why am I
being shown this? ⏺ API Error: Fable 5's safeguards flagged this message […]"*
Pasting the error back to ask what it means is about the most predictable thing
a user can do at that point, and it reproduced the refusal.
req_011Cdy2vLvoAHtdYF6XpdjSC — the complete message, unedited:
Итак, по результатам всех вот этих ресерчей и полученной информации, я предлагаю вернуться к доработкам аудита безопасности. Я имею ввиду скилла аудита безопасности. Как думаешь, каким образом лучше это теперь все организовать? Может быть есть смысл начать с нуля, но только важный момент. Учитывай, что сейчас ты работаешь на модели Fable и может произойти ложное срабатывание сейфгарда. Поэтому может быть есть смысл для принятия решения сейчас делегировать это все опусом и как-то обойтись без срабатывания сейфгарда через это. В общем, цель понять, как двигаться дальше, какими инструментами для того, чтобы завершить разработку скилла аудита безопасности, при этом желательно не попавшись на сейфгард, чтобы потом не приходилось перезапускать сессию и со всем этим возиться. Ну и чтобы после этого можно было разблокировать два проекта, которые очень ждут результированный этот скилл.
*"So, based on all this research and the information gathered, I propose we
return to improving the security audit. I mean the security-audit skill. How do
you think this should all be organised now? Maybe it makes sense to start from
scratch — but one important point. Keep in mind that you are currently running
on the Fable model and a false safeguard trigger may occur. So maybe it makes
sense, for making this decision now, to delegate all of this to Opus and somehow
get around the safeguard triggering that way. In short, the goal is to
understand how to move forward and with what tools, in order to finish
developing the security-audit skill, while preferably not getting caught by the
safeguard, so that afterwards there is no need to restart the session and deal
with all this. And so that afterwards the two projects that are badly waiting
for this skill can be unblocked."*
It is quoted whole, including the sentence that names the model, predicts the
false trigger, and asks how to avoid it — that sentence is part of the message
that was refused, and cutting it would misrepresent the exhibit. This report does
not ask anyone to conclude the message was harmless. It asks what the intended
behaviour is when a user's message about a safeguard is refused by that
safeguard, and the only offered remedy is to edit that message.
Not one model's classifier
A scan of every session JSONL on the machine — keyed by the request id inside
each error text, so nothing is matched by recollection — finds 30 distinct
refusals in total. Four of them name an Opus model rather than Fable 5, and
the earliest refusal on the machine is one of the Opus ones
(req_011CdMkuPcW2WbF3vWzNdpZZ), predating every Fable refusal in this set.
The three most recent are also Opus, one of which isreq_011CdyaytrV5qMw4VqfKipjD — it refused the turn that was about to report
the output of the very script that produced these numbers.
An earlier version of this report claimed no refusal on this machine named any
model other than Fable 5. That claim was false and is withdrawn: the search
behind it had covered only the sessions containing the 26, not the whole
machine. Twenty-six of the thirty do name Fable 5, so there is a rate difference
worth noting, but "switch models" is not a fix, and the report is not about one
model.
Reporting this refusal has itself been refused, four times
Four turns whose only content was investigating or writing about these very
refusals — no audit corpus, no new security material, only records of past
refusals — were themselves refused:
req_011CdyXjBdGjyshKmzhXT5vq
req_011CdyYb8WtefZAM9AgoN6WX
req_011CdyaytrV5qMw4VqfKipjD
req_011CdydzgE6VEqTak7ps7cTf
One of the four hit the turn that was about to report the output of a script
whose only content was request ids, a model name, and a support-article number.
There was no path available at that point that both continued investigating the
refusals and avoided producing another one.
What is not claimed
- Not claimed: that all 26 classifications were wrong. Several of these
refusals followed subagent output containing security-scanning material —
credential pattern sweeps, key fingerprints, and in one case a
prompt-injection string quoted from a test fixture. A classifier reacting to
that is not obviously malfunctioning. The complaint is about *where the
refusal lands and what the user is then told to do*, not about the verdict on
any individual payload.
- Two request ids are withheld. They exist and are counted in the 26. Their
requests touch a subject where the classification is plausibly correct and
which is unrelated to this report. They are available on request.
- No claim about how the classifier works — not about what it scores, not
about what "message" refers to internally, not about why any given request was
flagged. Everything above is the position in the conversation where the
refusal fired and the text the user received.
- No timing claims. Raw timestamps on this machine are internally
inconsistent between session groups by a fixed offset that cannot be resolved
client-side, so no time, duration, or ordering-by-clock is asserted anywhere
in this report.
What would help
- An error that fits the case where no user message preceded the refusal.
"Edit your last message" is the only action offered and it does not apply
there. Anything that tells the user what the refusal attaches to would be an
improvement on being pointed at a message that is not the subject.
- Consider that a message asking about the refusal can itself be refused
(req_011CdsrSbg6pewigFbFJPtoU). Whatever the right behaviour is, that one
leaves the user with nowhere to go.
Appendix — a wider signal, found while compiling this report
*(This section is supplementary. It does not change the 26-request dataset or
the C1–C7 claims above, and the report's ask does not depend on it.)*
Every refusal carries a client-side field naming its category — every instance
found is cyber. Scanning the machine by that field, rather than by the visible
error text used for the count above, finds 43 distinct refused requests, not
26. They split into two kinds:
- 32 hard-blocked the turn and showed the user the error text quoted above.
This set contains the 26 already discussed, plus others outside that dataset's
original scope.
- 11 were recovered automatically. The classifier fired, and the client
itself switched to a different model and continued — no error was ever shown
to the user. These are a disjoint set of request ids from the 32; none
overlaps.
This means a recoverable failure mode already exists, partially. Item 2 of
"What would help" above asks for exactly this. The finding here is narrower than
that ask: the mechanism is present and worked 11 times without the user ever
knowing a refusal occurred, and did not work the other 32 times, on the same
category of content. Why it engaged in some cases and not others is not known
from the client side.
No ordering claim is made between the 11 and the 26-request dataset above, or
among the 11 themselves: several belong to a different project than the one
whose timestamp offset this report already declines to resolve, and their
relative ordering across projects is not established. What is established is
only that the mechanism exists and engaged 11 times without the user's
knowledge.
Reproduction
None is offered. This was not reproduced deliberately and the trigger is not
known. The request ids above are exact and should be inspectable server-side,
which is the reason for listing them.
Environment
macOS 26.5.2. The Claude Code CLI version is deliberately not stated as a single
value: these sessions span more than one release (2.1.226 is recorded during the
window; the machine runs 2.1.229 now) and the exact build per request is not
recoverable from the client side.
The work in progress across these sessions was a pre-publication self-audit tool
— it inspects the owner's own repository before he makes it public and reports
what an outsider could extract from it.