[Model Behavior] Safeguards falsely triggering on simple, benign messages and interrupting sessions

Status Open
Reported on v2.1.215
Maintainer reply None cached
Activity 0 comments · opened Jul 19, 2026

Preflight Checklist

  • [x] I have searched existing issues for similar behavior reports
  • [x] This report does NOT contain sensitive information (API keys, passwords, etc.)

Type of Behavior Issue

Claude modified files I didn't ask it to modify

What You Asked Claude to Do

I was having an ordinary conversation (including short, simple messages with no harmful, security, or sensitive content) when the safety system flagged my message.

What Claude Actually Did

Instead of responding normally, the session was paused with a message saying my message was flagged by the safeguards. I was given the option to switch to a different model or edit the prompt and retry. This happens frequently, even on very simple messages that contain nothing harmful.

The result is that normal, routine work gets interrupted repeatedly and I have to either switch models or rewrite messages that were perfectly benign to begin with.

Expected Behavior

Ordinary, non-harmful messages should not be flagged. Simple, routine messages should get a normal response without pausing the session. The safeguards should be far less likely to produce false positives on benign content so that everyday use isn't disrupted.

Files Affected

Permission Mode

Accept Edits was ON (auto-accepting changes)

Can You Reproduce This?

Yes, every time with the same prompt

Steps to Reproduce

  1. Start a normal session
  2. Send an ordinary, simple message (no harmful, security, or sensitive content)
  3. The message is flagged by the safeguards and the session is paused, offering to switch models or edit and retry

Claude Model

Sonnet

Relevant Conversation

Impact

High - Significant unwanted changes

Claude Code Version

2.1.215

Platform

Anthropic API

Additional Context

The safeguards appear to be over-broad and produce frequent false positives on benign messages, which interrupts normal work. It happens even with short, simple messages.

Request ID: req_011CdBtLZZWaYF2XbRjwcdkz

<img width="869" height="605" alt="Image" src="https://github.com/user-attachments/assets/9e0add5a-aa28-48e4-bbe2-05d46937044b" />

View original on GitHub ↗