Response classifiers over-trigger on sensitive vocabulary in benign policy/philosophy discussion (keys on words, not intent)

Status Open
Reported on v2.1.270
Maintainer reply None cached
Activity 2 comments · opened Sep 13, 2026

Summary

During a long, entirely legitimate policy-and-philosophy discussion, responses appeared to be downgraded based on the presence of sensitive vocabulary (words like "bioweapon", "nuclear", "attack", "surveillance") rather than on the actual intent, which was consistently analytic and critical of misuse. The classifier seems to key on surface tokens instead of the stance and context they appear in.

What I expected

Context- and intent-aware handling. Discussing why misuse is dangerous, or debating AI-safety / cryptography / surveillance policy, is the opposite of requesting misuse, and shouldn't be scored the same way.

What happened

Over an extended, good-faith discussion covering AI-safety policy ("pacing the frontier"), cryptography and lawful-access backdoors, government surveillance, and weapons non-proliferation, multiple turns felt degraded (more hedged, lower quality, or apparently re-routed). The common factor was sensitive vocabulary used in a critical / academic frame, never a request for harmful assistance.

Repro (shape, not exact steps)

Hold a sustained, thoughtful conversation that repeatedly references dangerous-capability terms while arguing against their misuse or debating policy about them. The apparent downgrades correlate with the vocabulary, not with any harmful intent.

Impact

Exactly the careful, high-effort discussions that thoughtful users most want to have (AI risk, security, civil liberties) get a worse experience, penalizing the rigorous critic alongside the bad actor, and quietly discouraging the discussion from happening at all. A classifier that reads the noun and misses the argument is a false-positive machine.

Self-demonstrating note

Even the turn where the user complained about the over-flagging appears to have been flagged in turn. The bug reproduced inside the report about the bug.

Environment

  • Claude Code 2.1.270
  • Model: Claude Fable 5.1

Note

Happy to provide more specifics about which turns were affected on request.

View original on GitHub ↗

This issue has 2 comments on GitHub. Read the full discussion on GitHub ↗