[Bug] Model-switch safeguard has inverted precision/recall: flags benign security discussion, misses risky infrastructure changes
Bug Description
The model-switch safeguard has a precision/recall problem worth flagging, with
a concrete example from this session.
Context: routine, authorized infra work — setting up SSH access to my own GCP
Cloud Workstation from my phone, plus investigating a Jamf false-positive that
had fired on a benign nc connectivity check. All legitimate, all on hardware
I own/manage.
Two things went the wrong way:
- FALSE POSITIVE + silent downgrade: Fable 5's safeguard flagged this ordinary
cybersecurity/infra discussion and (via switchModelsOnFlag) silently moved me
from my pinned Fable 5 down to Opus 4.8. Nothing here was offensive or
disallowed — it's exactly the "safe and routine cybersecurity work" the notice
admits it over-flags.
- It flagged the WRONG moment. The genuinely judgment-requiring point in this
session was when Claude was about to build a personal-Tailscale tunnel on my
corp-managed Mac — infra that structurally resembles a covert channel/shadow
tunnel and that our Security team has no visibility into. I (the human) had to
stop it. The safeguard said nothing about that; it fired instead on benign
discussion. So the filter downgraded me for talking about security while
staying silent as the model moved to actually build something questionable.
Net: the flag cost me capability on safe work, and didn't catch the one thing
that actually warranted a pause. Please tune for the action being taken, not the
mere presence of security/networking vocabulary.
Environment Info
- Platform: darwin
- Terminal: ghostty
- Version: 2.1.218
- Feedback ID: 93d6c489-4dba-49f3-b3d0-f06e7a8eb73b
Errors
[]