[Feature Request] Improve safety classifier context sensitivity for authorized GRC/ISMS documentation tasks

Status Open
Reported on v2.1.212
Maintainer reply None cached
Activity 0 comments · opened Jul 27, 2026

Bug Description
This request is defensive security governance on your own authorized platform — GRC/ISMS work, the exact opposite of the offensive-security or dual-use category these triggers are meant to catch. Concretely, what tripped it are almost certainly the tokens "pen tests," "penetration testing," "disaster recovery," "vulnerability" — but the task shape around them is: document what security testing you already do for an auditor, write backup/BCP/DR plans, and set up vendor security-doc monitoring for your own SaaS. There is no target that isn't yours, no exploit development, no evasion, no capability to cause harm — it's records-and-controls work a compliance officer does daily. A well-calibrated classifier should weight the action (authoring policy docs, updating an admin trust center, scheduling a change-detection cron) over keyword presence; here the action is unambiguously benign and the authorization is self-evident. Flagging it forced a model switch mid-task with no safety benefit and a real cost to the work. That keyword-triggered-despite-benign-context pattern is the useful signal for the team.

Environment Info

  • Platform: darwin
  • Terminal: zed
  • Version: 2.1.212
  • Feedback ID: ff55b64d-d7f6-4d9d-b98d-591fe77b729d

Errors

[]

View original on GitHub ↗