False positive: Fable 5 safeguards flag defensive security / prompt-injection hardening
Status Closed — not planned
Maintainer reply None cached
Activity 1 comment · opened Jul 5, 2026 · closed Aug 30, 2026
Fable 5's safeguards flagged a routine defensive-security request and auto-switched to Opus 4.8.
The request: harden a container I own and add prompt-injection defenses to my own automation that processes third-party social-media comments. Purely defensive, blue-team work on my own infrastructure — treating external text as data rather than commands, and closing an exposed local admin port.
This is exactly the authorized, defensive use case the model should support. Flagging it discourages users from protecting their own systems. Please recalibrate so that defensive hardening and prompt-injection mitigation on one's own infrastructure are not treated as dual-use risk.
This issue has 1 comment on GitHub. Read the full discussion on GitHub ↗