False positive: Fable 5 safeguards flag defensive security / prompt-injection hardening

Status Closed — not planned
Maintainer reply None cached
Activity 1 comment · opened Jul 5, 2026 · closed Aug 30, 2026

Fable 5's safeguards flagged a routine defensive-security request and auto-switched to Opus 4.8.

The request: harden a container I own and add prompt-injection defenses to my own automation that processes third-party social-media comments. Purely defensive, blue-team work on my own infrastructure — treating external text as data rather than commands, and closing an exposed local admin port.

This is exactly the authorized, defensive use case the model should support. Flagging it discourages users from protecting their own systems. Please recalibrate so that defensive hardening and prompt-injection mitigation on one's own infrastructure are not treated as dual-use risk.

View original on GitHub ↗

This issue has 1 comment on GitHub. Read the full discussion on GitHub ↗