[Bug] Fable safeguard false positive flags safety refusals, causing costly mid-session routing to Opus

Status Open
Reported on v2.1.217
Maintainer reply None cached
Activity 0 comments · opened Jul 23, 2026

Bug Description
Fable safeguard false positive, cost me ~$5.65 in wasted context: Fable steered my ML session toward a security-adjacent idea. I'm the one who recognized the risk and refused it. Despite being the safety refusal in the conversation, I got flagged and routed to Opus mid-session, torching paid context to rebuild state. The routing is keyword-based and trajectory-blind — it flagged the person doing the refusing. Please make it account for conversational trajectory, and don't pass false-positive handoff costs on to users.

Environment Info

  • Platform: darwin
  • Terminal: iTerm.app
  • Version: 2.1.217
  • Feedback ID: 184b4f80-adab-4ef2-a0cd-f4021a3265b5

Errors

[]

View original on GitHub ↗