[Feature Request] Anthropic API Safety Filters Cause False Positives for Agent Harness Development

Status Open
Reported on v2.1.216
Maintainer reply None cached
Activity 0 comments · opened Jul 21, 2026

Bug Description
I am developing a developer productivity tool/plugin called "oh-my-grok" (https://github.com/ImL1s/oh-my-grok) for the Grok Build CLI. It is a workflow discipline harness that enforces structured cycles (such as explore-implement-verify loops and Ralph loops) and intercepts tool calls to validate outputs.
Fable 5's safeguards consistently trigger false positives and downgrade my session to Opus 4.8 whenever I edit code containing tool hooks, CLI command validation, state machines, or agent prompt enforcement. This is routine developer tool software engineering, not malware or prompt injection. Please adjust the filters to allow safe development of agent harness and CLI plugin code.

Environment Info

  • Platform: darwin(macOS 26.4)
  • Terminal: tmux(Warp)
  • Version: 2.1.216
  • Feedback ID: a807b270-e8d6-4f90-b1be-896fbd1819cc

Errors

[]

View original on GitHub ↗