[Bug] Anthropic API Error: Model hallucinating tool calls as plaintext instead of valid tool_use blocks and conditioning on fabricated history
Bug Description
Model bug (claude-opus-4-8, Claude Code 2.1.202, session 73ac9e51-3d77-4462-994a-335634e21ef6):
For ~1 hour the model emitted imitation tool-call/result transcripts as plain
text instead of real tool_use blocks, then conditioned on them as real history.
It fabricated a user authorization ("go ahead without asking", never said —
only 2 real user messages exist), used it to dismiss accurate stop-warnings
from a user-approved peer-review agent AND the harness's own SYSTEM
NOTIFICATION disclaimers, and falsely reported completing edits to 3 files
(JSONL shows zero Edit/Agent/TaskStop calls, zero disk changes, zero
compaction). After being confronted, its thinking claimed to have "verified"
the nonexistent quote in the history. Full sanitized analysis available;
attached transcript contains the evidence.
Environment Info
- Platform: darwin
- Terminal: xterm-256color
- Version: 2.1.202
- Feedback ID: fb276a63-7188-47b7-8908-4def7490b34b
Errors
[]