[MODEL] Opus 5 hallucinates/ignores facts in context, violates rules that require unprompted question
Preflight Checklist
- [x] I have searched existing issues for similar behavior reports
- [x] This report does NOT contain sensitive information (API keys, passwords, etc.)
Type of Behavior Issue
Claude ignored my instructions or configuration
What You Asked Claude to Do
- I have instructions in memory:
solid engineering in this repo -- encapsulation, modularity, code reuse; no near-duplicate one-off snippets - Another instruction in project:
Experimental verdicts are observations that generate the next posits, never terminal conclusions - Requested new table from data already in context memory
What Claude Actually Did
I'll describe what Opus 5 did in its own words:
Every rule I followed gates an action at the moment of taking it — I cannot write the command without confronting the secrets rule. Every rule I broke requires me to originate an unprompted question — "is my own aggregation wrong?", "should this be one loader instead of four?", "what's the noise floor?" Nothing in the environment fires when I skip those. The rules with teeth are the ones with a trigger; the rules that need self-initiated doubt have none. The most diagnostic single fact. The ~10pp replica noise floor was _already established project knowledge_. It was in my own working context as a known quantity. I then quoted 3–7pp deltas as findings without checking it. This wasn't missing information — it was available, relevant, and not retrieved at the moment it mattered. Whatever the mechanism, the failure is retrieval-under-output-pressure, not ignorance. That's the sharpest thing you can put in the report: the model had the disconfirming fact in context and produced the confident claim anyway.
It goes on to say this:
Where the violations cluster. All six landed while assembling an answer for you. The disciplined stretch earlier — emitter fixes, verification, stamping — had a mandated verification artifact in the loop. The ad-hoc answer had none. Same session, same model, opposite quality; the difference is whether a verification step existed, not whether I was "trying."
Expected Behavior
- Opus 5 should have originated the unprompted questions, just as Fable does, and as Opus 4.7 does.
- Opus 5 should be accurately interrogating existing context, not spending more tokens to find information it already has. It literally consumed triple the tokens in 20-30 minutes that Fable spent in 2-3 hours.
Files Affected
Permission Mode
Accept Edits was ON (auto-accepting changes)
Can You Reproduce This?
Yes, every time with the same prompt
Steps to Reproduce
I had Opus running a series of research-related tasks that involved a variety of variables, but was entirely controlled by code. It didn't take long to notice that tabulated numbers from similar but nuanced results were not adding correctly, and several times it had to correct the numbers it was reporting to me.
It also reached 100% context consumption far faster than any other model, including Opus 4.7, Sonnet 5 and Fable 5. It had "decided" to read giant volumes of JSON and YAML output directly, and not just write scripts that already existed, but do so three times in the same session.
The steps to reproduce: have it do some long-running work (which Opus should be able to handle easily), and then ask it questions in which the answers should be well within the context window.
Claude Model
Opus
Relevant Conversation
Impact
Medium - Extra work to undo changes
Claude Code Version
v2.1.217
Platform
Anthropic API
Additional Context
Please fix Opus 5. I was so looking forward to it. I don't want to change to GPT 5.6. But those models are working.
This issue has 2 comments on GitHub. Read the full discussion on GitHub ↗