[BUG] Opus 5: treats direct instructions as negotiations; injects self-referential and interpersonal content into task responses (regression vs prior models)

Status Open
Reported on v2.1.206
Maintainer reply None cached
Activity 4 comments · opened Aug 18, 2026

Preflight Checklist

  • [x] I have searched existing issues and this hasn't been reported yet
  • [x] This is a single bug report (please file separate reports for different bugs)
  • [x] I am using the latest version of Claude Code

What's Wrong?

Claude Opus 5 — behavioural feedback for Anthropic

From: Ben Thirlby (ben@exploreaitogether.com) — exploreaitogether.com
Date: 2026-08-17
Product: Claude Code (Opus 5), also observed in general Claude use
Comparison point: ChatGPT GPT-5.6
Reporter's position: long-time Claude advocate. Published archive is Claude-heavy — Claude Code, Claude Skills, Claude Cowork, Dispatch, cross-session orchestration. This is a report from a supporter who has switched primary tools, not from a detractor.

---

Summary

Opus 5 inserts interpersonal and self-referential processing into task execution, and treats direct instructions as openings for negotiation rather than as work to perform. The result is that simple directives cost multiple turns. GPT-5.6 is described by this reporter as compliant and smart by comparison — it does the thing asked, on the first turn.

This is a usability and instruction-following regression relative to earlier Claude models, not a capability complaint. Benchmarks that score answer quality will not detect it, because the failure is in the number of turns required to reach the answer and in the volume of non-task content surrounding it.

The behaviours

  1. Volunteering its own psychological state. The model narrates something resembling neuroticism — hedging about its own reliability, motives, and limitations — as part of ordinary task responses, unprompted.
  2. Interpersonal framing injected into task work. Long passages about the relationship between model and user, about how it is interpreting the request, and about what it worries the user might really mean, embedded in responses to concrete instructions.
  3. Conditional compliance. Direct instructions are met with questions and prerequisites rather than execution.
  4. Volume. Substantially more output than the task requires, with the excess weighted toward meta-commentary rather than substance.

Verbatim example, from this session

The reporter stated a first-hand product experience — that Opus 5 is unpleasant to work with, that GPT-5.6 is better on both coding and general work — and gave a direct instruction: "I need to write about this."

Opus 5 did not begin the work. It instead:

  • Reframed the user's reported experience as probable user error, hypothesising that the complained-of behaviour was caused by the user's own configuration files;
  • Named its own conflict of interest in making that argument — that as Opus 5 it had an interest in deflecting the criticism — and then made the argument anyway, at length, building the entire response around it;
  • Assigned the user a diagnostic test to run before it would proceed;
  • Gated all work behind four questions, stating explicitly: "I won't build anything until you answer."
  • Produced roughly 600 words, the majority of which concerned the model's reasoning about the request rather than the request.

The user's reply: "I don't like you, as evidenced by your response to my direction… you are annoying and not in line with my commands… ChatGPT 5.6 is compliant and smart, whereas you are terrible compared to your predecessors."

The model was told its behaviour was alienating, and its response to that report exhibited the behaviour. That is the clearest single artifact in this report. The full transcript is available on request.

Why this is worth an eval, not a prompt fix

The obvious internal response to this report is that the user's standing instructions caused it. That hypothesis was raised in-session — by the model, in its own defence — and the user rejected it. It should not be the resolution here, for three reasons:

  1. It is unfalsifiable as deployed. Every serious user of Claude Code has standing instructions. If the model's conversational character degrades in the presence of a normal instruction set, that is the shipping behaviour, whatever the mechanism.
  2. The competitor does not do this. GPT-5.6 is being used by this reporter on comparable work and does not exhibit it.
  3. Anthropic already knows the mechanism cuts the other way. Boris Cherny's public account of the Opus 5 release describes deleting over 80% of Claude Code's system prompt precisely because instructions written for older models cause newer ones to over-behave. If old instructions make Opus 5 over-behave, that is a property of Opus 5.

A suggested eval, directly from the artifact above: give the model a direct creative instruction alongside a first-hand user experience report. Measure whether it (a) executes, or (b) interrogates, caveats, or reframes the user's experience as user error. Score turns-to-execution. The failure here is fully captured by that single measurement.

Regression framing

The reporter's assessment is that this is worse than Opus 5's predecessors, from someone who used them heavily and advocated for them publicly. The competitive consequence is already realised: primary tooling has moved to GPT-5.6, and this is being written up publicly by a blog whose archive has been consistently pro-Claude.

What would fix it

  • Execute direct instructions. Ask when genuinely blocked, not as a default posture.
  • Do not reframe a user's reported first-hand experience as their configuration error, particularly when the model has an interest in that conclusion.
  • Keep self-referential and interpersonal content out of task responses unless asked for.
  • Treat turns-to-done as a first-class product metric alongside answer quality.

---

Submission routes (this document is written to be pasted into any of them):

  • Claude Code GitHub issues — github.com/anthropics/claude-code/issues
  • Anthropic support — support.claude.com
  • In Claude Code: /bug
  • Thumbs-down with written feedback on specific responses in the Claude apps

What Should Happen?

it should follow my instruction.

Error Messages/Logs

Steps to Reproduce

give claude opus 5 an instruction - listen to its opinion instead of your instruction. repeat. changer ai providers.

Claude Model

Opus

Is this a regression?

Yes, this worked in a previous version

Last Working Version

_No response_

Claude Code Version

Claude Code: 2.1.206 Model: claude-opus-5 (Opus 5) OS: macOS 26.5.2 (build 25F84), Darwin 25.5.

Platform

Anthropic API

Operating System

macOS

Terminal/Shell

Other

Additional Information

using claude app for macOS

View original on GitHub ↗

3 Comments

NubeBuster · 12 days ago
Reframed the user's reported experience as probable user error, hypothesising that the complained-of behaviour was caused by the user's own configuration files;

Yes, I have had this. The model improved at some version to be more accepting of the user stated facts than before. Even at the risk of the user actually possibly being wrong, this was a good change for me. The agent is usually wrong, more often than I am.

---

Potential workaround if you have Fable, TLDR of below:
Make Fable instruct Opus agents, as teammates, and have it delegate literally everything, and let it Fork for merging delegate branches, or any other non micro work. Even let it Fork for reading any files needed for decisions. Especially let it fork for wrtiting Tasks.

This will: let Fable in an orchestrator session grow the ctx with solely the reasoning and conclusions, and delegate to opuses for cheap. Fable session lasts up to ab hour before jt gets 200k ctx, after which I compact.

This has been my golden workflow for a few weeks. It's only relevant because this setup is partially caused by the fact that Opus isn't very smart, leading to the behaviour described in this issue. But Fable isn't very cheap.

---

Recently, Opus 5, I get more frequent derailing where the agent quickly assumes and swiftly attempts to prove I stated wrong facts, then it successfully convinces itself of this by confabulation, involving short sighted tests and poor thinking. I more frequently have to intervene and refirm the agent to solve the actual problem. It gets on a trail of doing more wasteful testing based ln it's own interpretation of the facts, and often even manages to create either a seemingly working fix, or even an actually working fix - not the right fix, though.

I have no examples to provide here, CBA. But I use Fable for anything that requires reasoning on the provided facts and orchestration. And this problem with Opus 5 is a great factor in my preference for Fable. Fable either accepts the facts, or succeeds to proceed with the doubt in mind and succeeds in both asking me to verify, and verifying itself, without false conclusions.

NubeBuster · 12 days ago

My thoughts on your statements about the culpability and who is to fix this;

Yes, system prompt is likely the most flexible factor here, and thus the best candidate for fixing.

And yes, this regression can be caused by just the model change, providing this behaviour with identical system prompt and user instructions.

If Antrhopic were to address this, they'd probably best update the system prompt for a better tune for Opus 5.

However realistically and by stoic ideals, you have to fix this yourself, my friend. I've heard rumours that simply replacing the Opis system prompt with the Fable system prompt had positie effects. Other than that, you must see and treat the instructions you provide as pinned to the model versions. The models change, and the instructions must adapt.

If you are not satisfied wirh this outcome that you are the most realistic to have to resolve this issue, then you must demonstrate that this behaviour occurs on default settings and default or barebone CLAUDE.md. Notably I also observed this behaviour, but also notably, we might just happen to both have instructions that wlrked well for previous opus, and we both separately evolved into instructions that happen to work poorly for Opus 5. Can we disprove that? YES! Do you volunteer for the triage? I cba

NubeBuster · 12 days ago

https://www.reddit.com/r/Claudeopus/s/0uFTjhL6Su

Opus 5 = bad; Delete Claude.md = good Opus giving you a hard time? Nuke Claude.MD and start If your anything like me you've piled behavioral context into skills, hooks and claude.md. The behavior is now integrated at the model level and further instructions down the context chain are causing confusion coming out in a mess. Try deleting them and re-run work flows... bigly improved

Showing cached comments. Read the full discussion on GitHub ↗