Opus 5.0 nerfed: terrible quality and does not deliver work even after 5 retries

Status Open
Reported on v2.1.220
Maintainer reply None cached
Activity 9 comments · opened Jul 29, 2026

Preflight Checklist

  • [x] I have searched existing issues for similar behavior reports
  • [x] This report does NOT contain sensitive information (API keys, passwords, etc.)

Type of Behavior Issue

Claude modified files I didn't ask it to modify

What You Asked Claude to Do

Opus 5.0 nerfed: terrible quality and does not deliver work

Environment

  • Claude Code (VS Code extension)
  • Model: Opus 5.0 (claude-opus-5)
  • Reasoning effort: Max
  • Task: add a blog page to an existing Next.js marketing site, write one article, produce a cover image

Summary

Opus 5.0 is terrible quality and doesn't deliver work. Basic tasks take 5-6 prompts and retries by the user before anything usable comes out. I am paying for the top model at Max effort and getting output that needs correcting on almost every turn.

Anthropic is seriously nerfing its new models. The gap between what these releases are announced as and what they actually do in daily use keeps widening, and Opus 5.0 is the clearest example of it so far.

In my assessment its performance is equivalent to Gemini 2.5 Pro. That is not a compliment for a flagship model in 2026, and it is not what the pricing or the "Max effort" setting implies.

What actually happened

Single session, one straightforward feature. Every item below is a retry caused by the model, not a change of mind on my part.

  1. Blog index design. First attempt was visually dated. I had to tell it to look at how competitors do it and that its page looked like it came from 1991 before it produced anything modern.
  2. Cover artwork, five rounds. Options round, hybrid round, then four more rounds of corrections (background too dark at the edges, act name not legible, stamp placement, law name missing on the card, label position, removing elements). I eventually gave up and made the image myself.
  3. Using my own image, retry one. I told it to use the left side of the image for the smaller render. It applied that to one of two render paths and declared it finished. The path it missed was the only one visible on the page, so nothing changed.
  4. Using my own image, retry two. The crop it produced cut the text off. It had taken a blind percentage slice of the image without checking where the text was.
  5. The article. First draft was padded, used wrong tenses for a news piece, and was full of obvious AI phrasing ("Here is the part that matters", "Read that list again", "The honest read on this"). Required a complete rewrite, which came in 48% shorter.

The pattern

It reports work as complete without verifying it. It checked HTTP status codes and grepped for a filename, then said the image was done. It has the ability to open and look at the image it produced and did not do so until I complained. Same failure on the render path: it edited one of two code paths and did not look at the resulting page.

Expected vs actual

Expected: a flagship model at maximum reasoning effort measures before it cuts, checks its own output, and finishes a small feature in one or two passes.

Actual: 5-6 rounds of user correction for basic work, with the user supplying the deliverable themselves partway through.

Impact

Time cost is worse than doing the work myself or handing it to a cheaper model. If Max effort produces this, the setting is not doing what it claims, and the difference between Opus 5.0 and a mid-tier model is not visible in the output. Combined with the pattern of new releases underperforming their announcements, the value proposition for paying for the top tier is gone.

What Claude Actually Did

DIDN'T DO ANYTHING THAT I ASKED PROPERLY

Expected Behavior

PROPERLY DO WHAT I ASK IT TO DO, not that difficult is it?

Files Affected

Any

Permission Mode

Accept Edits was ON (auto-accepting changes)

Can You Reproduce This?

Yes, every time with the same prompt

Steps to Reproduce

_No response_

Claude Model

Opus

Relevant Conversation

Impact

High - Significant unwanted changes

Claude Code Version

anthropic.claude-code-2.1.220-linux-x64

Platform

Anthropic API

Additional Context

_No response_

View original on GitHub ↗

3 Comments

Azbesciak · 28 days ago

Agreed. It is hard to believe that you publish benchmarks, compare models, pop champagne, and then, a week later, it turns out you have made the product worse. How reliable are your guarantees, then? You make promises, only for them to turn out to be lies later.

ngill307 · 28 days ago

I don't trust the 'Trust me bro' benchmarks, even if they are somehow not already providing the answers to the model, I bet they somehow try to recognise any accounts linked to famous people, benchmark testers, news orgs, influencers and give them some special 10m token reasoning window model, for normies like us, they slip in a low reasoning model, whenever they can. Choose whatever setting you want on your Claude Code, it's all for show.

Not so surprisingly, my service quality has been a lot better since I made this public post from the email linked to my Anthropic account.

Let's keep pubicly shaming them for their switch-and-bait tactics, and keep a GLM 5.2 or Deepseek V4 ready in the background.

jlcp89 · 20 days ago

Opus 5 is not following the rules created inside the .claude folder at pc and project level.

Showing cached comments. Read the full discussion on GitHub ↗