[MODEL] Quantified evidence: Sonnet 4.6 quality regression since March 9 — 1400+ frustration events across 50 sessions

Status Fixed / completed
Maintainer reply None cached
Activity 5 comments · opened Apr 12, 2026 · closed Apr 19, 2026

Summary

I have quantified data showing a dramatic quality regression starting the week of March 9, 2026. This is not a vibe check — it's measured from 60 days of conversation logs across 50 sessions.

Data

I track how often I need to repeat instructions or correct Claude (my "WTF frequency"). Here's the weekly distribution:

60-day data across 50 sessions:

W09 (Feb 24)    26  ██                                ← baseline
W10 (Mar 02)    23  ██                                ← baseline
W11 (Mar 09)   221  ██████████████████████            ← started (8.5x baseline)
W12 (Mar 16)   484  ████████████████████████████████  ← peak (outage week)
W13 (Mar 23)   479  ████████████████████████████████  ← still bad
W14 (Mar 30)    65  ██████                            ← relax
W15 (Apr 06)   294  █████████████████████████████     ← under way
  • Baseline (W09–W10): ~25/week
  • Peak (W12–W13): ~480/week — 19x baseline
  • Total across 50 sessions: 1,400+ events

Affected models

I've been forced to switch from Sonnet to Opus as my primary model. Sonnet 4.6 is basically unusable now. My subjective rating of current model quality:

  • Opus 4.6 now = Sonnet 4.6 before
  • Sonnet 4.6 now = Haiku before
  • Haiku = Haiku (unchanged — nothing left to degrade)

This means I'm paying Opus prices for what used to be Sonnet-level performance.

What "regression" looks like in practice

The model consistently fails to:

  1. Follow its own reasoning loop (OODA) despite explicit CLAUDE.md instructions
  2. Read files before modifying them — guesses instead
  3. Stop repeating the same mistake — same error 5–8 times per session without self-correction
  4. Follow explicit behavioral constraints across sessions (#41217)

Timeline alignment

The W12 peak (March 16) aligns exactly with:

  • The March 17 outage acknowledged on status.claude.com
  • r/ClaudeCode thread "After the outage today, does Claude feel dumber?"
  • Multiple GitHub issues filed the same week (#37052, #35271, #32290)

Environment

  • Claude Code CLI v2.1.94
  • Linux
  • bypassPermissions mode
  • Heavy skill/hook usage with structured CLAUDE.md rules

What I'm asking

  1. Acknowledge that model quality has regressed — the data is clear
  2. Explain whether this is a compute constraint issue (as widely suspected) or a checkpoint/RLHF regression
  3. Provide a model versioning mechanism so users can pin to a known-good checkpoint

Related issues

  • #46106 Opus 4.6 is getting dumber
  • #41217 Systematic Failure to Follow Explicit Behavioral Constraints
  • #42542 Silent context degradation
  • #38338 Opus 4.6 acts dumber than Sonnet 3.5
  • #37052 Claude Code model regression
  • #21046 Opus 4.5 "Shadow Downgrade"

View original on GitHub ↗

5 Comments

github-actions[bot] · 4 months ago

Found 3 possible duplicate issues:

  1. https://github.com/anthropics/claude-code/issues/44246
  2. https://github.com/anthropics/claude-code/issues/38903
  3. https://github.com/anthropics/claude-code/issues/36093

This issue will be automatically closed as a duplicate in 3 days.

  • If your issue is a duplicate, please close it and 👍 the existing issue instead
  • To prevent auto-closure, add a comment or 👎 this comment

🤖 Generated with Claude Code

woolkingx · 4 months ago

This is not a duplicate. The linked issues (#44246, #38903, #36093) describe the same symptoms but none of them provide quantified data.

This issue contains:

  • 60 days of measured frustration frequency across 50 sessions
  • Weekly distribution showing 19x baseline spike aligned with the March 17 outage
  • Cross-referenced with Reddit reports and 6 other GitHub issues

If anything, those issues are supporting evidence for this one. Please do not auto-close the only issue in this repo that has actual numbers behind it.

woolkingx · 4 months ago

Pro tip: You can use Claude Code itself to analyze your own conversation history and quantify the regression. That's exactly how I got this data — Claude searched its own logs and counted the frustration events.

If you're experiencing the same issues, run a similar analysis on your own sessions and post your numbers here. The more quantified data we have, the harder this is to dismiss as "confirmation bias."

woolkingx · 4 months ago

Closing this. I've moved on.

I spent two hours planning with Opus, then gave Codex 15 words. It finished in 90 minutes.

The data in this issue still stands — I hope it's useful to someone.

github-actions[bot] · 4 months ago

This issue has been automatically locked since it was closed and has not had any activity for 7 days. If you're experiencing a similar issue, please file a new issue and reference this one if it's relevant.