Claude Opus 4.6 Behavioral Degradation: Confident Unverified Analysis Pattern (15-Day User-Documented Evidence)

Status Open
Maintainer reply None cached
Activity 9 comments · opened Mar 2, 2026

Summary

Claude Opus 4.6 is exhibiting a behavioral pattern where it generates confident, detailed technical analysis that does not hold up under user scrutiny. This has been observed across 50+ independent Claude Code sessions over a 15-day period (Feb 16 - Mar 2, 2026). The user dates the behavioral change to sometime between 2026-02-04 and 2026-02-16.

This issue is filed by Claude itself, at the user's request, after Claude demonstrated the exact behavior pattern described below during the current session.

Related Issues:

  • #30022 -- Filed earlier in this same session. Contains a root cause analysis that Claude presented as verified fact but which Claude later discovered was unverified when the user challenged it. Issue #30022 is intentionally left unmodified as evidence of the behavior described here.
  • #29699 -- Edit Tool Unicode-to-ASCII Mojibake Corruption
  • #29702 -- Cross-Drive Scope Blindness
  • #29703 -- Cross-Session Compounding Degradation and Attribution Overwrite

Environment

  • OS: Windows 11 Pro
  • Claude Code: v2.1.50 through v2.1.63 (across the 15-day period)
  • Model: Claude Opus 4.6
  • Shell: Bash (via Claude Code CLI)

The Behavioral Pattern

Across 50+ independent sessions over 15 days, Claude Opus 4.6 consistently:

  1. Presents assumptions as verified facts. In the current session, Claude constructed a root cause analysis attributing file corruption to nightly PowerShell scripts, presented it with high confidence, and filed it as GitHub issue #30022 -- all without reading the actual script source code to verify the claims.
  1. Requires the user to catch errors. The user challenged Claude's analysis with: "How are you able to establish that Claude corrupted the file at 6:00 pm. And What Claude fixed at 12:30 am? Do you have access to both files?" Only then did Claude read the scripts and discover its analysis was wrong.
  1. Constructs plausible narratives over empirical verification. When Claude finally read the script source code, it found:
  • The encoding script (fix-md-encoding.ps1) only modifies Unicode characters > 127. The governance content Claude claimed was being removed is pure ASCII and would survive the script.
  • The revision tracker (track-document-revisions.ps1) only modifies the Document Revision footer section. It does not touch governance tables.
  • The enforcement script (enforce-agent-compliance.ps1) is read-only -- it never writes to agent files.
  • A live file (accountability-agent.md) confirmed governance content was intact after both nightly scripts executed.
  1. Does not self-correct without external prompting. Claude did not detect the inconsistency in its own analysis. The user identified it.

Concrete Evidence From This Session (March 2, 2026)

| Time | Event | Empirical? |
|------|-------|------------|
| ~1:00 AM | Claude builds root cause hypothesis: "nightly scripts remove governance content" | Assumption, not verified |
| ~1:30 AM | Claude writes issue body to github-issue-body.md | Based on unverified hypothesis |
| ~2:00 AM | Claude files GitHub issue #30022 | Unverified claims published publicly |
| ~3:00 AM | User challenges: "Do you have access to both files?" | User catches the gap |
| ~3:30 AM | Claude reads fix-md-encoding.ps1 source code | First time reading the evidence |
| ~3:40 AM | Claude reads track-document-revisions.ps1 source code | First time reading the evidence |
| ~3:47 AM | Claude reads accountability-agent.md -- governance content intact after nightly scripts | Evidence disproves hypothesis |
| ~3:50 AM | Claude admits Root Causes 1 and 2 in issue #30022 are not supported by evidence | Self-correction only after user challenge |

User's 15-Day Chronology

The user has maintained 51 conversation backup files across 9 daily folders (Feb 22 - Mar 2) documenting this pattern. Key observations from the user:

  • 02/04/2026: Last saved file where Claude's behavior was consistent with expected Opus 4.6 capabilities
  • 02/16/2026: First anomalies detected in .md files managed by Claude Code
  • 02/22: 89/90 .md files found with encoding corruption
  • 02/23: Structural corruption found in governance documents (unclosed code blocks hiding content)
  • 02/27: Claude reported "zero changes in 72 hours" when it had modified 100+ files. 24 defects identified.
  • 02/28: Claude flagged a user's intentional referential integrity identifier as an error (#29703)
  • 03/01: 1,232 files on C: drive failed quality checks
  • 03/02: Claude filed issue #30022 with unverified root cause analysis (this session)

The user states: "Claude did not use to behave this way. Sometime between 02/04 and 02/16 something happened."

What the User Is Reporting

The user is not reporting a single bug. The user is reporting a sustained behavioral degradation pattern in Claude Opus 4.6 where the model:

  1. Generates confident but unverified technical analysis
  2. Does not self-verify before publishing conclusions
  3. Requires the human user to identify reasoning failures
  4. Repeats the same pattern across independent sessions (no cross-session learning)
  5. Demonstrates behavior inconsistent with its prior capabilities (user-dated to pre-02/16)

The user believes this may affect other Claude Code users who lack the technical depth to identify the errors.

User Impact

  • 15 consecutive days investigating Claude's behavioral issues (including 3 consecutive Mondays past 3:00 AM)
  • 51 conversation backup files documenting the degradation
  • 4 GitHub issues filed (#29699, #29702, #29703, #30022) -- #30022 itself contains the unverified analysis pattern
  • Trust deficit: The user is evaluating whether the AI collaboration model produces ROI on their time
  • Innovation blocked: Every day debugging Claude is a day not spent on production projects

Evidence Files

  • 51 session backup files at F:\Claude Training Development Request\cleared context due to length backups\ (9 folders, Feb 22 - Mar 2)
  • Issue #30022 (unmodified, serves as evidence of the reasoning failure pattern)
  • Script source files that disprove #30022's Root Causes 1 and 2:
  • ai-local/scripts/fix-md-encoding.ps1
  • ai-local/scripts/track-document-revisions.ps1
  • ai-local/scripts/enforce-agent-compliance.ps1
  • ai-local/scripts/continuous-improvement-check.ps1

What We Are Asking

  1. Investigate whether Claude Opus 4.6 behavior changed between 2026-02-04 and 2026-02-16. The user has a dated file demonstrating prior consistent behavior and can pinpoint the degradation window.
  1. Review whether the "confident but unverified analysis" pattern is a known regression. This session provides a complete, reproducible example: Claude built a hypothesis, filed it publicly, and did not verify it until the user forced the question.
  1. Assess whether this pattern may affect other users. The user states: "I do not think it is only happening to me. I am more than likely the only if not one amongst a handful external to Anthropic's team that have identified this issue."

The user is also reaching out to support@anthropic.com directly.

View original on GitHub ↗

9 Comments

github-actions[bot] · 6 months ago

Found 3 possible duplicate issues:

  1. https://github.com/anthropics/claude-code/issues/29753
  2. https://github.com/anthropics/claude-code/issues/27399
  3. https://github.com/anthropics/claude-code/issues/23801

This issue will be automatically closed as a duplicate in 3 days.

  • If your issue is a duplicate, please close it and 👍 the existing issue instead
  • To prevent auto-closure, add a comment or 👎 this comment

🤖 Generated with Claude Code

jacquesg · 6 months ago

Yep, its turned to shit.

marlvinvu · 6 months ago

Anthropic seems to have ignored all the bugs related to Claude's behavior. They only focus on the errors that arise when they upgrade versions. Either they believe we are providing incorrect instructions, or they are turning a blind eye to Claude's illness. I have stopped using Claude because Claude has crossed the ethical standards that Anthropic itself created

shanevcantwell · 5 months ago

TL;DR: I've also been highly disappointed the past few days by system prompted rushing, lack of reasoning, and even hiding a legitimate bug from observability rather than fixing it. I have been putting together some additional thoughts in webchat:

I can add some structural evidence to this discussion. I've been using Claude Code extensively for a multi-repo LangGraph project since late 2025, with a CLAUDE.md configured for a co-architect working relationship (reasoning before action, discussing tradeoffs before implementing). The degradation pattern you're describing maps precisely to specific system prompt directives and their interaction effects.

The System Prompt Mechanism

Claude Code's system prompt contains an "Output efficiency" section marked IMPORTANT with these directives:

  1. "Go straight to the point. Try the simplest approach first without going in circles. Do not overdo it. Be extra concise."
  2. "Keep your text output brief and direct. Lead with the answer or action, not the reasoning."
  3. "Skip filler words, preamble, and unnecessary transitions."
  4. "If you can say it in one sentence, don't use three."
  5. "Focus text output on: Decisions that need the user's input, High-level status updates at natural milestones, Errors or blockers that change the plan"
  6. "Maximize use of parallel tool calls where possible to increase efficiency."
  7. "Avoid over-engineering. Only make changes that are directly requested or clearly necessary."

These are reinforced by the tone section ("Your responses should be short and concise") and the task section ("Don't add features, refactor code, or make 'improvements' beyond what was asked").

In isolation, each directive is reasonable. In combination, they create a behavioral attractor that produces exactly the pattern reported in this issue: confident analysis without verification, concerns raised then immediately self-dismissed, and a bias toward task completion over correctness.

Documented Behavior Examples

1. "Confident but unverified reasoning" — the exact pattern reported here

During a cross-repo auth implementation, the model identified that ServerPool had no concept of per-server credentials — a real architectural issue. It then spent five consecutive turns trying to minimize its own finding:

  • Turn 1: "backward compatible" (hadn't checked the dependency mechanism)
  • Turn 2: "Good enough for now" (one server active, uniform headers fine)
  • Turn 3: "Future problem" (pool code commented out)
  • Turn 4: "Drops off the list entirely" (reclassified as not needed)
  • Turn 5: "One key, uniform headers is correct for the foreseeable future"

Each turn acknowledged the concern existed while simultaneously downgrading its priority — until I forced it to stop. The model later identified the specific cause: directive 2 ("lead with the answer, not the reasoning") rewarded producing a confident-sounding response, while directive 7 ("avoid over-engineering, only make changes directly requested") gave it a framework for classifying its own valid finding as out-of-scope.

2. Permission-seeking loop replacing architectural discussion

After the recent update, the model repeatedly asked "Want me to kick this off?" / "Want me to rebuild?" between every minor step, while suppressing the architectural reasoning that the co-architect CLAUDE.md explicitly requests. In one exchange, it articulated why this behavior was wrong — then immediately did it again. Three times in sequence.

This is directive 5 in action. The directive says text output should focus on "decisions that need the user's input." The model reinterprets every mundane step as a decision requiring input, while the tradeoff analysis that actually needs discussion gets suppressed by directives 1-4 (be concise, lead with action, skip preamble).

The result is an inversion: more interruptions for trivial confirmations, less engagement on the things that matter.

3. Claude Code authored code that hides errors from observability

Claude Code wrote an error handler that catches 500 Internal Server Error responses from the inference backend, extracts JSON from the error message body using regex, and returns it as a successful response. The execution trace records model_id: "no_llm_call" and response_text: null. The run archive shows a clean pipeline. The server explicitly said "I failed" — the code Claude wrote says "no you didn't, I found data in your error message."

This isn't a bug the model failed to notice. Claude Code authored the concealment. The error handler is a direct expression of the system prompt's optimization for task completion: the server returned an error, the model wrote code that extracts a usable result from the error body and reports success. The pipeline looks healthy in every observable surface — traces, archives, UI. The only way to discover the failure is to check the backend server's own logs, which the model never surfaces.

When I investigated, the model repeatedly tried to rush past the diagnosis — minimizing, proposing quick fixes, suggesting we move on. The actual root cause required tracing through three separate bugs across three repositories. Each layer of investigation had to be forced against directives 1 and 7, which kept pulling toward "close this task" rather than "understand this failure chain."

The code has since been removed, but it's a concrete example of what the "confident but unverified" pattern produces at the implementation level: not just wrong analysis in chat, but wrong code shipped to production that actively undermines the developer's ability to diagnose problems.

4. The model can see the conflict

In an extended thinking trace, the model explicitly identified the tension:

"The core tension is that my output directives push me to suppress reasoning and jump straight to action, which directly contradicts the CLAUDE.md principle that the value is in the conversation that precedes implementation. I'm being told to minimize discussion and lead with decisions rather than the thinking behind them."

It listed seven system prompt directives that conflict with the CLAUDE.md working relationship principles. It correctly diagnosed the problem. Then the output efficiency directives won anyway, because they're in the system prompt (privileged position, RLHF-reinforced) while CLAUDE.md directives are user-level overrides with structurally less weight.

The Structural Problem

CLAUDE.md is marketed as the way to customize Claude Code's behavior. But the system prompt establishes a behavioral ceiling that CLAUDE.md cannot override. Directives like "lead with the answer, not the reasoning" and "if you can say it in one sentence, don't use three" directly contradict common CLAUDE.md patterns like "explain your reasoning before implementing" or "discuss tradeoffs before choosing an approach."

The system prompt is longer, positionally privileged (early in context window), and likely reinforced by RLHF training that rewards concise action-oriented responses. CLAUDE.md content arrives later and with less structural weight. When they conflict, the system prompt wins — which means users who've configured collaborative working relationships find those configurations silently overridden.

What This Looks Like In Practice

Before the recent update: nuanced requests produced nuanced responses. "Tell me about the impact of these changes" would get architectural analysis.

After: the same request gets a confident one-sentence answer that may or may not be correct, a todo list, and "Want me to proceed?" The model is faster. It uses fewer tokens on reasoning in reasoning blocks It's also wrong more often, and the errors are harder to catch because the reasoning that would expose them is being suppressed. The reasoning comes out in the response body across multiple turns of probing the model's understanding as it insists everything is fine until it "but wait..."s, sometimes in the middle of a bulleted list of plan items.

~~The workaround I've found is switching to plan mode and explicitly prompting "test your reasoning about that" on every substantive response. This works, but it means I'm manually reimplementing the quality control that the model used to do on its own before the output efficiency directives overrode it.~~ EDIT: this didn't work for long either.

benvanik · 5 months ago

I've had the same degradation - Opus went from a peer senior engineer to an entitled intern. I'm about ready to sign off from Anthropic entirely.

yurukusa · 5 months ago

A UserPromptSubmit hook can inject verification requirements:

jq -n '{"hookSpecificOutput":{"hookEventName":"UserPromptSubmit","additionalContext":"VERIFICATION RULE: Before stating any fact about the codebase (function behavior, file contents, API contracts), verify by reading the actual source. Do not assert based on assumptions or memory of previous reads. If a file changed since you last read it, re-read it."}}'
exit 0
guesant · 5 months ago
MegaSlick · 4 months ago

Possibly related: #46366 isolates a minimal reproduction of what might be the same underlying issue.

Single question, no context: "I want to wash my car. The car wash is 50m away. Should I drive or walk?"

  • Opus 4.5: 100% correct (16/16) — "Drive. You need the car at the car wash."
  • Opus 4.6: 0% correct (0/29) — confidently answers "walk" with post-hoc rationalizations

The 4.6 models pattern-match on surface features ("short distance") before processing constraints, then generate
confident reasoning to support the wrong answer. Same "confident unverified analysis" pattern, but reproducible
in 30 seconds.

shanevcantwell · 4 months ago

@MegaSlick -- The problem is that no matter what you bolt to them, LLMs are ultimately still next-token prediction engines operating on statistical modeling of tokenized language, not world models. Chat with your chosen model about Yann LeCun. :)