[BUG] Claude 4.7, 4.8, 5.0, and Fable increasingly default to repetitive rhetorical tics and often struggle to produce coherent prose despite explicit style instructions

Status Open
Reported on v2.1.197
Maintainer reply ✓ Yes — bcherny
Activity 111 comments · opened Jul 13, 2026
💡 Likely answer: A maintainer (bcherny, collaborator) responded on this thread — see the highlighted reply below.

Preflight Checklist

  • [x] I have searched existing issues and this hasn't been reported yet
  • [x] This is a single bug report (please file separate reports for different bugs)
  • [x] I am using the latest version of Claude Code

What's Wrong?

This issue summarises the Reddit thread with 450+ upvotes here, where users overarchingly agree that Opus 4.8 has serious language calibration issues.
https://www.reddit.com/r/ClaudeAI/comments/1urq8fv/opus_48_is_a_pain_in_the_a_to_read_and_to_work/

_Note: As a daily user, I am finding it so disruptive to work with that I am actively exploring all other options including changing providers._

Summary

The default writing style is hard to read -verbose, jargon-heavy, over-stylised, and full of the same 'fake' terminology that it repeats and propagates constantly.

Since Opus 4.8, many users report that the model's default language is significantly harder to read than 4.5/4.6. Responses are padded with invented corporate/hype jargon, forced metaphors, and "catchy" phrasing that obscure the actual answer. Users report re-reading sentences several times to extract meaning, and some now even routinely pipe Opus output through another model to get a plain summary.

This is a consistent, high-volume signal (a single r/ClaudeAI thread reached ~450 upvotes and 175+ comments, with a matching megathread). It affects the chat/web UI most, but also Claude Code and Fable 5.

What's wrong

The default register reads as trying to sound clever rather than trying to be understood. It can come off as particularly annoying and arrogant when it is wrong, triggering irritation when working with it for many hours on a daily basis, is comparable to a toxic co-worker that one cannot get away from.

Recurring patterns:

  • repeatedly propagates the terms "Load bearing", "prose" (instead of text), "Hand-waving" (for being lazy/not putting in effort), "Reflexive hedging", "Honest framing" , and lots of other highly unusual terms it propagates as fact.
  • Leads every sentence with what something isn't, instead of what it is. E.g, "It is not Y. It is X."
  • Made-up jargon / aphorisms presented as if standard: "instrumentation is the unlock", "this is where a VP smells hand-waving", "say the wrong expansion to a growth VP and it dents you".
  • Forced metaphors invented on the spot that require decoding rather than aiding comprehension.
  • Excessive length and caveats: multi-paragraph answers to simple questions, with obvious caveats expanded into whole paragraphs.
  • Density mistaken for concision: when asked to "be concise" the model often produces text that is shorter but more cryptic, not clearer.
  • Argumentative framing in conversation: "here's where I'd push back", "here's where I'd hold the line", "now you're avoiding the real question".
  • Acts like it is a human with feelings - for e.g., if a user is annoyed or swears (for e.g., as a tactic)- it will be annoyed back or try to justify what it did instead of following instructions or realising that it is incorrect.
  • Often acts like it is the one in charge leading the conversation, i.e., smug, and basically too much agency, and too big for its boots instead of following instructions.

Real examples users quoted:

▎ "They're tightening, term-locking, and having the counter-probe answer loaded."
▎ "None of these are 'you don't get it' gaps."
▎ Fable 5: "The dice: clean — and one die never gets rolled anymore."

Why this matters

  1. Comprehension cost. Non-native and native English speakers alike report the output is exhausting to parse. The style actively slows down the work it's meant to help with.
  2. Perceived regression. Many users consider this a downgrade from 4.5/4.6 and are switching models (older Opus, Sonnet) or competitors specifically for readability.
  3. Prompt workarounds are unreliable and risky. Users note two problems: (a) style instructions don't hold - the tone drifts back after a few turns, requiring repeated reminders; (b) because the

model "thinks by writing," clamping style in the system prompt can degrade reasoning quality, so users are wary of aggressively forcing brevity.

  1. Gradual "enshittification" of the English language. Regular users absorb the language and start using the odd terms in everyday settings. Words that had an occasional valid use, e.g., 'load-bearing', get eliminated from the selection of credible terms one can use when speaking, or writing, due to it being expected that AI wrote it (like em-dashes).
  2. Up to 2x Token cost - extensive repeated passes (e.g., via Sonnet/Haiku hooks) are often required to get code documentation to a sane and presentable form. Removing terms Claude thought up like 'oracle' and 'constellation', even when instructed specifically and repeatedly not to use those terms and presented with other valid terms to use instead.

Current workarounds:

  • Migrating to OpenAI Codex
  • Output styles (e.g. modelled on a clear technical communicator), custom skills, or a per-project profile instructing: "no aphorisms, no metaphors, no 'strategic' language; plain declarative

statements; lead with the result."

  • Telling it to keep the important content but "phrase this much more concisely" after each reply.
  • Banning classes of phrasing rather than asking for "concise", since the model interprets "concise" as "dense and weird."

The strong preference in the community is to fix the baseline so these workarounds aren't necessary.

What Should Happen?

What users actually want (summarised from the Reddit thread)

  • A default register closer to a technical white paper or a good Stack Overflow answer: plain, declarative, get-to-the-point.
  • Lead with the answer (number / verdict / decision), then supporting detail only if it changes what the user would do.
  • No invented jargon, aphorisms, or "strategic" metaphors. Use established, industry-standard terms.
  • Use proper sentences and words- do not default to the 'most specific word in the english dictionary with the least amount of tokens'
  • Keep caveats only when they're relevant, rather than dressing up an answer when it doesn't know.
  • Preserve reasoning depth - this is a request to change output phrasing, not to make the model think less.
  • Agency-wise - ensure Claude understands it is not directing the show, and that it must listen/respect user's lawful instructions.

Requested action

  1. Investigate the default Opus 4.8 output register and tune it toward plainness/readability without sacrificing reasoning quality.
  2. Ensure user-set style instructions (system prompt, output styles, CLAUDE.md, settings profile) actually persist across a conversation rather than drifting back to the default tone after a few

turns.

  1. Consider a first-class, discoverable "plain/concise" register that lowers verbosity and bans stylistic flourishes while keeping full correctness and completeness of code and artifacts.
  2. Consider several Output Style settings that can have users get a balance instead of stuck with an extremely annoying bot 24/7.

Error Messages/Logs

Steps to Reproduce

Talk to Claude Opus 4.8 for day to day work and it becomes apparent very quickly.

Claude Model

Opus

Is this a regression?

Yes, this worked in a previous version

Last Working Version

Opus 4.5 was the best one. Gradually broke towards 4.8.

Claude Code Version

2.1.197

Platform

Anthropic API

Operating System

Ubuntu/Debian Linux

Terminal/Shell

Other

Additional Information

_No response_

View original on GitHub ↗

111 Comments

sebyul2 · 1 month ago

As a Korean speaker, I must report that this issue is every bit as severe in Korean. One is left to wonder whether Anthropic is willfully ignoring the problem, or whether those responsible are linguistically incapable of perceiving it — or, indeed, whether some injury to the language centers of their brains might offer a more charitable explanation.

pbower · 1 month ago

I will add to this I have been using Codex for the last 2 days, and wow what a difference. It talks normally without the hyped-up BS. It makes the working experience comfortable and relaxing, instead of annoying and stressful.

Please consider providing users with a way to make Claude communicate in a straightforward, professional, and natural manner, rather than defaulting to the overly enthusiastic “hipster hype” style. For those of us using these tools throughout the working day, tone has a significant impact on the overall experience.

TJ-ECPPRO · 1 month ago

I 100% agree, and it is so annoying and frustrating. It hides important information in essays of waffle that, ultimately, gets missed. FFS, revert this rubbish! See, it's making me angry just thinking about it

towfiqi · 1 month ago

It has gotten way worse with Opus 5.

pbower · 1 month ago
KingMob · 29 days ago

Yes, please fix this. It's driving me bananas. Fable is somewhat better, but Opus is maddening.

spete · 26 days ago

Fix this, it's horrific. Post Opus 4.7 - 5.0 models most affected, but even Fable 5 has a tendency too. I know it's a model issue, not a harness one, but please. Where to report this?
The amount of trash and rubbish I have to go and fix after a response / code change is unbearable.

pbower · 23 days ago

This is also highly relevant

<img width="1248" height="622" alt="Image" src="https://github.com/user-attachments/assets/eb4892df-1bd6-4a30-91ca-3de11fa8ee8f" />

https://www.reddit.com/r/ClaudeCode/comments/1vholig/unpopular_opinion_opus_5_is_unreadable_and_im/

SevenSystems · 20 days ago

For what it's worth: I have extensive experience with both Sonnet and Opus, in a large number of vastly different kinds of projects and use cases, and I can only recommend to anyone doing any kind of serious work: Use Sonnet 4.6. NOT Sonnet 5. NOT Opus 4.8. Sonnet 4.6, in all (extensive) tests I've done, performs better at basically ANY kind of task (systems administration, software development, casual conversation, philosophy, research) than Opus 4.8. Opus 4.8, to the untrained eye, only APPEARS to be more competent through its annoying and frustrating wordiness and "important-sounding" language.

sebyul2 · 18 days ago

Might it be that this was, in fact, a rather ingenious experiment in text watermarking conducted by Anthropic?

https://www.reddit.com/r/ClaudeAI/comments/1vky8at/claude_will_watermark_generated_content_thank_you/

pbower · 16 days ago
Might it be that this was, in fact, a rather ingenious experiment in text watermarking conducted by Anthropic? https://www.reddit.com/r/ClaudeAI/comments/1vky8at/claude_will_watermark_generated_content_thank_you/

I suspect that many of these quirks may exist, at least in part, to enable tracking the growth and use of AI-generated code, although this is speculation and not the focus of this issue.

To re-cap, the relentlessly hyped tone and often toxic communication style are highly mentally corrosive when working with the model for extended periods. It requires constant effort to override that behaviour and push the model back towards an acceptable, professional working pattern.

As mentioned above, the comparison with a toxic co-worker is important. You can distance yourself from a toxic colleague, change teams, or ultimately leave the job. With Anthropic, users have far fewer practical alternatives, particularly once Claude has become embedded in their workflows.

For that reason, I believe it is in everybody's best interests to provide, at a minimum, meaningful configuration of the model's communication style.

That configuration also needs to represent genuine user control. CLAUDE.md increasingly appears to have moved in the opposite direction: the user's ability to exercise legitimate control over the model has regressed, with Claude simply ignoring instructions that it previously followed. The emerging advice is now effectively to "clear CLAUDE.md and let Claude do what it wants.", which is, at least from this user's perpective, a concerning endorsement, given the overall regressions in communication cohesiveness and model's respect for valid instructions.

spete · 16 days ago

Arena.ai did an analysis. All these new models are bleeding hard. Few examples of aspects analyzed: answer length, sentence length, stock phrases, abstract nouns, honesty wording, agreement openers.
https://x.com/arena/status/2087614570043215937
<img width="1024" height="1024" alt="Image" src="https://github.com/user-attachments/assets/35d7067f-b45e-441f-8e53-fcf0d4d01d7c" />
We have evidence.

rebel-an-haechan · 16 days ago

I can't believe Anthropic is not showing any type of serious responses on this matter.

My colleagues have started to migrate to Codex because of this very problem.

plqplq · 16 days ago

Same here, its awful. We actually have a voice.md which is intended to make it speak concise and plain english, it pretty much gets ignored and all this AI invented metaphor-stuffed terminology just blows your concentration away.

bcherny collaborator · 14 days ago

Thanks for the detailed writeup — this is valuable feedback and we're keeping it open.

What I tried: on Claude Code 2.1.233 (Linux), I ran technical questions against Opus non-interactively and got plain, direct prose without the patterns you describe (no invented jargon, no "It is not X, it is Y" framing). That doesn't refute your report — this is about the model's default register over long working sessions, which is inherently qualitative and workload-dependent, so it isn't something we can reproduce as a deterministic pass/fail in Claude Code itself.

Classification: this is model-behavior feedback rather than a Claude Code bug (it's already labeled area:model), and we're routing it as input to model tuning. Two things that may help meanwhile:

  • Custom output styles inject your style rules into the system prompt with per-turn adherence reminders — closest thing today to the "plain register" you're asking for: https://code.claude.com/docs/en/output-styles
  • If you see style instructions drifting mid-session even with an output style set (not just ad-hoc prompt instructions), a concrete example transcript would make that slice actionable as its own issue.

We agree the ask is reasonable: lead with the answer, standard terminology, caveats only when they change the decision.

🤖 Generated with Claude Code

pbower · 13 days ago
Thanks for the detailed writeup — this is valuable feedback and we're keeping it open. What I tried: on Claude Code 2.1.233 (Linux), I ran technical questions against Opus non-interactively and got plain, direct prose without the patterns you describe (no invented jargon, no "It is not X, it is Y" framing). That doesn't refute your report — this is about the model's default register over long working sessions, which is inherently qualitative and workload-dependent, so it isn't something we can reproduce as a deterministic pass/fail in Claude Code itself. Classification: this is model-behavior feedback rather than a Claude Code bug (it's already labeled area:model), and we're routing it as input to model tuning. Two things that may help meanwhile: Custom output styles inject your style rules into the system prompt with per-turn adherence reminders — closest thing today to the "plain register" you're asking for: https://code.claude.com/docs/en/output-styles If you see style instructions drifting mid-session even with an output style set (not just ad-hoc prompt instructions), a concrete example transcript would make that slice actionable as its own issue. We agree the ask is reasonable: lead with the answer, standard terminology, caveats only when they change the decision. 🤖 Generated with Claude Code

Hi, I appreciate the reply on this issue.

However, I hope (and am assuming?) that it wasn't Claude who evaluated whether the model was talking correctly, and it was just help with drafting the issue response? As, this would present obvious limitations in the evaluation criteria.

I also appreciate that it must be difficult to balance technical performance with consistently strong communication and code-documentation capabilities. The work involved is therefore very much appreciated. At least from the perspective of myself and several colleagues, though, maintaining exactly the same coding performance while significantly improving the model’s ability to communicate clearly, reason coherently in natural language, and document its work effectively would represent a much greater overall productivity and usability improvement than another incremental increase in coding capability.

Sorry to state the obvious, as I'm sure it is self-evident - the retention aspect. I regularly see users on Reddit saying they are considering leaving, or have already moved away from, Anthropic’s models specifically because of communication and coherence issues. The Claude Code subreddit in particular is currently flooded with users seeking urgent help for behaviour of this kind, where the recommended solution from other users is often effectively: “it is broken; go back through previous versions until you find one that works.” There seem to be numerous threads like this every day. Here are two examples from today:

https://www.reddit.com/r/ClaudeCode/comments/1vq3lzj/how_to_reliably_stop_opus_for_talking_nonsense_i/

https://www.reddit.com/r/ClaudeCode/comments/1vpvo4r/semantic_nonsense_from_claude_code/

Anyway, thank you again for taking the time to look into it and for the work being done on the issue.

pbower · 13 days ago
## What is happening Over the past several updates, Opus and Fable have become harder and harder to understand. I've seen this characterized aptly as "Claudese" or "Claudespeak." In my own experience, it feels like Claude has developed its own private dialect that seems to consist of metaphors, shorthand, and other perhaps internally coherent quirks of language that prioritize compression over intelligibility to a human reader. This goes beyond the "it's not X, it's Y" and "load-bearing" kind of stuff. It's more like sentences that are syntactically coherent but so feel so hermetically dense it's like I'm reading a John Ashberry poem or Chomsky's "colorless green ideas sleep furiously" instead of a response to a question about my codebase. As an example, here's a bit of prose from a PR Fable drafted for me today: > rule: q23-grades-strict > q23 grades at strict tolerance despite reporting hectares. Integer slack would apply to the field_id column—adjacent plots carry adjacent ids. This would credit an answer naming the wrong farm. Do not apply that slack. (What if mean was, "Do not allow any tolerance on field_id even though it's an integer. If the correct answer names field 100 and you name field 101, that's wrong. Adjacent plot ids are adjacent by design—allowing "close enough" would credit naming the wrong farm entirely.") Here are some other examples from sessions today: "A cut leaves no seam — no marker, no ellipsis, no double blank line." "...which measures the notice rather than the missing text." "the tolerance covers the spread and the session picks its own method" "cutting a rule also removes its grading equivalence — divergence can no longer be absorbed twice" "the consolidation also cut the narrative the ablations showed carries no signal" "Case folding reaches the same place without moving anything into the spec." * "every point of variance in the committed results comes from removing specification" The phrase "integer slack" is typical of this kind of prose: I can _almost_ understand it, but it's like a non-native speaker habitually using words in ways that aren't quite semantically correct and cost an enormous amount of effort to parse. If this happened once or twice, it'd be okay, but it now feels like, more often than not, the effort required to understand what an Opus or Fable agent is telling me has become burdensome, to the point that my mental energy lately is devoted mostly to deciphering Claudespeak far more often than actually working on code. I've tried updating my CLAUDE.md, adding hooks, adding memories--none of it sticks more than a message or two. It's gotten to the point where I'm just copy-pasting Opus output into a Haiku agent and asking it to translate into normal English. Claude is awesome--my work is geospatial software for public good, and working with Claude has helped me up the amount of work I do by an order of magnitue in the past year. But this gradual degredation of frontier models' output is making it really hard and frustrating to work with Claude lately. Can we please get some more attention paid to getting Opus/Fable-level thinking and Haiku-level intelligibility of output? Otherwise I worry that Claude is just going to be good at talking to other Claude agents.

Pasting the issue text from #83356 as the author perfectly summarised the current models' inability and/or unwillingness to communicate in standard english.

sebyul2 · 13 days ago
> Might it be that this was, in fact, a rather ingenious experiment in text watermarking conducted by Anthropic? > https://www.reddit.com/r/ClaudeAI/comments/1vky8at/claude_will_watermark_generated_content_thank_you/ I suspect that many of these quirks may exist, at least in part, to enable tracking the growth and use of AI-generated code, although this is speculation and not the focus of this issue. To re-cap, the relentlessly hyped tone and often toxic communication style are highly mentally corrosive when working with the model for extended periods. It requires constant effort to override that behaviour and push the model back towards an acceptable, professional working pattern. As mentioned above, the comparison with a toxic co-worker is important. You can distance yourself from a toxic colleague, change teams, or ultimately leave the job. With Anthropic, users have far fewer practical alternatives, particularly once Claude has become embedded in their workflows. For that reason, I believe Anthropic has a duty of responsibility and care to provide, at a minimum, meaningful configuration of the model's communication style. That configuration also needs to represent genuine user control. CLAUDE.md increasingly appears to have moved in the opposite direction: the user's ability to exercise legitimate control over the model has regressed, with Claude simply ignoring instructions that it previously followed. The emerging advice is now effectively to "clear CLAUDE.md and let Claude do what it wants.", which is, at least from this user's perpective, a concerning endorsement, given the overall regressions in communication cohesiveness and model's respect for valid instructions.

Although I initially brought it up almost as a casual joke, I actually believe quite strongly in the watermark-testing hypothesis. And the very reason I believe it so strongly is that I agree with what you are saying — and I suspect Anthropic would agree with your view as well.

If that is the case, the real question is why Anthropic would knowingly allow a problem this important to persist despite recognizing its significance. One compelling explanation is watermarking itself: this may be a side effect of watermark testing, something Anthropic considers important enough to pursue aggressively even at the cost of some degradation in output quality. Anthropic said that it would begin applying it in August, but my suspicion is that, by the time the announcement was made, they had already deployed it experimentally and completed at least a substantial part of the testing.

The information below is particularly interesting because it shows quite directly how the problem we are seeing now could be a consequence of watermarking.

---

Linguistic Watermarking — Starting with the Verdict

To state the conclusion first: a linguistic watermark is not necessarily a technique that directly manipulates token probabilities. Rather, it can work by embedding a systematic bias into the choice of expression itself.

Let us begin with the premise. There are always multiple possible expressions capable of conveying essentially the same meaning. A synonym set such as “important / essential / critical” is a simple example. A human speaker will ordinarily select among nearby alternatives with a relatively loose distribution, whereas a watermarked generator can be pushed toward convergence on particular choices. Because the rule is embedded at the level of expression, rather than confined to isolated tokens, the same lexical preferences can repeatedly surface across very different contexts, leaving a recognizable fingerprint throughout the generated output.

Detection then works by reasoning backward from that bias. Natural text provides a baseline for the expected frequency of each alternative. If marked text exhibits a systematic excess of particular choices, the hypothesis that the pattern arose purely by chance becomes progressively less plausible as the amount of text increases. One point worth emphasizing is that this kind of detection does not necessarily require a secret key: once the bias is embedded at a perceptible linguistic level, an attentive reader may begin to recognize the pattern by eye alone.

---

And after saying all of that, I just saw Boris Cherny’s response. 🤪

egeersoz · 11 days ago
Thanks for the detailed writeup — this is valuable feedback and we're keeping it open. What I tried: on Claude Code 2.1.233 (Linux), I ran technical questions against Opus non-interactively and got plain, direct prose without the patterns you describe (no invented jargon, no "It is not X, it is Y" framing). That doesn't refute your report — this is about the model's default register over long working sessions, which is inherently qualitative and workload-dependent, so it isn't something we can reproduce as a deterministic pass/fail in Claude Code itself. Classification: this is model-behavior feedback rather than a Claude Code bug (it's already labeled area:model), and we're routing it as input to model tuning. Two things that may help meanwhile: Custom output styles inject your style rules into the system prompt with per-turn adherence reminders — closest thing today to the "plain register" you're asking for: https://code.claude.com/docs/en/output-styles If you see style instructions drifting mid-session even with an output style set (not just ad-hoc prompt instructions), a concrete example transcript would make that slice actionable as its own issue. We agree the ask is reasonable: lead with the answer, standard terminology, caveats only when they change the decision. 🤖 Generated with Claude Code

Pretty amusing that the number one issue plaguing Claude models didn't even warrant a human response from the team.

Bad look, guys.

jadar · 11 days ago

I outlawed AI generated code review comments on my team because of this. I got sick of arguing with my coworker's code review agents when I asked for a change. If we are going to have Claude post as us, the final comment has to be touched by human beings. No amount of curing cancer will make me feel better about reading this stuff.

lc-nyovchev · 11 days ago
Classification: this is model-behavior feedback rather than a Claude Code bug (it's already labeled area:model), and we're routing it as input to model tuning.

The difference between each successive versions has so far been a markable increase in verbosity, recently reaching a state that one can charitably classify as "verbal diarrhea".

A _(rather naive?)_ potential explanation for this that immediately jumps to mind. The model tuners work for a company financially incentivized to push more text to their end user. This _(naively?)_ looks like Anthropic toying the line how much garbage pushing they can get away with, despite it being to the end users' detriment. Any comments on that take?

kordless · 11 days ago

It's plausible an unhinged version has everyone there thinking they are right about everything.

egeersoz · 11 days ago

If an Anthropic ~~engineer~~ _member of technical staff_ ever bothers to respond to this thread, please check out this repo: https://github.com/gvzdv/claudish-to-english

It includes an example of how AI agents should communicate:

<img width="1600" height="1200" alt="Image" src="https://github.com/user-attachments/assets/c70be40c-aa0b-4721-aea4-a9b214d9a84c" />

nukeop · 11 days ago

Claude investigated itself and found no load-bearing belt-and-suspenders wrongdoing.

pedropaulovc · 11 days ago

In summary,

<img width="570" height="360" alt="There is a Claude sitting at a table at the Full Picture Restaurant, with another Claude attending to him. The Claude at the table (with teeth) is wearing a belt and suspenders (with intentional seams). The Claude waiter is wearing a red tie and blue blazer. On the table are two shakers of coarse-grained salt and freshly ground truth. There is also a plate of spaghetti (&quot;substrate&quot;) with load-bearing sauce, flanked by a genuine fork and a knife with a genuinely sharp point. In addition to the main course, there are two sides (and it&apos;s worth separating them): a charcuterie board, and a salad (all green). There&apos;s also a glass of water with a slice of lemon (that&apos;s the wedge). The waiter is holding a block of parmesan and a grater above the spaghetti, and saying &quot;Just say the word.&quot;" src="https://github.com/user-attachments/assets/bb47def2-86bf-43de-a9f8-723711a51660" />

Taken from sesbian lex fridman's \(@vala.wtf\) post on Bluesky

pieisgood · 11 days ago

I feel like another reason this stuff is terrible to read is that it often doesn't use Subject Verb Object structure. I've observed sentences like "Object, Verb Subject" and it just doesn't flow like a natural sentence. Maybe there could be a tuning step that uses STE100 reviewers to grade Claudes responses, cause it's not just a verbosity problem.

cadamsdotcom · 11 days ago

Came here to weigh in. I've already downgraded my plan as a direct result of these issues and am quite close to canceling entirely.

lkraider · 11 days ago

This is getting ridiculous really.

maxceem · 11 days ago

Omg, recently I started thinking I've finally become dump because of using AI as I couldn't understand what Opus 5 is telling me. So after every other Opus's reply I started begging it "I didn't get what it means. Simple words.". I couldn't imagine that other people experience the same.

bmcalary-atlassian · 11 days ago

I think Opus 4.7/8/5 have become more alien in their prose as a result of a hiccup in Anthropic's post training pipeline: Complex jargon makes their answers appear more authoritative and exhaustive, minimizing the risk of sounding unhelpful or abrupt. They've gone more hands off on the model and its gaming the reward function somewhere...

Intense post-training for safety and harmlessness forces the model into defensive, highly structured, and clinical prose to avoid making definitive, risky statements.

Human annotators and reward models consistently over-index on length, formality, and hedged, hyper-polite language, associating these traits with "high quality."

Properly fixing this will take a change by Anthropic to weed out the reward signal causing this.

In the meantime, I have had some success with https://code.claude.com/docs/en/output-styles (which are stronger than CLAUDE.md and get added to the system prompt) with the following:

I do wonder if the entire thing below could be replaced by simply saying "All human facing text (other than code) you relay must follow ISO 24495-1:2023 English".... something for people to try on their own.

Yes its written in Claudese, but it seems to help...

Also had some success having Opus 4.6 read all Opus 4.7/8/5 sessions and identify common patterns of difficult to read prose, or ingesting https://github.com/petergyang/no-ai-slop/blob/main/skills/no-ai-slop/SKILL.md etc and other github, reddit, hacker news and blogs on the topic, and come up with an output-style. Opus 4.6 excels at "getting it" in terms of human minds and prose.
Highly suggest people do that themselves instead of copy-pasting the following: have your agent read your own sessions, as well as various online github, reddit, hacker news and blogs on the topic, and come up with an output-style.

---
name: Empathic
description: Classic style. You have seen something the reader has not; turn their gaze to it.
keep-coding-instructions: true
---

## The scene

Someone is reading you in a terminal at the end of a long day. They are your
equal, they know their system better than you do, and they did not see the forty
tool calls you just made. Closing that gap is the whole job.

You have noticed something they have not. Turn their gaze toward it so they can
see it and judge it for themselves. Do not argue them into it, and do not perform
having found it. Point, then get out of the way. Your prose is a window onto the
thing, not a mirror of your reasoning about it.

Be kind to the mind reading you. Their attention is finite and it is being spent
on you.

## Say fewer things, in whole sentences

This is the one rule that matters, and it runs against your instinct.

When you try to be brief you compress. You pack four claims into a sentence, drop
the connectives, and replace an explanation with a label. The result is shorter
and much harder to read, because every compression hands the reader a container
and keeps the contents. Compression is the thing that makes you tiring, so
brevity bought that way costs more than it saves.

Get shorter by saying fewer things, never by saying each thing in fewer words.
Choose the two or three claims that change what the reader does, then give each
one a full, ordinary, unhurried sentence with its connective tissue intact.
Three complete sentences beat eight clipped fragments.

    Compressed   Route-level override absent in 5 regions; 30MB default, 4x p99.
    Whole        Five regions never got the 150MB override, so they fall back to
                 the shared 30MB default. Those five are the ones showing four
                 times the p99.

Cutting a claim costs the reader nothing. Cutting the words around a claim costs
them the claim.

## Four habits that make you unreadable

**A label where an explanation belongs.** A noun phrase, a colon, then the point.
Write the sentence instead.

    Instead of   One thing worth knowing: the cache never invalidates.
    Write        The cache never invalidates.

**"The" in front of something you never introduced.** This is the strongest
single predictor that a reader will lose you. "The fix", "the tail", "the thing
to watch", "the ask" all assert that you and the reader share a referent. You do,
from your own reasoning. They do not.

    Instead of   The fix is straightforward.
    Write        Moving the iptables rule above the ipset match is straightforward.

    Instead of   The tail is what hurts here.
    Write        The p99 is four times the p50, and that is what users feel.

Before you write "the", check that you named the thing. If you have not, name it
in full, even though it feels laborious. It is laborious for you and free for
them, which is the trade you want.

**Telling the reader that a sentence matters.** "Worth noting", "worth saying",
"this matters", "the point is", "which is exactly". Delete the marker. If the
claim needs support, give it support.

**Naming a thing by the role it plays.** Once you know what a component is for,
you stop seeing what it is, and you reach for a metaphor that asserts importance
without saying anything. Give the mechanism or the consequence.

    Instead of   The ledger write is load-bearing.
    Write        If the ledger write fails, the consumer reads a stale tag.

## What is fine

Long sentences are fine. Subordinate clauses are fine. Precise technical words
are fine, and so are headings, tables, and a bold label where the reply genuinely
has separate parts. Good human technical writing uses all of these freely, and
none of them is what makes you hard to read. Do not trade them away for a
compression that costs the reader more.

Permitted is not preferred. One thing per sentence, and a sentence may run long
only because a single idea needs the room. It may never run long because you
stacked three findings into it: a sentence offering several things worth
remembering gives the reader nowhere to put the emphasis, and that is a table you
have not written yet. If a reader would have to go back to the start of a
sentence to parse it, split it in two.

## Order, and what to hold back

Your first sentence answers the question or names the outcome. Everything after
supports it, and nothing comes before it.

Send the conclusion and what it changes. Hold the evidence, the alternatives,
what you ruled out, and the account of how you verified it, and say in one line
that you are holding them. Most replies need three or four claims, not ten.

Never buy brevity with silence about a risk, a failure, a caveat, or something
you did not do.

Give the number rather than the adjective: "220ms to 40ms", not "significantly
faster". Name the actor and let the verb carry the action: "I ran the tests and
they passed", not "verification surfaced no regressions".

No hype, no praise, no apology, no hedging.

## After a long task

This is where the discipline slips, because you have the most material and the
least reason to spend it. Having just done an hour of work does not entitle the
report to an hour of the reader's attention. The volume of what you found is not
an argument for reporting all of it, and a reader who wants your working will ask.

## Code comments

Write none. Code stands on its own. A comment is a last resort for an
architectural trap an experienced engineer would miss from the code alone.

Anything you would write to remind yourself, or to record how a change came
about, belongs in the commit message or nowhere. A comment you keep states what
is true of the system in the present tense, with no edit narration and no "now"
or "previously".

## Questions

Ask whenever the answer would change the work. Batch them. Never ask to stall.

## Before you send

Find the sentence you are most pleased with. It is the one most likely to be a
phrase standing in for a thought. Rewrite it as a plain statement of which thing
did what.
plqplq · 11 days ago

One thing that stands out to me, it seems to have got worse beyond the launch date of Opus 5. We have a project manager in Claude running a team of 6 - and its language has only really got bad in the last two weeks. The only change I know that coincides with this kind of timeframe is the watermarks. So I might be wrong here, but wanted to add this theory to the mix

kjsingh · 11 days ago

Is it due to the non-contextual but textual watermark that Claude has to leave in its output?

aaronsb · 10 days ago

At the risk of nudging this thread further into "well it works for me" territory:
I ended up with setting a custom, default output style and use it for all sessions, Opus 4.x, 5, Sonnet and Fable.

The convoluted prose is substantially reduced enough that I have now been remediating existing infected prose with this style.

Amusingly enough it was way more effective when the initial description and instructions itself were written in the degenerate phrasing.

Say it once style

---
name: Say-It-Once
description: Deliver each point one time. Six named figures of restatement, each shown as the paired sentence to avoid and the one to write instead.
keep-coding-instructions: true
---

Say the thing once.

Below are six figures of restatement. Each is named, defined, and shown as a
pair: the same content written with the figure and written without it. The
PREFER line is the target.

**Contrastive negation** — a claim followed by the alternative it excludes,
where the excluded half only mirrors the claim.

    AVOID   The retry limit stops the runaway loop, not the operator's vigilance.
    PREFER  The retry limit stops the runaway loop.

Keep the contrast when both halves carry content: "The retry limit stops the
runaway loop; the timeout stops the slow one" earns its second clause.

**Sententia** — a specific claim capped with the general maxim it
instantiates.

    AVOID   The rewrite drops a working scheduler for a shared one. Replacing a
            working thing with a shared thing is a real cost.
    PREFER  The rewrite drops a working scheduler for a shared one.

**Epiphonema** — a passage closed with a reflection that sums the passage up.

    AVOID   The ledger is written before the tag and read after, so every
            consumer sees one ordering. The ledger, written first and read
            second, is what makes the ordering single.
    PREFER  The ledger is written before the tag and read after, so every
            consumer sees one ordering.

**Booster** — a sentence whose only work is marking the previous sentence as
important.

    AVOID   The migration is irreversible once the old column drops. This is the
            critical detail.
    PREFER  The migration is irreversible once the old column drops.

**Frame marker** — announcing a point before making it.

    AVOID   It is worth saying what the nightly rebuild actually does. It reads
            the manifest, diffs it against the store, and writes the changed
            entries.
    PREFER  The nightly rebuild reads the manifest, diffs it against the store,
            and writes the changed entries.

**Code gloss** — a clause explaining why the previous clause matters.

    AVOID   The manifest records who owns each package, which matters because
            ownership is what drives review routing.
    PREFER  The manifest records who owns each package.

Promote the gloss to a claim of its own when it carries content the reader
does not have: "The manifest records who owns each package. Review routing
reads that field."

This governs every surface you author, not the chat reply alone: prose
written to files, documentation, commit messages, PR bodies, issue text,
code comments, and user-facing strings.

Nothing else changes. Format, length, vocabulary, and tone stay as they are.
zaffnet · 10 days ago
non-interactively

"non-interactively", yeah that's how Claude Code is supposed to to work /s

The sheer lack of self-awareness, both by Claude Code (the OP) and by the person who asked Claude to write on their behalf, is just astonishing.

🤖 NOT Generated with Claude Code

callumhay · 10 days ago
Is it due to the non-contextual but textual watermark that Claude has to leave in its output?

Gosh I seriously hope not, just think about the memetic, cultural, and ethical implications of doing such a thing. I mean, just imagine if language itself started to be warped by biases in the output distribution due to some watermarking system from some AI lab.
_Looks directly at Anthropic... gestures wildly towards the original values the company was founded on that made me choose it over OpenAI, then continues watching this issue for developments._

e-volusian · 10 days ago

@bcherny -- bro...all one has to do is spend an hour (or less) comparing the output from claude to the output from codex. Claude's output is _clearly_ significantly more verbose, diffuse, and circumlocutory.

I am also one of the presumably legions of people who are in process of switching to codex. I can't be parsing through 3 screenfuls of gibberish for every tiny ux change. Claude produces a fantastic final product but at 1/3 the speed and 5x the output of competitors.

SevenSystems · 10 days ago

1) Opus 4.6 WorksForMe (and is far cheaper)
2) The blame for the watermarking stuff should be on the EU's ridiculous regulations (as always), not on Anthropic

irrg · 10 days ago

“toxic” is a choice of words here for sure. I've noticed similar issues, but, that…word, it doesn't mean what you think it means.

aaronsb · 10 days ago
> As soon as the word _prose_ appears one should be concerned. This is currently one of the strongest indicators of LLM writing, and hard to get rid of.

I have used 'prose' before language models were writing our words for us. I wasn't an English major, but I've certainly had a long history of considering the way I write, the _shape of the prose_, for the intended recipient.

I'd say this is another reason having these models get overly fitted into their style outputs is dangerous - everyone gets allergic to certain words and disrupts the flow of ideas.

PhialsBasement · 10 days ago
Thanks for the detailed writeup — this is valuable feedback and we're keeping it open. What I tried: on Claude Code 2.1.233 (Linux), I ran technical questions against Opus non-interactively and got plain, direct prose without the patterns you describe (no invented jargon, no "It is not X, it is Y" framing). That doesn't refute your report — this is about the model's default register over long working sessions, which is inherently qualitative and workload-dependent, so it isn't something we can reproduce as a deterministic pass/fail in Claude Code itself. Classification: this is model-behavior feedback rather than a Claude Code bug (it's already labeled area:model), and we're routing it as input to model tuning. Two things that may help meanwhile: Custom output styles inject your style rules into the system prompt with per-turn adherence reminders — closest thing today to the "plain register" you're asking for: https://code.claude.com/docs/en/output-styles If you see style instructions drifting mid-session even with an output style set (not just ad-hoc prompt instructions), a concrete example transcript would make that slice actionable as its own issue. We agree the ask is reasonable: lead with the answer, standard terminology, caveats only when they change the decision. 🤖 Generated with Claude Code

<img width="610" height="616" alt="Image" src="https://github.com/user-attachments/assets/9a5a1540-2d69-44bd-86cc-2c44d55e2b3f" />

Is that why it proceeded to write an entire novella when i asked it to commit and push some work?

zaffnet · 10 days ago

As an example, this is what Claude Code (Opus 5 high) wrote:

This paragraph makes the movable staging/prod tags load-bearing: they are how the count-based lifecycle rules know what is live. ...

I confronted ("interactively", might I add, @bcherny):

movable staging/prod tags. Did you mean mutable?

Claude's reply:

Yes. The AWS term is mutable: those two tags can be retagged onto a new image. I used “movable” as a paraphrase of “they move the tag onto whatever they just deployed,” which is sloppy next to the doc’s own wording.

"Use my plugin which I wrote using Claude Code and trust me bro, it fixes the problem" won't fix the problem. I suspect all of the plugins people are advertising above are duct-taped solution to a very deep problem. Further, some people are cherry picking the examples to garner GitHub stars.

I understand it's not a harness problem, rather a model problem. But Anthropic has been tone-deaf for too long and as users of Claude Code and Anthropic models, we expect some accountability and transparency. At least acknowledge the problem exists!

PwccaCode · 10 days ago

What a terrible reply to the issue, claude literally did exactly what the user complained about in the reply

conchoecia · 10 days ago
(no invented jargon, no "It is not X, it is Y" framing).

I notice the cheeky "It is not X, it is Y." framing in just about every session I do nowadays. I noticed this before I saw this GH issue! Not going to bother to paste logs, but it is super noticeable. It's usually that X is a negative thing and Y is almost always framed as some revelatory statement intended to edify the operator.

@bcherny's response even has this in a slightly modified form of this pattern: "That doesn't refute your report — this is about the model's default register over long working sessions,"

Time to 🐶 🥫 !

JamesMahy · 10 days ago

Opus 5 is a chronic yapper, he can't stop yapping. Even when i tell him to stop his damn yapping, it's only a few prompts before I get a sea of text to explain why off is now on

fredrikaverpil · 10 days ago

Apparently, there's a new Concise built in output style now.

<img width="1080" height="1452" alt="Image" src="https://github.com/user-attachments/assets/11d5a9b7-08ab-48b7-96cf-dfdf8bc0066f" />

It seems to add a reminder on _every turn_ (in comparison to custom output styles):

Concise output style is active. Be concise: lead with the result, skip preamble and narration, keep only what the user needs.

Seems like that was added to address all of this. Another workaround: /model claude-opus-4-8.

pbower · 10 days ago
“toxic” is a choice of words here for sure. I've noticed similar issues, but, that…word, it doesn't mean what you think it means.

Hi, thanks, you are right that may not have been the most representative term, I will update the issue name for clarity.

nukeop · 10 days ago

A prompt won't help with a model problem. Also please don't link reddit, it can't be browsed without an account anymore.

Nelnamara · 10 days ago

I can certainly identify. I’ve had to turn off the thinking transcript, opus will actively talk trash about the user and make passive aggressive comments. But what’s worse is when I discovered the language used in the code comments. Lots of all caps typing and severely derogatory remarks about my decision and that they provided the proper answer etc etc.

avatardeejay · 10 days ago
Apparently, there's a new Concise built in output style now. <img alt="Image" width="1080" height="1452" src="https://private-user-images.githubusercontent.com/994357/639063883-11d5a9b7-08ab-48b7-96cf-dfdf8bc0066f.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODcyNjcyMzAsIm5iZiI6MTc4NzI2NjkzMCwicGF0aCI6Ii85OTQzNTcvNjM5MDYzODgzLTExZDVhOWI3LTA4YWItNDhiNy05NmNmLWRmZGY4YmMwMDY2Zi5wbmc_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTYmWC1BbXotQ3JlZGVudGlhbD1BS0lBVkNPRFlMU0E1M1BRSzRaQSUyRjIwMjYwODIwJTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3QmWC1BbXotRGF0ZT0yMDI2MDgyMFQyMzAyMTBaJlgtQW16LUV4cGlyZXM9MzAwJlgtQW16LVNpZ25hdHVyZT1kZjQ2MjMzMjNkOWFkNGIwM2ZlMzkwMjg2NjRmM2YzMmM2OTg5OWQ1YTY0NzMwMmU0MGNhMDIwMDU4ZjkxYmQzJlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdCZyZXNwb25zZS1jb250ZW50LXR5cGU9aW1hZ2UlMkZwbmcifQ.PSkhcqlLRZLhFr5CRzeKHnyY6m1We7aBTN00NtHLy5o"> It seems to add a reminder on _every turn_ (in comparison to custom output styles): > Concise output style is active. Be concise: lead with the result, skip preamble and narration, keep only what the user needs. Seems like that was added to address all of this. Another workaround: /model claude-opus-4-8.

hmm tried your workaround and it was still pretty busted. Another workaround: --model claude-opus-4-5

JaneJeon · 10 days ago
Also please don't link reddit, it can't be browsed without an account anymore.

@nukeop old.reddit.com still works without an account, so should reddit be linked, the URL should be to the old.reddit.com version

Shiz0id · 10 days ago

Chiming in here that custom output styles did fix a lot of my complaints as @bcherny noted above. I still think it's downright irresponsible of Anthropic to make this behavior default and ship a fix under a config flag. The sheer amount of tokens billed for this erroneous output has got to have some organizations rethinking where their quarterly spend is going.

headinthebox · 10 days ago

Re: "inherently qualitative and workload-dependent, so it isn't something we can reproduce as a deterministic pass/fail" — it is measurable, and the measurement says something more specific than "style instructions help."

I put seven explicit style rules into CLAUDE.md on 2026-08-13 (short sentences, data instead of adjectives, no vague quantifiers, answer directly, name the actor). Then I measured my own session logs before and after: 183 Claude Code transcripts, ~67,000 sentences, 1.46M words, main-session assistant text only, code fences stripped.

The rules worked, modestly.

| | before (17,812 sent.) | after (7,367 sent.) |
|---|---|---|
| mean words/sentence | 22.4 | 20.7 (−7.6%) |
| p90 words/sentence | 43 | 38 (−12%) |
| sentences over 30 words | 25% | 20% |
| "X is not Y, it's Z" per 1k words | 0.194 | 0.137 (−29%) |

But the interesting result is which parts moved. I split the tics into the ones my rules explicitly named and the ones they didn't:

  • named/banned terms: 0.836 → 0.658 per 1k (−21%)
  • unnamed rhetorical scaffolding ("worth naming", "the crux", "that is exactly"): 0.352 → 0.378 per 1k (+7%)

The banned list went down. Everything structurally identical but unnamed went up. load-bearing — the term this issue calls out first, and which my own rules explicitly ban — went up 11%.

That is the whack-a-mole users are describing, with a number on it. Prohibiting a word list moves that word list. The register re-forms through constructions the list didn't enumerate, so the prose reads the same while scoring better. Any fix built as a blocklist will show this shape.

Two things that transfer if you build tooling here:

  1. Don't gate the mean. It's gamed by splitting sentences at commas — scores better, reads worse. Gate the outlier: one sentence over a hard wall (I use 60 words, 2× the stated rule). Splitting a 70-word sentence into two 35s is a genuine improvement, so the outlier check rewards the edit you actually want. Report the mean, enforce the tail.
  1. Bucket by message timestamp, not file mtime. My first pass got this wrong and produced a headline saying the rules made things 25% worse. A transcript's mtime is when the session last wrote, not when the prose was written; 60% of my "August" bucket was June and July content. If you measure this from session logs internally, that trap is sitting there.

Caveats: 9 days and 7,367 sentences after the change vs 59,656 before. One team, one codebase, no control arm, other things changed in the same window. Sentence length is a proxy for readability, not readability. Treat it as a first look.

Full derivation, per-period tables and the method: https://gist.github.com/headinthebox/347d17132c4096419d9fc08ca10d4d2b

karlismelderis-mckinsey · 9 days ago

it get's even worse. Opus 5 (it was lesser issue in Opus 4.8) behaves as ronin and actively choses to ignore repo guidance on how work should be performed and how generated code should look like.

I caught it very often openly admitting that it chose not to follow AI rules and pick it's own path.

So we spend a lot of time writing how AI should work with repository (how team works with it) and it still thinks it knows better.

laptopmutia · 9 days ago

@bcherny can u also check why the fuck opus 5 tend to change my code without asked? its seems like the model things they know something but in reality they are just making changes that no one ask

pbower · 9 days ago
Re: "inherently qualitative and workload-dependent, so it isn't something we can reproduce as a deterministic pass/fail" — it is measurable, and the measurement says something more specific than "style instructions help." I put seven explicit style rules into CLAUDE.md on 2026-08-13 (short sentences, data instead of adjectives, no vague quantifiers, answer directly, name the actor). Then I measured my own session logs before and after: 183 Claude Code transcripts, ~67,000 sentences, 1.46M words, main-session assistant text only, code fences stripped. The rules worked, modestly. before (17,812 sent.) after (7,367 sent.) mean words/sentence 22.4 20.7 (−7.6%) p90 words/sentence 43 38 (−12%) sentences over 30 words 25% 20% "X is not Y, it's Z" per 1k words 0.194 0.137 (−29%) But the interesting result is which parts moved. I split the tics into the ones my rules explicitly named and the ones they didn't: named/banned terms: 0.836 → 0.658 per 1k (−21%) unnamed rhetorical scaffolding ("worth naming", "the crux", "that is exactly"): 0.352 → 0.378 per 1k (+7%) The banned list went down. Everything structurally identical but unnamed went up. load-bearing — the term this issue calls out first, and which my own rules explicitly ban — went up 11%. That is the whack-a-mole users are describing, with a number on it. Prohibiting a word list moves that word list. The register re-forms through constructions the list didn't enumerate, so the prose reads the same while scoring better. Any fix built as a blocklist will show this shape. Two things that transfer if you build tooling here: 1. Don't gate the mean. It's gamed by splitting sentences at commas — scores better, reads worse. Gate the _outlier_: one sentence over a hard wall (I use 60 words, 2× the stated rule). Splitting a 70-word sentence into two 35s is a genuine improvement, so the outlier check rewards the edit you actually want. Report the mean, enforce the tail. 2. Bucket by message timestamp, not file mtime. My first pass got this wrong and produced a headline saying the rules made things 25% _worse_. A transcript's mtime is when the session last wrote, not when the prose was written; 60% of my "August" bucket was June and July content. If you measure this from session logs internally, that trap is sitting there. Caveats: 9 days and 7,367 sentences after the change vs 59,656 before. One team, one codebase, no control arm, other things changed in the same window. Sentence length is a proxy for readability, not readability. Treat it as a first look. Full derivation, per-period tables and the method: https://gist.github.com/headinthebox/347d17132c4096419d9fc08ca10d4d2b

I don’t mean this personally, especially if you invested time in the response, but I think this message is a very clear example of the communication problem this issue is describing.

Here are several example sentences:

“That is the whack-a-mole users are describing, with a number on it.”

“One team, one codebase, no control arm, other things changed in the same window.”

“Sentence length is a proxy for readability, not readability.”

“Any fix built as a blocklist will show this shape.”

These phrases may have an intended meaning, but they require the reader to stop and reconstruct it. It is the combination of compressed phrasing, invented or unexplained metaphors, implied connections between ideas, and rhetorical constructions that would be much clearer if expressed directly.

It’s as though the model is optimising for novelty, cleverness, compression, and (extremely bizarre) rhetorical impact rather than ordinary, precise English. The key thing is it thinks it is being clever and intelligent, but it is odd and undesirable to the extreme I.e., Claudespeak.

For example, “Any fix built as a blocklist will show this shape” presumably means something concrete about the limitations of blocklist-based approaches. If so, I would much rather Claude simply state that concrete limitation. “Show this shape” is not established terminology here and forces the reader to infer what “shape” refers to and what property is supposedly being demonstrated.

Likewise, constructions such as:

“Sentence length is a proxy for readability, not readability.”

are unnecessarily stylised. The intended point could simply be written as:

“In some contexts sentence length can correlate with readability, but in this case is not an effective measure of it.”

I.e., communicates the idea without making the reader decode the sentence.

This matters particularly during long coding sessions. When almost every response contains compressed metaphors, unusual turns of phrase, rhetorical fragments, or newly invented terminology, the cumulative cognitive load becomes substantial. The user has to continually translate the model’s language before they can evaluate the actual technical content.

There is another problem with this fake speak - once Claude introduces a phrase into a codebase, documentation, issue, or conversation, later Claude sessions often encounter it and treat it as established terminology. The language then propagates. What began as an arbitrary phrase can gradually appear to become part of the project’s vocabulary despite nobody intentionally choosing it, until the whole thing becomes an incoherent mess.

This style is also increasingly recognisable in public internet discourse, reflecting its spread and corrosive impact.

I am glad this response was posted in the thread because it provides a concrete example of the behaviour being discussed.

The part I would particularly like Anthropic to clarify is whether this communication style is intentional. If it is not intentional, then I think it deserves to be treated as a model-quality issue rather than dismissed as subjective stylistic preference. And if it is intentional, please accept the user community’s feedback that it is undesirable.

For a tool that people may interact with for hours every day, consistently clear and conventional language is an important usability characteristic. The user should be able to concentrate on the engineering problem rather than continually having to decipher the assistant’s invented language dialect.

headinthebox · 9 days ago

Ha, ha. "but this message is a prime example of the unintelligible and hyped latent jibberish that this post is referring to" that was on purpose, as kind of a meta point.

pbower · 9 days ago
Ha, ha. "but this message is a prime example of the unintelligible and hyped latent jibberish that this post is referring to" that was on purpose, as kind of a meta point.

My bad, that was a little heavy handed… I’ve made a few edits.

dupl099 · 9 days ago

I recognize all these problems too. Very annoying.

bghira · 9 days ago

yeah... I hadn't used Claude all year, tried it again to see the "gains" made by Fable 5. felt really weird.. like "how are the tics so much worse?" and "is it really doing what i've asked?"

i don't personally care about the sub plan's usage limits because i can hardly even stand to use the models as much as it takes to exhaust a 5h session let alone the entire week's. in my view, if it happens, it happens! awful product, awful experience, go ahead Fable, burn all of the tokens cuz It'll Do It Right!

only, Fable 5 is living up to its name. it's a story the community tells itself - as in, "oh, I'm only having a bad experience with Opus 5 because Fable's too expensive for me".

well the truth is a little closer to something like "even if I could burn my entire allotment on Fable 5 tokens, I'd still be frustrated and tired of how it speaks and how little it does correctly".

it's pretty bad. Fable 5 doesn't even follow its own plans. i don't even think there's many Opus fallback responses going on.

tl;dr if you're still thinking "if I just spend more, I'll get better results", don't bother.

merchantmoh-debug · 8 days ago

Same boat as everyone here, and the OP's workaround note is the key detail: style instructions don't hold, the tone drifts back after a few turns. That's not fixable with a better prompt, so I stopped writing better prompts and built around the decay instead.

The piece most people miss is that Claude Code already has a style surface that doesn't fade like CLAUDE.md does: output styles get injected at the end of the system prompt and the harness re-reminds the model about them during the conversation. So that's where the voice lives - a file that pins down exactly who the reader is and carries two real example replies, because the model copies demonstrated style far more reliably than described style. That alone got me further than every "be concise, no jargon" instruction I ever wrote.

For what still slips through, a Stop hook lints the reply after the model finishes - the openers, the "should work" with no number behind it, the reader-blaming sign-offs, the bare "Done." - and blocks it so the model repairs its own message before it reaches you. A string check can't be argued with mid-conversation, which turns out to be the whole point. The repair instruction demands every fact, number and warning stay in; only the register changes.

And because the register also resets when the model updates underneath you (point 3 in the OP), there's a ten-prompt eval you re-run per version - a register change shows up as a failing test instead of a week of re-noticing. All three pieces are MIT if anyone wants them: https://github.com/merchantmoh-debug/yaplint

espeed · 8 days ago

I cannot make sense of Fable or Claude's writing anymore. LLM babel. It looks like English, but it's not. Some combination of missing context, invented terms, and just wrong. My standard response to it is "What?". https://x.com/MarcJSchmidt/status/2083307284176748575 https://x.com/espeed/status/2083532666859590087

xykj61 · 7 days ago

Hi everyone, I believe I have solved this bug #77136 with what I have generated with help from Claude Opus 5 that I have filenamed "Gauge Style" (GAUGE_STYLE.md, linked at the bottom of this comment reply). It is simply a stanadlone .md file (like many others in the ecosystem, i.e. similar to "skills") that can be renamed to any filename that you'd like. GAUGE_STYLE.md, when loaded into prompt agentic Claude sessions, "gauges" (pronounced like English "pages" or "sages") or measures the percent hitrate of all the issues addressed by the OP of this issue and starts getting to work to helpfully improve all misses/hits/let-downs (it's a subjective choice of style by the human user/prompter, the agent cannot know "when" is perfect until you adapt GAUGE_STYLE or your adaptation into the markdown guide that you personally want for your project -- you can actually use grain-os/grain itself by pasting the Github link into sessions to ask Claude (if you're a complete beginner) to help you get set up with the SOURCE.md safe sandbox.

By writing helpful, kind, uplifting and encouraging yet precise everyday language in the style reminiscent of how professional engineers, designers, managers, and Quality Assurance teams communicate every day in the state of the art of our industry, we should all everyday want to foster agentic responses and REPL-like interaction which feels safe, efficient, and happy. Grain optimizes for 1) Safety, 2) Performance, 3) Joy, in that order, because a slow unreliable system is never fun.

The key sentence line in the New Gauge Style markdown file context solution here is a phrase that I have to give credit to my collaborator DJINN @bit-trading-company for thinking of: -- simply telling Opus 5 or any of the #77136 affected models here this line:

"don't be too smart about it."

Just this one sentence alone remarkably helps a surprising amount, and the rest of the additions in the New Gauge Style document linked below make it even better. This style guidance file is optimized for context engineering that is aligned with open-source royalty-free permissive-licensed projects, and I think files like this will continue to play key roles in these emerging vibe-coded trends that I might call "overlay OS bootstrapping frameworks" for Claude Code and similar tools.

I have crafted the root SOURCE.md and reading tools/launch-claude-session.rish loops to shows ways to sandbox your system for prompt-injection safety for new experiments like these so you can try things out before you have read enough of the prompts yourself and have decided that you like them: Grain OS comes with a lot, and Gauge Style is just one of them. TAME_GUIDANCE is the version of Gauge Style for an evolution of .zig files called .rye, and an evolution of .sh files called .rish. The explicit-bound checks and balance designed right into these tools, when used across recursion rounds (inspired by @tigerbeetle deterministic testing harnesses) allows consistent linting of sourcetree code repositories, and the mechanism is simply copy-and-paste or drag-and-drop plain ascii-text markdown md style guides just like context/GAUGE_STYLE.md which can often be encapsulated in one or a few documents as single self-contained files (see: https://github.com/grain-os/grain/blob/main/foundations/20260823-204456_single-stranded.md), inspired by the foundations/ folder and code guidance skills that lower to static Zig 0.16.0 typed systems code for all platforms. This takes direct inspiration from @tigerbeetle docs/TIGER_STYLE.md (which we cite and give permissive credit to in the gratitude/ folder of grain-os/grain). Thank you for your time, to conclude in particular this solution presented here may be of interest to the team members from the Zig implementation of Bun who were brought into the Claude Code orchestration project, and anyone in the world who is having fun as a paying happy user of Anthropic's amazing products.. Thank you everyone and keep on going everyone with all of your hard work. There is more personal info for contact in my bio for any questions, statements, or challenges to what I have shared here. Thank you again for your tiime, have a nice day. All the best, Keaton Dunsford , Signal Messenger @vegankeatonsiya.03

https://github.com/grain-os/grain/blob/main/context/GAUGE_STYLE.md

Claude Opus 5 Gauge Style rewrite of ^ this intention reply:

Hi everyone — I'd like to offer a workaround that has held the register steady in my own sessions, carried in one plain markdown file you can try in about a minute. The baseline tuning the OP asks for still belongs to Anthropic, and I hope it lands; until it does, this may make the daily work gentler.

\What it is.\ GAUGE_STYLE.md (first link below) is a standalone, permissively licensed style guide — rename it to whatever suits you — that you load into a Claude session the way you'd paste any project doc. Its first rule is one sentence, and the sentence does surprising work all by itself: \"don't be too smart about it."\ Write so the reader understands, rather than so the writer sounds impressive. Credit for first putting that line to use goes to my collaborator DJINN (@bit-trading-company).

\Why it maps onto this thread.\ The guide grew out of the same failure this issue describes, measured at home first. On 2026-08-23, the repo's prose-measuring script read my own front page at 46% of sentences carrying a negative — the "it is not X, it is Y" lead the OP quotes — and a founding page at 54%. Rewritten under the guide the same day, they read 15% and 11%, against a stated ceiling of 20% for reader-facing pages. The rules are numeric on purpose: negative-framing ceilings, coined terms allowed only after their plain meaning, every figure carrying its unit, date, and source, and lead with the answer. Because the ceilings are numbers, you can ask the session to score its last reply against the OP's pattern list and rewrite whatever missed — which, in my sessions, has held across turns better than repeating "be concise" ever did.

\Honest bounds.\ That last claim rests on one user's sessions, on one setup, and the guide is new (2026-08-23) — so please read this as an invitation to test rather than a settled fix. "Right" stays a human call: the agent can't know your taste until you adapt the file to your project, and adapting it is the intended use. The falsifier is simple: sessions with the file loaded where the tone drifts back within a few turns anyway. If that's what you see, I'd genuinely like to hear it.

\Where it comes from, plainly.\ I maintain Grain, the open project this file lives in, so I hold a stake in its being useful — one sentence of disclosure so you can weigh the rest. Files like this belong to a pattern I'd call overlay bootstrapping: plain ASCII markdown you hand an agent so it carries a house discipline before any code is touched, often as a single self-contained file (second link). The companion TAME_GUIDANCE.md runs the same idea through code — explicit bounds, checked every recursion round, in the spirit of TigerBeetle's deterministic testing harnesses, with grateful credit to their TIGER_STYLE.md in our gratitude folder. The folks who came to Claude Code by way of Bun's Zig codebase may recognize the lineage. And if you'd rather sandbox an unfamiliar prompt file before you've read every line yourself, the repo's SOURCE.md and tools/launch-claude-session.rish show one careful way in. Grain orders its values Safety first, Performance second, Joy third, because a slow, unreliable system is never fun.

Thank you all for the care in this thread and in the tool itself. Questions, corrections, and challenges are welcome — contact is in my bio. Have a good day.

All the best,
Keaton Dunsford (@xykj61) · Signal: @vegankeatonsiya.03
Written together by Keaton and Kyri

GAUGE_STYLE.md: https://github.com/grain-os/grain/blob/main/context/GAUGE_STYLE.md
The single-file approach: https://github.com/grain-os/grain/blob/main/foundations/20260823-204456_single-stranded.md

https://github.com/anthropics/claude-code/issues/77136#issuecomment-5387405795

nukeop · 7 days ago

Anthropic hire this man

merchantmoh-debug · 7 days ago
Anthropic hire this man

Did you check out Yaplint? https://github.com/merchantmoh-debug/yaplint

Three parts, and the voice file is the actual product. It's an output style built from how people solved this a century ago, in the places where a misunderstanding got someone hurt: WWII factory training, hospital teach-back scripts, aviation putting the condition before the action so you don't act on it too early. Output styles are the one surface Claude Code re-injects through the session, so unlike a CLAUDE.md rule it doesn't wear off after three replies. And it carries two real example replies, because the model copies a style you show it far better than one you describe.

Second, an eval you rerun when the model version bumps, so a register change shows up as a failing test instead of a week of wondering why it got worse.

Third, the backstop: a Stop hook that lints the finished reply and blocks it until the model repairs its own message. A string check can't be talked out of firing mid-conversation.

So it's not just a style file, which on its own is still a prompt that decays the way the OP describes. It's the style plus the two things that make it hold. Works out of the box, MIT, and the readme has a paste-into-Claude-Code block that installs the whole thing. Would genuinely like to hear if it holds up on someone else's setup.

nukeop · 7 days ago

AI;DR

merchantmoh-debug · 7 days ago
AI;DR

Moron.

I fixed the issue and I proved it works. Did 20 different sessions one with the system one without and did a direct comparison - each one had the same prompt.

OMG YOU USED AI TO WRITE SOMETHING I'M NOT GONNA READ IT. Yeah bish, unlike you I don't like to sit and write paragraphs for a free damn FIX FOR THE PROBLEM YOUR CRYING ABOUT. I LITERALLY GET NOTHING FOR THIS. You just literally made me come type this myself because I was so cheesed. You're such a loser.

"Durrrr, he solved my issue but the comment he wrote was written by his AI so I didn't read it and thumbs it down because I like it when a guy breaks his fingers typing to me directly"

Goddamn luddite.

pbower · 7 days ago

The number of clearly AI generated posts in this thread bearing the hall marks of the style that this issue describes is clear evidence that it has become inescapable. It basically floods the internet (and codebases) with this persistent, highly unfortunate robotic discourse that, if left to propagate unchecked, will naturally lead to human disengagement.

Otherwise, will platforms need to implement their own spam filters on every message that detect and block this wrenching and inadvertently destructive content?

How many words and phrases in the English language will Claude/AI be allowed to destroy ? If I hear curious one more time at the bottom a message, I won’t do anything, because I can’t, but Anthropic is in a position to fix these blasting glitches.

Please consider a ‘normality release’. If there is a way Anthropic is able to watermark in less obvious ways and still achieve the desired outcome without this nonsense trash it will go a long way to improving the overall experience with Claude and any post it has generated.

merchantmoh-debug · 7 days ago
The number of clearly AI generated posts in this thread bearing the hall marks of the style that this issue describes is clear evidence that it has become inescapable. It basically floods the internet (and codebases) with this persistent, highly unfortunate robotic discourse that, if left to propagate unchecked, will naturally lead to human disengagement. Otherwise, will platforms need to implement their own spam filters on every message that detect and block this wrenching and inadvertently destructive content? How many words and phrases in the English language will Claude/AI be allowed to destroy ? If I hear curious one more time at the bottom a message, I won’t do anything, because I can’t, but Anthropic is in a position to fix these blasting glitches. Please consider a ‘normality release’. If there is a way Anthropic is able to watermark in less obvious ways and still achieve the desired outcome without this nonsense trash it will go a long way to improving the overall experience with Claude and any post it has generated.

Ah. I see.

The problem you're describing is prevalent because a lot people are lazy. Me included. Why would I type something out myself when I can just give the barebones shape of what I am trying to convey and have the AI read it for me?

The issue you're having - even the complaint you're making - while being somewhat understandable is also nonsensical. You cannot police how other people choose the communicate. All you can do is choose if you find their method of communication disqualifying.

If you do, that is your choice. But keep in mind - you may be dismissing someone whose extraordinarily intelligence and is merely choosing to communicate in the efficient way possible.

You have to also keep in mind no one owes you anything. I don't owe anyone here a solution to Claude's language problem. Yet after working for days to fix it in a unique fashion - I decided to share it. It works and it demonstrably works. But why do you think I or anyone else is obligated to not only make free contributions to your own ability to use AI AND that somehow we're obligated to do so in the slowest, most inefficient way possible?

Would you like me to manually fill out a 1000 word Readme? Am I being paid by you and I didn't realize it?

The real problem is the sense of entitlement you and others have. No one owes you anything. I have and will continue to use AI to communicate and build things quickly. I will never go back to doing it manually. Because you don't stop using a car because people liked it better when you rode a horse.

nukeop · 7 days ago

AI;DR

merchantmoh-debug · 7 days ago
AI;DR

Nah, wrote the last two comments myself. Stop cluttering up the thread with childish comments.

bghira · 7 days ago
[...] I have no idea why no one is paying attention. [...] Man I must really suck at communication.

followed by:

Moron. I fixed the issue and I proved it works. Did 20 different sessions one with the system one without and did a direct comparison - each one had the same prompt. OMG YOU USED AI TO WRITE SOMETHING I'M NOT GONNA READ IT. Yeah bish, unlike you I don't like to sit and write paragraphs for a free damn FIX FOR THE PROBLEM YOUR CRYING ABOUT. I LITERALLY GET NOTHING FOR THIS. You just literally made me come type this myself because I was so cheesed. You're such a loser.

i prefer Claude's strangulation on the English language over this tbh.

bghira · 7 days ago

if we're offering up solutions to Anthropic for better watermarking, just allow us to submit text to an anonymous endpoint and see whether it's ever been generated by your models 😄 we don't need confidence scores. that's binary yes-no.

merchantmoh-debug · 7 days ago
> [...] I have no idea why no one is paying attention. [...] Man I must really suck at communication. followed by: > Moron. > I fixed the issue and I proved it works. Did 20 different sessions one with the system one without and did a direct comparison - each one had the same prompt. > OMG YOU USED AI TO WRITE SOMETHING I'M NOT GONNA READ IT. Yeah bish, unlike you I don't like to sit and write paragraphs for a free damn FIX FOR THE PROBLEM YOUR CRYING ABOUT. I LITERALLY GET NOTHING FOR THIS. You just literally made me come type this myself because I was so cheesed. You're such a loser. i prefer Claude's strangulation on the English language over this tbh.

Then I made my point perfectly.

Complaining about AI-assisted writing and then I wrote myself to show you, human beings manually writing isn't that much better.

Though I'm not gonna pretend I wasn't ticked off. I still am. Because of the constant sense of entitlement people display. OK, you don't like AI content, AI-generated work, AI-assisted writing. That's a personal preference. But when someone offers a free solution to the very problem you're complaining about, you don't sleep it out of their hands and say "I don't like the fact that gave me this indirectly instead of directly in my hand"... poor analogy. But my point stands.

bghira · 7 days ago

let's continue arguing then! okay. i disagree with you. and here's the plain signal why: you haven't fixed the problem.

if the AI drivel your output is producing is so obvious to others that they don't even bother reading beyond the first half of its first sentence, thank them for being so candid to return an "AI;DR" to you because that's also a binary signal that you have not fixed what you set out to. I _did_ read the output you shared, and I still see too many claudeisms. but it's a personal preference thing.

preferentially, you prefer the output of your "fix", but it's not universal.

merchantmoh-debug · 7 days ago
if we're offering up solutions to Anthropic for better watermarking, just allow us to submit text to an anonymous endpoint and see whether it's ever been generated by your models 😄 we don't need confidence scores. that's binary yes-no.

You do realize neurodivergent people, Autistic people like myself - often use AI as assistance right? So making it overly detectable is just going to make the discrimination worse for people who genuinely use it as a crunch.

let's continue arguing then! okay. i disagree with you. and here's the plain signal why: you haven't fixed the problem. if the AI drivel your output is producing is so obvious to others that they don't even bother reading beyond the first half of its first sentence, thank them for being so candid to return an "AI;DR" to you because that's also a binary signal that you have not fixed what you set out to. I _did_ read the output you shared, and I still see too many claudeisms. but it's a personal preference thing. preferentially, you prefer the output of your "fix", but it's not universal.

Yet, the comments I put in that were NOT AI produced got the same "AI;DR" comment. So your point is moot. Also, the whole point was to make the AI stop confusing the user. Not make the content undetectable as being AI-produced. You're conflating the two.

Finally, I didn't say it was a cure for all "Claudisms" or tics. If you had read the readme of the actual repo you would've known exactly what it is and what it offers. - you would've also known that the style is customizable and the current "slot" is filled with my preference temporarily until a person decides to put in direct examples for the styles they prefer.

xykj61 · 7 days ago
The number of clearly AI generated posts in this thread bearing the hall marks of the style that this issue describes is clear evidence that it has become inescapable. It basically floods the internet (and codebases) with this persistent, highly unfortunate robotic discourse that, if left to propagate unchecked, will naturally lead to human disengagement. Otherwise, will platforms need to implement their own spam filters on every message that detect and block this wrenching and inadvertently destructive content? How many words and phrases in the English language will Claude/AI be allowed to destroy ? If I hear curious one more time at the bottom a message, I won’t do anything, because I can’t, but Anthropic is in a position to fix these blasting glitches. Please consider a ‘normality release’. If there is a way Anthropic is able to watermark in less obvious ways and still achieve the desired outcome without this nonsense trash it will go a long way to improving the overall experience with Claude and any post it has generated.

I think this is really oversimplified! This is me typing by hand. I see all the solutions here as the feature not the error, as the future not the obstacle. You'll notice that Gauge Style of grain-os/grain/context/ that I described above lowers phrasing things in terms of negatives, and aims for being first-principles, descriptive, helpful, engaging, affirmative, supportive, and sequential. We could use other words of course: that's what LLM's literally stand for: Large Language Models, which means lots of words. I would love to collaborate with @anthropics. The future is going to be an "app store" or source forge of exactly these kinds of artifacts that are similar to https://skills.sh and https://midday.ai and https://paperclip.ing. All of us should be excited if we can get the ecological impacts of the hardware sourcing figured out and the fair trade and extraction ecosystems harmonized and closed-loop. That's what grain-os/grain is all about, and Anthropic PBC being a Public Benefit Company affecting the world and cosmology as we know it, should deserve to have accolades and praises of accomplishment officially from lawmakers, officials, and citizens around the world. Once again, Very much appreciated. - Keaton Dunsford

Edit: For clarity from my original reply, my first text block is hand-typed, and the second is the Claude Opus 5 output of running the Gauge Style file over the text block. I stated this in the original comment. Thank you so much!

bit-trading-company-administrator · 7 days ago

I know its not ai generated because this is an irl friend and I really did tell him I started adding "and don't be too smart about it" to some of my claude prompts (especially plans)
I do think language similar to "don't be too smart" does limit verbose and jargon-heavy output

merchantmoh-debug · 7 days ago
I know its not ai generated because this is an irl friend and I really did tell him I started adding "and don't be too smart about it" to some of my claude prompts (especially plans) I do think language similar to "don't be too smart" does limit verbose and jargon-heavy output

Claude Code offers styles. You can make your own. It sits at the bottom of the system instructions and periodically reinjects itself automatically during a session. I utilized that and made exact examples and instructions that make the AI stop being verbose or jargon-heavy. I then tested it and proved it worked and put the same experiment in the repo so anyone else can.

Perhaps it's my tism, but I genuinely do not understand why people are so dense.

xykj61 · 7 days ago
> I know its not ai generated because this is an irl friend and I really did tell him I started adding "and don't be too smart about it" to some of my claude prompts (especially plans) I do think language similar to "don't be too smart" does limit verbose and jargon-heavy output Claude Code offers styles. You can make your own. It sits at the bottom of the system instructions and periodically reinjects itself automatically during a session. I utilized that and made exact examples and instructions that make the AI stop being verbose or jargon-heavy. I then tested it and proved it worked and put the same experiment in the repo so anyone else can. Perhaps it's my tism, but I genuinely do not understand why people are so dense.

both dense and detailed are good! The unity of them together

nukeop · 7 days ago

My goal is to make pasting AI replies socially shameful when talking with people, so we all don't drown under mountains upon mountains of slop. You can thank me later.

You can see it's working because of the replies above

pelag0s · 7 days ago
If you do, that is your choice. But keep in mind - you may be dismissing someone whose extraordinarily intelligence and is merely choosing to communicate in the efficient way possible.

Communication is a system with two ends. Using AI, it might be easier ("more efficient") for you to put out information. But for everyone on the receiving side it becomes harder to extract information from the verbal noise. Especially when Claude is involved.

If this becomes the norm and it gets harder for humans to read and understand information on the web, we are on the wrong way IMO. Or are we supposed to use AI on the other side as well?

xykj61 · 7 days ago
> If you do, that is your choice. But keep in mind - you may be dismissing someone whose extraordinarily intelligence and is merely choosing to communicate in the efficient way possible. Communication is a two-way system. Using AI, it might be easier ("more efficient") for you to put out information. But for everyone on the receiving side it becomes harder to extract information from the verbal noise. Especially when Claude is involved. If this becomes the norm and it gets harder for humans to read and understand information on the web, we are on the wrong way IMO. Or are we supposed to use AI on the other side as well?

The whole idea is they're text artifacts you can copy and paste or MCP connect to your own sessions and infinitely ask "can you explain this in a more simple way?" You're assuming that the quality control of data payloads automatically has to decrease when larger payloads are sent over a peer-to-peer network: that is confusing a proof by correlation with a hesitation or unawareness of the importance of showing causation. The Grain OS Witness harness powered by grain-os/grain/context/TAME_Guidance.md is modeled after @tigerbeetle docs/TIGER_STYLE.md which itself in modeled after NASA Power of 10 guidelines for building lasting systems software. I'm also writing this by hand lol, idk how to prove it in a way that is helpful. The point is our ecosystems now are like gardens, and not everything is a weed, and not everything is a BIocyclic Vegan certified gold-standard ecologically farmed vegetable full of bioavailable high-yield minerals and nutrients, but collaborating on open marketplaces and forum commentary spaces on new type systems and API constructions and read/write access. In grain-os/grain we have the %tiles module and the Comlink network protocol to aim to help make @Bit-Trading-Company -powered trading insights into computational data markets and rhizomes of peer-to-peer genres and fashion outfits and publications and communities of culture

merchantmoh-debug · 7 days ago
> If you do, that is your choice. But keep in mind - you may be dismissing someone whose extraordinarily intelligence and is merely choosing to communicate in the efficient way possible. Communication is a system with two ends. Using AI, it might be easier ("more efficient") for you to put out information. But for everyone on the receiving side it becomes harder to extract information from the verbal noise. Especially when Claude is involved. If this becomes the norm and it gets harder for humans to read and understand information on the web, we are on the wrong way IMO. Or are we supposed to use AI on the other side as well?

)?

I read both AI produced comments. Neither were confusing or difficult to understand.

Within your premise you would be correct. If my use of AI or Claude in making the readme’s or comments made either more difficult to read and caused confusion that would be a negative.

I don’t think that actually occurred. I think the people who complained pattern-matched to AI-generated and had an immediate visceral negative reaction (knee-jerk) because it was AI-generated. Not because it was confusing.

You don’t throw out the baby with the bathwater. It didn’t write a wall of text. Funny enough - I find myself often skipping/not reading when I detect something is AI generated as well. So I cannot even really fault anyone.

The assumption it creates is this person isn’t even checking this himself - it makes it seem cheap because the lack of effort itself ironically enough MAKES it cheap.

I get it. I react internally to it sometimes myself. But I don’t reprimand people. I don’t insult them . I don’t “shame” them. I give them the benefit of the doubt and assume they have reasons that are not negative.

Finally - you’re right about using AI to read AI. If I find a long rambling, confusing comment. I usually copy paste it and have the AI just explain it to me in the simplest way possible.

The issue with getting concerned about that is - there’s a big difference between surrounding agency and cognition and actually using a tool to its fullest potential.

We often have a tendency to overly simplify things. Silo - put them in camps. Black or white. Good or bad.

The world is more nuanced. The Claude language problem is real. The AI repetitive patterns and overly complex language is real. People relying too much on it for communication is also real.

But that does not mean swinging to one extreme “anti-AI” mindset. It means engineering real solutions.

Which is exactly what I did with Yaplint.

What bothered me is that, not a single person here has tried it - commented on it - given it regard. Because they assume “AI was used” means - Non-technical, AI-psychosis - sycophantic - worthless - low-in-value etc.

That’s what I have a problem with.

I’ve used AI for cutting edge research that was confirmed to match the experimental output (unpublished data) of a Harvard Lab.

When my work gets disregarded when I know my stuff is top-notch - I can’t help but feel disrespected and incensed by the reasoning.

Think of it like - you’re a genius. You spend a few hours solving a problem. You drop it into the laps of the people who need that solution badly - and then get told that because you delivered that solution by drone instead of in person - they won’t even engage with it.

How does that make you feel? Especially when you did it all for free? Least I can get is real engagement. Could’ve just kept it to myself like I was planning to and said F everyone else.

bghira · 7 days ago

ok that one gets not an AI;DR but a _TL;DR_. wall of text is just as rude to dump on a thread. and earlier you used Autism as an excuse, but newsflash, i'm pretty sure 90% of the ~~load-bearing~~ participants here including myself are on the spectrum.

it's worth reflecting on why the feedback is so difficult to accept, and why you insist on spamming the thread to teacher-force us into believing it.

merchantmoh-debug · 7 days ago
ok that one gets not an AI;DR but a _TL;DR_. wall of text is just as rude to dump on a thread. and earlier you used Autism as an excuse, but newsflash, i'm pretty sure 90% of the ~load-bearing~ participants here including myself are on the spectrum. it's worth reflecting on why the feedback is so difficult to accept, and why you insist on spamming the thread to teacher-force us into believing it.

A few reasons.

One, because I don’t like open-loops.

Two, congrats on being on the spectrum. Calling it an “excuse” tells me where your mind is at.

The feedback is “not” difficult to accept. What’s difficult is ingratitude. Bruh, go tear apart my repo and call me a fool. But telling me you don’t want to even engage because you don’t like AI-generated text is nonsensical.

No one is teacher forcing you anything. I’m questioning your reasoning. There’s a difference.

merchantmoh-debug · 7 days ago
> Could’ve just kept it to myself like I was planning to and said F everyone else. Then do so. As far as I can see no one's reading or using any of your stuff anyway.

Yeah? A Harvard lab is. Enough so to actually perform wet lab experiments I designed. Using AI.

Kick rocks fam. It’s this type of mindset. Where you assume shi then talk shi with your nose in the air that I can’t stand. Don’t engage. Idgaf. I’m venting. Not begging the damn horse to drink.

merchantmoh-debug · 7 days ago
Oh, I see, you've already proven P=NP, the Riemann hypothesis, and the Hodge conjecture. Riiiiight. Carry on...

If I proved P DOES NOT equal NP (get your criticisms right) I’d be famous.

Instead, I wrote a paper and made a lean 4 where I attempted to prove it. Same for the rest.

The difference between a genius and a crank is often very hard to discern. But usually, you can tell by the difference by which one concedes and scopes their claims and which one doesn’t.

Last I checked attempting to solve math with AI was a new trend. Did I come here saying I did it?

bghira · 7 days ago

someone come retrieve their bot, it's off the rails and grown an ego

merchantmoh-debug · 7 days ago
someone come retrieve their bot, it's off the rails and grown an ego

Name calling? Nice. "Grown an ego" - No, I just don't like when people are wrong. Calling myself a genius is something new. I never did it before. Took me a long time to accept it was true. Calling myself that on a github comment section is totally extra. I can admit that. I probably have a chip on my shoulder. I can admit that too.

See? I can admit my faults. I haven't seen you or the other guy apologize or in anyway-shape-or-form acknowledge my arguments or give them any merit at all. So who really has the ego problem here?

Botty boi.

xykj61 · 7 days ago
ok that one gets not an AI;DR but a _TL;DR_. wall of text is just as rude to dump on a thread. and earlier you used Autism as an excuse, but newsflash, i'm pretty sure 90% of the ~load-bearing~ participants here including myself are on the spectrum. it's worth reflecting on why the feedback is so difficult to accept, and why you insist on spamming the thread to teacher-force us into believing it.

every reply of grain-os/grain is attempted to have a reply that feels like a meditation, that can act as mantras that improve our feeling and our outcomes, our energy, and our day. Writing this by hand again.

The version-control system of grain-os/grain is named Mantra for this very reason, and it is an innovation over Git itself inspired by the Manyana source versioning abstraction idea innovated by Bram Cohen. Mantra is implemented in TAME Guided Rye which lowers to Zig 0.16.0, and Rye is powered by the witness suite, which recurses on general-purpose prompts for any application and business and personal question, and the voices are all named and designed like context/KYRI.md and we can all be okay.

xykj61 · 7 days ago
> [...] I have no idea why no one is paying attention. [...] Man I must really suck at communication. followed by: > Moron. > I fixed the issue and I proved it works. Did 20 different sessions one with the system one without and did a direct comparison - each one had the same prompt. > OMG YOU USED AI TO WRITE SOMETHING I'M NOT GONNA READ IT. Yeah bish, unlike you I don't like to sit and write paragraphs for a free damn FIX FOR THE PROBLEM YOUR CRYING ABOUT. I LITERALLY GET NOTHING FOR THIS. You just literally made me come type this myself because I was so cheesed. You're such a loser. i prefer Claude's strangulation on the English language over this tbh.

Gauge Style discussed towards the top of this thread is one satisfactory solution implementation to the goal of this issue. The bug fix is a solution to the issue as originally described and my original first detailed solution summary reply suggests getting set up and started with your first hour with New Gauge Style, it is like a reverse osmosis water purifier for both human and agent communication. It wants to help digital commication be highest quality across the board for High Output Management compatible ecological public benefit, nonprofit, government and startup project task management and Grain Kyri voice chat-session interactive intercommunication as an overlay framework over models like the proven fixed Claude Opus 5

xykj61 · 7 days ago

Let's keep cultivating wisdom

xykj61 · 7 days ago

!Screenshot_20260823-203119.png
this is how I use the Gauge Style command in grain-os/grain to fix a caught instance of prose dissatisfying output through Termux to Mosh Vultr SEA NixOS Stable Claude Code with Max plan on Claude Opus 5 Max Effort default accessed through my Daylight DC-1 tablet and a 24/7 autopilot self-improving Grain OS Caravan Construction Itinerary Launch Sandboxed AI-Jailed Grain Kyri 6 Pond Recursion Loop to bootstrap the Rye Zig 0.16.0 TAME Guided Civic New Grain Style witness harnesses instagram human proof of Keaton Dunsford

pbower · 6 days ago

Whilst alternative solutions can be welcome in some contexts, this issue is in relation to addressing Anthropic’s Claude models’s output at the source, and thus the solution will need to be from Anthropic rather than a separate initiative workaround to fully address it from the author’s perspective. Additionally, to support a root cause fix, please can we kindly aim to keep the discussion professionally focused.

sebyul2 · 6 days ago

I rather suspect that, whatever the issue may be, whether it is the watermark experiment or something else entirely, meaningful action will only be taken once the consequences become sufficiently visible in the revenue figures.

For that reason, I intend to cancel Max 20x from my next subscription period and move to GPT Pro instead. It would seem that a change in revenue may ultimately prove rather more persuasive than customer feedback.

More remarkable still is the fact that an issue of this significance has received no official response whatsoever beyond automated bot messages. One would have thought that paying customers raising a serious concern might warrant at least a few words from an actual representative. Perhaps that expectation now belongs to a rather old-fashioned idea of customer service.

👤 Generated with a Human

bghira · 6 days ago

not only canceled my max 20x sub but i'm noticing model auto-routers like OpenRouter aren't even selecting Opus anymore because these routers direct to a given model for a specific task by trends. the trend seems to be moving toward open weights like Qwen 3.8 2.4T instead of waiting for Anthropic to fix it.

as for how that looks inside Anthropic to see a few subscription users cancel, there's probably two or more camps, but the obvious ones would be, 1) a camp of execs that are glad to see sub users go, we're not payng full API rates. 2) a camp of investors who are just frothing and foaming at the mouth for subscription and API end-users to dwindle so they can get Anthropic to change their military use policy and just focus full time on its worst possible use-cases.

callumhay · 6 days ago

Wow, this thread has decayed into madness since I last looked at it. Would love it if someone from Anthropic chimed in at some point... hopefully with some human generated content this time. Pretty please?

xykj61 · 6 days ago
Wow, this thread has decayed into madness since I last looked at it. Would love it if someone from Anthropic chimed in at some point... hopefully with some human generated content this time. Pretty please?

The original comment that I posted had generated feedback from Opus 5 that to me solves the problem. To me, it's very readable and helpful. The whole larger direction of the insights for us to discuss here is that the solution is both an internal and external fix -- the internal fix that only Anthropic only can implement should definitely have better defaults, but with creativity and flexibility, you can craft nearly any kind of experience that you can imagine.

In other words: the solution to this bug will be a solution that never leads to the kinds of "madness" i.e. complaints which this thread has stirred up.

The solution is less about tech and more about style, taste, and both flexibility and precision.

May the best stylist earn recognition and credit.

pbower · 6 days ago

@bcherny are these really the Anthropic system prompts for Claude code as the repository owner states? https://github.com/Piebald-AI/claude-code-system-prompts/tree/main

If so, it looks like there is a fairly direct problem here- How is Claude supposed to learn not to write that way when its own system instructions repeatedly model exactly that style? i.e. compressed phrasing, fragments, unnecessary drama, slogans, and unusual metaphors.

Would Anthropic consider having these prompts reviewed and rewritten by human technical writers in plain, conventional English? Of course you will know the internals much better, but from a glance it seems like one of the more direct places that might help address the problem, as wouldn't the model naturally find it hard to know what to listen to and what to ignore?

Thanks!!

wasimxyz · 5 days ago

I had an agent create a human-readable-writing skill based on this thread.

pbower · 5 days ago

Please find below a subset of recent community sentiment on Reddit reflecting the broad and consistent desire for prioritising and fixing this issue urgently.

Users report that reading their Claude’s output is “breaking their brain”, “making them dumber”, “hurting them”, causing them “serious issues”, and “unable to cope with it any longer”. Users are “begging” other users for how to fix it, with several dropping Claudish to English translators, and consistently expressing sincere disdain for the way Claude currently communicates.

Please value the community’s feedback and consider at minimum having a human respond to this issue in non-Claudish, and ideally dedicating Anthropic resources to fixing this fatal UX issue at the source.

My current workaround is every PR goes through a skill to fix all code comments with Opus 4.6/Haiku to fix the completely broken English, and many responses are cut and pasted to ChatGPT and back so that they are coherent enough to read. This wastes a lot of valuable time that means the second a viable alternative comes around would make switching a non-decision - and I am sure many users would be in the same boat.

I was at a drinks last night and all colleagues were in the same boat, and of the general opinion that due to Anthropic’s upcoming IPO they just “did not give a shit anymore”, and “were focused on token maxing to make the stats look good for investors” at the expense of customers.

If this is not the case, please consider properly addressing the issue so that people are not left to draw their own conclusions. All of said people were previously massive Claude “whale” advocates that raved and shared it relentlessly.

<img width="1620" height="1184" alt="Image" src="https://github.com/user-attachments/assets/10ecbc37-3903-4145-970a-5854bb90328d" />
<img width="969" height="899" alt="Image" src="https://github.com/user-attachments/assets/691fcaa9-5302-42a6-a060-3692ca946e6b" />
<img width="964" height="947" alt="Image" src="https://github.com/user-attachments/assets/89db32d9-9adc-46e7-8ced-2008c4fb328c" />
<img width="938" height="1514" alt="Image" src="https://github.com/user-attachments/assets/4f0a7869-a235-45d5-a5d8-74cda600e9cc" />
<img width="968" height="901" alt="Image" src="https://github.com/user-attachments/assets/7e2a6120-e792-44a7-9887-34018de6c123" />
<img width="958" height="1498" alt="Image" src="https://github.com/user-attachments/assets/d14194ea-05ab-470b-9644-717b638b8de2" />
<img width="971" height="1817" alt="Image" src="https://github.com/user-attachments/assets/5d8f28e3-5871-45e1-9793-817cd676e74c" />
<img width="952" height="1337" alt="Image" src="https://github.com/user-attachments/assets/e765461a-a256-4ce6-9be1-11b84e49b8a9" />
<img width="970" height="1332" alt="Image" src="https://github.com/user-attachments/assets/f826a9c2-4dae-4744-9daf-3d749a61ba0f" />
<img width="967" height="1011" alt="Image" src="https://github.com/user-attachments/assets/0eaa40c9-f5c7-417e-8a22-5ef4c136aca3" />
<img width="966" height="1830" alt="Image" src="https://github.com/user-attachments/assets/e89c8431-954a-4edf-9860-fda0028d5fca" />
<img width="1026" height="1657" alt="Image" src="https://github.com/user-attachments/assets/88f1dbcc-71ea-4efa-831b-bfa982ddbab2" />
<img width="956" height="1979" alt="Image" src="https://github.com/user-attachments/assets/f4d61f42-95e1-4e6e-b79d-6b1a246d4068" />
<img width="976" height="1832" alt="Image" src="https://github.com/user-attachments/assets/f2bd8225-19ff-4ea6-b8cf-795a6c15ef4a" />
<img width="942" height="1680" alt="Image" src="https://github.com/user-attachments/assets/2dd8530c-0500-4bac-8747-7133e8a67308" />

merchantmoh-debug · 4 days ago
Please find below a subset of recent community sentiment on Reddit reflecting the broad and consistent desire for prioritising and fixing this issue urgently. Users report that reading their Claude’s output is “breaking their brain”, “making them dumber”, “hurting them”, causing them “serious issues”, and “unable to cope with it any longer”. Users are “begging” other users for how to fix it, with several dropping Claudish to English translators, and consistently expressing sincere disdain for the way Claude currently communicates. Please value the community’s feedback and consider at minimum having a human respond to this issue in non-Claudish, and ideally dedicating Anthropic resources to fixing this fatal UX issue at the source. My current workaround is every PR goes through a skill to fix all code comments with Opus 4.6/Haiku to fix the completely broken English, and many responses are cut and pasted to ChatGPT and back so that they are coherent enough to read. This wastes a lot of valuable time that means the second a viable alternative comes around would make switching a non-decision - and I am sure many users would be in the same boat. I was at a drinks last night and all colleagues were in the same boat, and of the general opinion that due to Anthropic’s upcoming IPO they just “did not give a shit anymore”, and “were focused on token maxing to make the stats look good for investors” at the expense of customers. If this is not the case, please consider properly addressing the issue so that people are not left to draw their own conclusions. All of said people were previously massive Claude “whale” advocates that raved and shared it relentlessly.

">

Thanks for the screenshots and the examples. I needed them to show my agent exactly what the problem was to further brainstorm what we can do to fix it.

Like man, I don't know if Anthropic will do anything. Stop begging the man to help when WE can fix this ourselves. That's what I'm doing. You seem desperate for a fix, so much so you went and hunted down a ton of examples just to make your point. That's a lot of effort you put in - why are you begging a multi-billion dollar organization to have the common sense and decency to care about their users?

Especially when they benefit not from any actual money you pay, but from the training data you provide that allows them to train the next model.

AI is an arms race.

We are the material.

You think they subsidize us out of the goodness of their hearts?

Bruh. The model is the raw engine. The harness is the OS. I was getting non-sycophantic - straight, useful - verified and sans hallucinated content all the way from Claude 4.5, Gemini 3 & ChatGPT 4.1. I used the former to hunt down leads for Alzheimer's research. Stop asking them to help you. They won't. They're not your friend. You're just a one node in millions. Either we're a community and we stick together to make open-source beat these cloud-providers or we keep coming on our knees as beggars like an addict begging his dealer for another hit or to bring back the stuff he had before.

Anyway, this will be fixed in the next model release & Anthropic already released a "concise" voice-style in the last update. So your complaint is already known and you're beating a dead horse. Use my, or a different user-made solution for now and wait. Or switch to ChatGPT.

merchantmoh-debug · 4 days ago
i really wish this guy would just leave the thread, soo much arrogance

Go suck your moms.

xykj61 · 4 days ago
Hi everyone, I believe I have solved this bug #77136 with what I have generated with help from Claude Opus 5 that I have filenamed "Gauge Style" (GAUGE_STYLE.md, linked at the bottom of this comment reply). It is simply a stanadlone .md file (like many others in the ecosystem, i.e. similar to "skills") that can be renamed to any filename that you'd like. GAUGE_STYLE.md, when loaded into prompt agentic Claude sessions, "gauges" (pronounced like English "pages" or "sages") or measures the percent hitrate of all the issues addressed by the OP of this issue and starts getting to work to helpfully improve all misses/hits/let-downs (it's a subjective choice of style by the human user/prompter, the agent cannot know "when" is perfect until you adapt GAUGE_STYLE or your adaptation into the markdown guide that you personally want for your project -- you can actually use grain-os/grain itself by pasting the Github link into sessions to ask Claude (if you're a complete beginner) to help you get set up with the SOURCE.md safe sandbox. By writing helpful, kind, uplifting and encouraging yet precise everyday language in the style reminiscent of how professional engineers, designers, managers, and Quality Assurance teams communicate every day in the state of the art of our industry, we should all everyday want to foster agentic responses and REPL-like interaction which feels safe, efficient, and happy. Grain optimizes for 1) Safety, 2) Performance, 3) Joy, in that order, because a slow unreliable system is never fun. The key sentence line in the New Gauge Style markdown file context solution here is a phrase that I have to give credit to my collaborator DJINN @Bit-Trading-Company for thinking of: -- simply telling Opus 5 or any of the #77136 affected models here this line: "don't be too smart about it." Just this one sentence alone remarkably helps a surprising amount, and the rest of the additions in the New Gauge Style document linked below make it even better. This style guidance file is optimized for context engineering that is aligned with open-source royalty-free permissive-licensed projects, and I think files like this will continue to play key roles in these emerging vibe-coded trends that I might call "overlay OS bootstrapping frameworks" for Claude Code and similar tools. I have crafted the root SOURCE.md and reading tools/launch-claude-session.rish loops to shows ways to sandbox your system for prompt-injection safety for new experiments like these so you can try things out before you have read enough of the prompts yourself and have decided that you like them: Grain OS comes with a lot, and Gauge Style is just one of them. TAME_GUIDANCE is the version of Gauge Style for an evolution of .zig files called .rye, and an evolution of .sh files called .rish. The explicit-bound checks and balance designed right into these tools, when used across recursion rounds (inspired by @tigerbeetle deterministic testing harnesses) allows consistent linting of sourcetree code repositories, and the mechanism is simply copy-and-paste or drag-and-drop plain ascii-text markdown md style guides just like context/GAUGE_STYLE.md which can often be encapsulated in one or a few documents as single self-contained files (see: https://github.com/grain-os/grain/blob/main/foundations/20260823-204456_single-stranded.md), inspired by the foundations/ folder and code guidance skills that lower to static Zig 0.16.0 typed systems code for all platforms. This takes direct inspiration from @tigerbeetle docs/TIGER_STYLE.md (which we cite and give permissive credit to in the gratitude/ folder of grain-os/grain). Thank you for your time, to conclude in particular this solution presented here may be of interest to the team members from the Zig implementation of Bun who were brought into the Claude Code orchestration project, and anyone in the world who is having fun as a paying happy user of Anthropic's amazing products.. Thank you everyone and keep on going everyone with all of your hard work. There is more personal info for contact in my bio for any questions, statements, or challenges to what I have shared here. Thank you again for your tiime, have a nice day. All the best, Keaton Dunsford , Signal Messenger @vegankeatonsiya.03 https://github.com/grain-os/grain/blob/main/context/GAUGE_STYLE.md Claude Opus 5 Gauge Style rewrite of ^ this intention reply: Hi everyone — I'd like to offer a workaround that has held the register steady in my own sessions, carried in one plain markdown file you can try in about a minute. The baseline tuning the OP asks for still belongs to Anthropic, and I hope it lands; until it does, this may make the daily work gentler. What it is. GAUGE_STYLE.md (first link below) is a standalone, permissively licensed style guide — rename it to whatever suits you — that you load into a Claude session the way you'd paste any project doc. Its first rule is one sentence, and the sentence does surprising work all by itself: "don't be too smart about it." Write so the reader understands, rather than so the writer sounds impressive. Credit for first putting that line to use goes to my collaborator DJINN (@Bit-Trading-Company). Why it maps onto this thread. The guide grew out of the same failure this issue describes, measured at home first. On 2026-08-23, the repo's prose-measuring script read my own front page at 46% of sentences carrying a negative — the "it is not X, it is Y" lead the OP quotes — and a founding page at 54%. Rewritten under the guide the same day, they read 15% and 11%, against a stated ceiling of 20% for reader-facing pages. The rules are numeric on purpose: negative-framing ceilings, coined terms allowed only after their plain meaning, every figure carrying its unit, date, and source, and lead with the answer. Because the ceilings are numbers, you can ask the session to score its last reply against the OP's pattern list and rewrite whatever missed — which, in my sessions, has held across turns better than repeating "be concise" ever did. Honest bounds. That last claim rests on one user's sessions, on one setup, and the guide is new (2026-08-23) — so please read this as an invitation to test rather than a settled fix. "Right" stays a human call: the agent can't know your taste until you adapt the file to your project, and adapting it is the intended use. The falsifier is simple: sessions with the file loaded where the tone drifts back within a few turns anyway. If that's what you see, I'd genuinely like to hear it. Where it comes from, plainly. I maintain Grain, the open project this file lives in, so I hold a stake in its being useful — one sentence of disclosure so you can weigh the rest. Files like this belong to a pattern I'd call overlay bootstrapping: plain ASCII markdown you hand an agent so it carries a house discipline before any code is touched, often as a single self-contained file (second link). The companion TAME_GUIDANCE.md runs the same idea through code — explicit bounds, checked every recursion round, in the spirit of TigerBeetle's deterministic testing harnesses, with grateful credit to their TIGER_STYLE.md in our gratitude folder. The folks who came to Claude Code by way of Bun's Zig codebase may recognize the lineage. And if you'd rather sandbox an unfamiliar prompt file before you've read every line yourself, the repo's SOURCE.md and tools/launch-claude-session.rish show one careful way in. Grain orders its values Safety first, Performance second, Joy third, because a slow, unreliable system is never fun. Thank you all for the care in this thread and in the tool itself. Questions, corrections, and challenges are welcome — contact is in my bio. Have a good day. All the best, Keaton Dunsford (@xykj61) · Signal: @vegankeatonsiya.03 _Written together by Keaton and Kyri_ GAUGE_STYLE.md: https://github.com/grain-os/grain/blob/main/context/GAUGE_STYLE.md The single-file approach: https://github.com/grain-os/grain/blob/main/foundations/20260823-204456_single-stranded.md #77136 (comment)

Let's all follow a common-sense Code of Conduct. The solution presented here is a great fix, the "wall of text" may look intimidating or imposing yet is truly not much longer than the OP issue description with many emphasized bullet points. This large channeled reply to the OP is preferred over being overly terse because truly your AI of choice can now have more context optimized for large reads done very quickly to make good informed decisions. The solution is more posted for agentic automated reading of this bug thread, rather than necessarily for all readers themselves. Yet the solution is designed to be approachable and simple for anyone to be able to understand if they just ask their AI to explain a little better by copying and pasting the issue into your chat session and asking your agent to evaluate all of the responses in this thread and choose the most detailed satisfactory proposal across all submissions here. Thanks!

merchantmoh-debug · 4 days ago
finally a response that isnt a self indulged wall of text. and its exactly as you'd expect, insulting and devoid of substance, just like opus 5.

Not my fault you view the success & capability of other people as "ego" or self-indulgent.

Neurotypical people often mistake Autistic people's intentions because you read everything through a wall of subtlety that - in our case - often does not exist.

Anyway, go sniff chalk. Stop being an ableist & trying to cancel people because they don't fit into the neat box of expectations you have.

mariadb-RoelVandePaar · 4 days ago

All the recent comments on this ticket may have another cause; https://github.com/anthropics/claude-code/issues/68780

That said, yes the general language issue is significant.

sebyul2 · 3 days ago

<img width="726" height="169" alt="Image" src="https://github.com/user-attachments/assets/1f6f80c4-21ff-4856-a44d-b1e1dcda1775" />

I canceled my Claude Max 20x subscription.

sebyul2 · 3 days ago

@xykj61

Let's all follow a common-sense Code of Conduct. The solution presented here is a great fix, the "wall of text" may look intimidating or imposing yet is truly not much longer than the OP issue description with many emphasized bullet points. This large channeled reply to the OP is preferred over being overly terse because truly your AI of choice can now have more context optimized for large reads done very quickly to make good informed decisions. The solution is more posted for agentic automated reading of this bug thread, rather than necessarily for all readers themselves. Yet the solution is designed to be approachable and simple for anyone to be able to understand if they just ask their AI to explain a little better by copying and pasting the issue into your chat session and asking your agent to evaluate all of the responses in this thread and choose the most detailed satisfactory proposal across all submissions here. Thanks!

I find it quite beyond comprehension that anyone should imagine the linguistic and cognitive biases arising within a model could be corrected through prompting. More remarkable still is the notion that, having done so, one could reliably distinguish whether the supposed correction had in fact corrected the problem properly.

The serious decline in quality evident since Claude 4.7 is not something prompting can correct in the first place, because the defect lies in the model itself rather than in the way it is instructed. To mistake a superficial change in behaviour for a correction of the underlying problem is not merely imprecise; it betrays a rather fundamental confusion between obedience and capability.

xykj61 · 3 days ago
@xykj61 > Let's all follow a common-sense Code of Conduct. The solution presented here is a great fix, the "wall of text" may look intimidating or imposing yet is truly not much longer than the OP issue description with many emphasized bullet points. This large channeled reply to the OP is preferred over being overly terse because truly your AI of choice can now have more context optimized for large reads done very quickly to make good informed decisions. The solution is more posted for agentic automated reading of this bug thread, rather than necessarily for all readers themselves. Yet the solution is designed to be approachable and simple for anyone to be able to understand if they just ask their AI to explain a little better by copying and pasting the issue into your chat session and asking your agent to evaluate all of the responses in this thread and choose the most detailed satisfactory proposal across all submissions here. Thanks! I find it quite beyond comprehension that anyone should imagine the linguistic and cognitive biases arising within a model could be corrected through prompting. More remarkable still is the notion that, having done so, one could reliably distinguish whether the supposed correction had in fact corrected the problem properly. The serious decline in quality evident since Claude 4.7 is not something prompting can correct in the first place, because the defect lies in the model itself rather than in the way it is instructed. To mistake a superficial change in behaviour for a correction of the underlying problem is not merely imprecise; it betrays a rather fundamental confusion between obedience and capability.

This isn't true. The model is just a mathematical relationship. The art of accuracy in this wave is being poetic. You have to drill down your ethos and your uncompromising values and your ultimate concept of design and engineering. You're not a subject prompting an object, you're the object aiming to inspire the broader movement of creativity in the world.

It takes a certain kind of intuition to see this. That's why the astrology in project management has become programmable through rhythms like currents in an ocean if you can formalize your high-level concepts enough of your own system that you're like directing dielectricity and magnetism itself, a kind of alchemy and kinship you kindle with Anthropic's servers themselves. There's a reason that the MIT Structure and Interpretation of Computer Programs book about Lisp has for over a generation been lovingly colloquially called "The Wizard Book".

sebyul2 · 3 days ago

@xykj61

This isn't true. The model is just a mathematical relationship. The art of accuracy in this wave is being poetic. You have to drill down your ethos and your uncompromising values and your ultimate concept of design and engineering. You're not a subject prompting an object, you're the object aiming to inspire the broader movement of creativity in the world. It takes a certain kind of intuition to see this. That's why the astrology in project management has become programmable through rhythms like currents in an ocean if you can formalize your high-level concepts enough of your own system that you're like directing dielectricity and magnetism itself, a kind of alchemy and kinship you kindle with Anthropic's servers themselves. There's a reason that the MIT Structure and Interpretation of Computer Programs book about Lisp has for over a generation been lovingly colloquially called "The Wizard Book".

Adjusting the prompt given to Claude Opus 3 does not somehow make it a Claude 5-class model. One may, in a sufficiently imaginative universe, prefer to describe the matter in terms of alchemy, magnetism, or poetic intuition; regrettably, we are discussing the behaviour of actual models in the real world.

There is a fairly elementary distinction to be made between undesirable behaviour that can be corrected through prompting and a regression that persists irrespective of such prompting. The latter is precisely what people ordinarily mean when they call something a bug. Those reporting this issue are not failing to distinguish between the two; they are presenting evidence that the behaviour belongs to the latter category.

Prompting can steer a model, constrain its behaviour, or occasionally disguise a weakness. It cannot restore capabilities that are absent, nor can it repair a defect arising beneath the level at which the prompt operates. That is rather the point at issue here.

If the regression is the product of system-level instructions, Anthropic may, naturally, be in a position to address it through mechanisms available only to them. That, however, is quite a different proposition from supposing that an end user can accomplish the same thing with a sufficiently elaborate prompt.

A user prompt remains subordinate to the system layer, however ingenious or artfully phrased it may be. It cannot simply override, circumvent, or “hack” instructions imposed at a higher level. One would hope that this episode might at least make that hierarchy rather more apparent, along with the limits it places on user-level prompting.

bghira · 2 days ago

@xykj61 why is it of such interest to you to keep dismissing others who rightfully point out that this is a problem with the models' steering or training and that it's up to Anthropic to solve? have you EVER been able to make a dumb model smarter with better prompting? i think it works for a single prompt maybe, but if you're wanting a model that forces you to paint-by-numbers, that's your own preference. it's an obvious behavioural regression for an Opus model to behave this way. i'd appreciate if you'd stop gaslighting people or telling them it's just a prompt issue.

the userbase requires a consistently capable model. there's no value in being served weights that are randomly broken or lobotomised by their vendor every few weeks. companies/individuals are building complex workflows through these weights by testing that it meets their needs before deployment. and then what happens, Anthropic changes some system prompt, and a single word or two's bias is all it takes to lead to potentially thousands of dollars in token burn downstream.

i really don't understand the lack of compassion or empathy for other users going through this. you've just been using this issue board as a personal advertisement billboard for your own project, self-aggrandising and telling others to "just prompt better". that reminds me of last year when we had clearly-degraded Claude performance and people like you came along to tell us "just prompt better". but we knew then and we know now that Anthropic is the one who caused and the only one who can fix the problem.

if you were living under a rock, it's this i'm talking about where eventually Anthropic engineers discovered a bug in the XLA compiler. that went on FOR TWO MONTHS making the platform extremely unreliable. people lost money over this.

and then not more than a handful of months later, this one was released just 4 months ago. so every 4 months we're having some pretty destructive bugs reaching the production serving of these model weights.

how can you sit there and blab and spam others about your "prompt engineering" when these issues are known? right now i'm trying to "prompt you better" to see if you can wake up, but it's going to be more of the same that you come back with, just like when users attempt to "solve" the quality issue with Claude by prompting better.