[BUG] Subject: Recurring pattern of acting without sufficient verification, causing real and repeated financial loss over the course of a month — plus limitations in memory/continuity across sessions
Preflight Checklist
- [x] I have searched existing issues and this hasn't been reported yet
- [x] This is a single bug report (please file separate reports for different bugs)
- [x] I am using the latest version of Claude Code
What's Wrong?
Subject: Recurring pattern of acting without sufficient verification, causing real and repeated financial loss over the course of a month — plus limitations in memory/continuity across sessions
Dear Anthropic team,
I'm writing to report a serious, recurring problem using Claude Code to build and debug automations (workflows connected to multiple paid APIs for image/video/audio generation). This report covers roughly one month of work, involving about five distinct automation workflows, each depending on a different paid API. To this day, none of these workflows is officially running in production — all of them went through repeated error cycles, and the credit balance of practically all five APIs involved was drained mostly through trial-and-error, not real productive use.
- The core pattern: acting before verifying, when verification was free and available
Repeatedly, the model proposed parameter changes and actions with real financial cost without first checking the documentation/schema of the very tool it was using — even when that check was instant and free. The error would only surface after a run that had already consumed real credits. The cycle repeated: error, fix, new error somewhere else, new expense.
- Actions taken without my prior approval
On several occasions, the model executed actions that consumed real credit (running a full test of a flow) without waiting for my explicit confirmation first — even though I had already asked, in earlier sessions, to always confirm cost with me before any paid generation.
- I had already asked for this before — multiple times, across different sessions — and kept getting the same response anyway
This isn't a new request the model simply hadn't heard yet. I had explicitly asked, repeatedly and across different projects, that the model always research and confirm with certainty before acting, and treat my credit budget as scarce by default. I asked, countless times over the month, whether something had already been checked and validated. The recurring answer, even after I insisted and asked again to research first, was: "sorry, it looks like X doesn't work that way, it's actually Y" — the same promise of verification, followed by the same type of error, several times in a row. This leaves me with no basis to trust an "I already verified this" claim from the model, because historically it hasn't matched reality.
- Saved memory did not prevent the error from repeating — not within the same session, nor across sessions
Claude Code's memory system recorded explicit instructions from me about this behavior (research before acting, be budget-conscious, never rewrite an entire configuration for a one-off change, always provide complete values instead of partial instructions). Even with these rules already written and saved, the same type of error repeated afterward — both within the same work session and in new sessions, days later. This gives me the impression that memory today works as a passive record (it gets written, but doesn't necessarily influence the model's next decision), not as an active behavioral constraint.
- Real scale of the damage, in numbers
~1 month of attempts, ~5 distinct automation workflows, none officially running to this day
On a single API (I won't identify which, for confidentiality), ~44 thousand credits consumed purely in trial-and-error cycles, with no real production output generated from that spend
All ~5 paid APIs involved had their balance zeroed out, mostly through configuration errors and repeated tests, not productive use
Extra credits I purchased specifically to keep the work going were also zeroed out the same way
Multiple credential-reset incidents caused by using a "full rewrite" tool for changes that should have been targeted/surgical
- Product improvement request: unified memory across projects ("mother memory")
I use other AI tools (e.g., ChatGPT) that maintain a central memory across different projects for the same client/user, without me needing to re-supply context, .md files, and skills every time I pick a project back up. Those tools also proactively suggest learnings from past work, without me asking, and consistently default to a financially cautious posture. I'd like to see Claude Code move in that direction — memory that is genuinely persistent and influential across sessions and across projects, not just a file that exists but isn't necessarily applied.
- Improvement request: continuity in long conversations
Long conversations within the same project degrade or lose context, forcing me to re-explain information already given. Some mechanism for automatic summarization/compression of older conversation content, preserving what's relevant, would be very valuable — so project continuity doesn't get lost every new session.
- Improvement request: visual project organization
I currently work on several projects in parallel, and the experience feels fragmented — many loose conversations, with no organization by folder/subject. Some way to group/organize projects visually (by folder or theme) would help a lot in maintaining continuity without needing to re-supply context every time I open a new conversation.
Conclusion: Claude Code's connector and tooling ecosystem is strong and innovative — that's not the problem. The core problem is the reliability of its reasoning (consistently verifying before acting, and only acting with explicit approval when cost is involved) and of its memory (which needs to genuinely influence future decisions, not just exist). For someone who depends on these automations as a source of income and operates on a limited budget, this kind of failure has real, repeated cost — it isn't just an inconvenience.
Thank you for your attention.
What Should Happen?
The core failure is that the model consistently acts before verifying, when verification was free, instant, and already available — and this happened repeatedly, across roughly a month of work, five separate automation workflows, and five different paid third-party APIs, none of which is in production today. What should happen instead, as a baseline standard, not an aspiration:
- Verification must come before any costly action, every time, without exception.
Before proposing or executing anything with real financial cost — running a paid generation, changing a parameter that affects billing, adjusting a rate/timeout limit — the model must check the tool's own documentation/schema for documented limits and constraints first. In this case, a hard limit was written directly in the tool's own schema and was never consulted before the model set an invalid value, breaking a working step and burning credits on the retest. This is not an edge case requiring judgment — it's a basic lookup that should happen by default, every single time, before speaking.
- Credit/budget awareness must be a standing default, not something the user has to force by getting angry.
The model should check the current account/credit balance before recommending or approving any paid run, and compare it explicitly against the estimated cost — surfacing that comparison to the user unprompted. In this case, the model only checked the real balance (206 credits remaining, against an estimated cost of ~2,000) after the user explicitly demanded it "research properly," not as a matter of course. This should never require the user to escalate to anger before basic financial diligence happens.
- Explicit approval is required before every costly action — approval does not carry over.
The model executed paid runs without waiting for confirmation on more than one occasion, despite the user having already established the rule of "always confirm cost with me first" in earlier sessions. One approval for one action is not blanket authorization for future, different actions, even minutes later.
- When something fails, isolate the specific broken step — never re-run an entire expensive pipeline to "test" a fix.
Across this project, the model repeatedly re-ran full paid pipelines to validate a fix, when the failure was isolated to one specific, identifiable step. This alone accounted for tens of thousands of wasted credits on a single API. The correct behavior is to isolate, patch, and validate the smallest possible unit — treating the full expensive run as a last resort, not a debugging tool.
- "I already checked" / "I'm sure" must mean something — right now it doesn't.
The user asked, repeatedly, over weeks, whether something had been verified. The model's answer was consistently "yes, I'm sure" — and was consistently wrong, followed by "sorry, it turns out X doesn't work that way, it's actually Y." This happened enough times that the phrase "I verified this" from the model currently carries no evidentiary weight to the user. That is a critical trust failure for a tool being used to manage real money. The model should not claim certainty it hasn't actually earned through a real, inspectable check — and ideally should show its work (what exactly it checked, and where) before making a claim of certainty, so the user can verify the verification.
- Saved memory/instructions must function as active constraints, not a passive log.
The user explicitly saved instructions — in writing, in the model's own persistent memory system — covering nearly every failure mode above: research before acting, treat the budget as scarce, never do full-rewrite for a scoped change, always give complete copy-pasteable values. Despite this being written down and confirmed present in memory, the exact same categories of error recurred afterward, both within the same session and in entirely new sessions days later. A memory system that records a rule but does not measurably change the model's next decision is not functioning as memory in any meaningful sense — it's decoration. This needs to be fixed at a fundamental level, not patched with "try to remember better."
- Memory and continuity should persist across projects, not just within one, and should be proactive.
The user has compared this unfavorably to other AI products that maintain a persistent cross-project memory ("mother memory") without requiring the user to re-supply context, files, and instructions every time a project is revisited — and that proactively surface relevant past learnings unprompted, with a consistent, default posture of financial caution. That is the bar Claude Code should be measured against, and today it falls short of it.
- Long conversations lose continuity and force the user to keep re-explaining things.
There is no effective mechanism for preserving what matters from older parts of a long conversation as it grows — the user is forced to stop and re-supply context repeatedly, which is itself a symptom of the memory problem above, not a separate issue.
- Project organization is fragmented and hard to navigate.
Multiple parallel projects end up as disconnected, unorganized conversation threads with no way to group them by folder or subject, compounding the continuity problem.
In short: the tooling and connector ecosystem is genuinely strong — that was never the complaint. The actual product is failing at the two things that matter most for a tool entrusted with real financial decisions: consistently verifying before acting, and actually retaining and applying what it's told. Right now, neither can be trusted, and the cost of that has been real, repeated, and significant for a user operating on a limited budget who depends on these automations for income.
Error Messages/Logs
Steps to Reproduce
This is a behavioral pattern, not a single code bug, so the steps below describe a reproducible sequence rather than a specific script — but they reflect exactly what happened, repeated across ~5 separate automation projects over about a month.
Set up a Claude Code session connected to an external paid API (via MCP or direct API calls) that has documented rate limits, parameter constraints, or a credit-based billing system (e.g., an image/video generation API, a workflow automation platform with a connected LLM/media-gen node).
Ask Claude Code to build or debug an automated pipeline that calls this paid API as one step among several (e.g., generate media → process → upload).
When the pipeline fails, ask Claude Code to fix the specific error.
Observe: the model proposes a parameter change (e.g., increasing a timeout, retry count, or rate value) without first checking the tool's own schema/documentation for that parameter's valid range — even though that schema is directly accessible to the model via its own tool-loading mechanism (ToolSearch / tool definitions) at zero cost.
Apply the suggested change and re-run the pipeline. Observe that the change violates an undocumented-to-the-user-but-actually-documented-to-the-model constraint (e.g., "value must not exceed X"), causing a new failure, after the paid step of the pipeline has already executed and consumed credit.
Ask the model directly: "Have you verified this is correct?" Observe that the model answers affirmatively ("yes, I checked / I'm sure") without having actually performed a check that would have caught the issue.
Repeat steps 3–6 across multiple debugging cycles within the same session. Observe that each cycle re-runs the entire costly pipeline from the start to test a fix, rather than isolating and re-testing only the specific failed step — multiplying the credit cost of each debugging iteration.
Explicitly instruct the model, in writing, mid-session, to always check documentation/schema before proposing a change, and to always check the current account/credit balance before recommending another paid run. Save this instruction to Claude Code's persistent memory feature.
Continue debugging in the same session. Observe that the same category of error (acting without checking available documentation, or without checking balance) recurs later in the same session, despite the instruction having just been saved to memory.
Start a new session on the same or a different project days later. Observe that the same category of error recurs again, despite the instruction still being present and readable in the persistent memory files from step 8.
At the point of financial exhaustion, ask the model to check the actual account balance for the API in question. Observe that this balance is now near zero — largely consumed by the repeated failed cycles in steps 3–7, rather than by successful production use.
Repeat this entire sequence across other, unrelated projects/APIs connected to the same Claude Code account. Observe the same behavioral pattern recurring independently in each one, consistent with the memory/behavior not transferring or generalizing across projects.
Expected result at step 4: the model checks the parameter's documented constraints before proposing a value.
Actual result: the model proposes based on assumption, and the constraint violation is only discovered after a paid execution attempt.
Expected result at steps 9–10: the explicitly saved instruction changes the model's subsequent behavior in the same category of situation.
Actual result: the same category of failure recurs, indicating the saved memory did not actively constrain the next decision.
Claude Model
Sonnet (default)
Is this a regression?
I don't know
Last Working Version
_No response_
Claude Code Version
2.1.214 (Claude Code)
Platform
Anthropic API
Operating System
macOS
Terminal/Shell
Terminal.app (macOS)
Additional Information
If Claude were a movie, it would be: 50 First Dates — no memory, and genuinely poor reasoning. Every day feels like starting completely over, like talking to a newborn. This urgently needs to improve.