[BUG] Inconsistent and Undisclosed Quota Accounting Changes in Claude Max Plan. Legal liability claim ready.
Preflight Checklist
- [x] I have searched existing issues and this hasn't been reported yet
- [x] This is a single bug report (please file separate reports for different bugs)
- [x] I am using the latest version of Claude Code
What's Wrong?
Bug Report: Inconsistent and Undisclosed Quota Accounting Changes in Claude Max Plan
Filed: 2026-02-01
Plan: Claude Max 20x ($200/month)
Affected Period: January 30 - February 1, 2026
Severity: Critical - Service degradation without disclosure
---
Executive Summary
Instrumented monitoring of Anthropic's rate limit headers reveals inconsistent quota consumption rates that cannot be explained by the "holiday bonus expiration" cited by Anthropic. The same user, same plan, same workload type shows burn rates varying from 5.6%/hour to 59.9%/hour within the same 48-hour period. This 10x variance constitutes either a bug in quota accounting or an undisclosed server-side change to rate limiting behavior.
---
Evidence
Methodology
All data was collected by intercepting HTTP responses from api.anthropic.com/v1/messages via a local mitmproxy instance. Rate limit values are extracted directly from Anthropic's own response headers:
x-ratelimit-5h-utilization: 0.81
x-ratelimit-7d-utilization: 0.23
x-ratelimit-5h-status: allowed
Data is stored in SQLite with 5,396+ samples spanning January 30 - February 1, 2026. No sampling bias - every API response is recorded.
Session Analysis (Quota Reset to Reset)
| # | Start | End | Start% | End% | Duration | Rate (%/hr) | Status |
|---|-------|-----|--------|------|----------|-------------|--------|
| 1 | Jan 30 18:41 | Jan 31 02:37 | 9% | 83% | 7.9h | 9.3%/hr | Normal |
| 2 | Jan 31 10:02 | Jan 31 20:24 | 0% | 100% | 10.4h | 9.6%/hr | Normal |
| 3 | Jan 31 20:28 | Jan 31 22:01 | 0% | 86% | 1.5h | 56.0%/hr | 2.8x fast |
| 4 | Jan 31 22:18 | Feb 01 15:00 | 0% | 94% | 16.7h | 5.6%/hr | Normal |
| 5 | Feb 01 15:00 | Feb 01 16:40 | 0% | 100% | 1.7h | 59.9%/hr | 3.0x fast |
| 6 | Feb 01 16:43 | Feb 01 18:30 | 0% | 100% | 1.8h | 56.1%/hr | 2.8x fast |
| 7 | Feb 01 20:01 | ongoing | 0% | 36% | 0.9h | 40.0%/hr | 2.0x fast |
Key Observation
Sessions 1, 2, and 4 all occur AFTER the holiday bonus expiration (Dec 31) and show normal ~10%/hr rates. If the holiday bonus expiration were the sole explanation, ALL sessions should show the same rate. Instead:
- Normal sessions: 5.6 - 9.6%/hr (consistent with advertised 5-hour window)
- Anomalous sessions: 40.0 - 59.9%/hr (3-6x the expected rate)
This variance occurs on the same account, same plan, same day, ruling out:
- Holiday bonus as the explanation (normal sessions exist post-holiday)
- User behavior differences (same operator, same workload type)
- Plan differences (same Max 20x account throughout)
Token-to-Quota Correlation
From 722 quota-increasing request pairs:
| Metric | Value |
|--------|-------|
| Median tokens per 1% quota | 2,517 |
| Mean tokens per 1% quota | 39,152 |
| Min (worst efficiency) | 12,300 tokens per 1% |
| Max (best efficiency) | 18,531,900 tokens per 1% |
The 1,500x spread between min and max tokens-per-percent is not explainable by cache behavior differences alone. This suggests the cost-per-token in quota terms is not deterministic.
---
Anthropic's Advertised Promises vs. Reality
What Anthropic Advertises
From claude.com/pricing and Claude Help Center:
"Max plan: 20x more usage than the Pro plan" "About 900+ messages per 5-hour window" "Rolling 5-hour windows rather than daily caps"
From Anthropic's official announcement:
"Introducing a new Max plan for Claude. It's flexible, with options for 5x or 20x more usage compared to our Pro plan."
What Users Actually Experience
During anomalous sessions (3, 5, 6, 7), the 5-hour quota window is exhausted in 1.5-1.8 hours. This means:
- Advertised: 5-hour window -> 900+ messages
- Actual: ~1.7-hour window -> significantly fewer messages
- Effective multiplier: Not 20x Pro, but approximately 6-7x Pro during anomalous periods
Users are paying $200/month for a 20x multiplier and intermittently receiving what amounts to a 6-7x multiplier with no disclosure, notification, or compensation.
---
Legal Analysis
1. Breach of Express Warranty (UCC 2-313)
Anthropic's marketing materials constitute express warranties under the Uniform Commercial Code:
- "20x more usage than Pro" is a specific, quantified promise
- "900+ messages per 5-hour window" is a specific performance guarantee
- "Rolling 5-hour windows" implies consistent, predictable behavior
When users intermittently receive 1.7-hour windows instead of 5-hour windows, the express warranty is breached.
2. California Unfair Competition Law (Bus. & Prof. Code 17200)
Anthropic is headquartered in San Francisco, California. California's UCL prohibits:
- Unfair business practices: Reducing service levels without notice or compensation while continuing to charge full price is "substantially injurious to consumers" (unfair prong)
- Fraudulent practices: Advertising "20x usage" and "5-hour windows" while delivering 1.7-hour windows during anomalous periods is conduct "likely to deceive" consumers (fraudulent prong)
- Unlawful practices: If the above also violates FTC Act Section 5, it's actionable under the UCL's unlawful prong as well
UCL imposes strict liability - Anthropic's intent is irrelevant. The fact that the practice is unfair or deceptive is sufficient.
3. FTC Act Section 5 - Deceptive Practices
The FTC prohibits deceptive acts or practices in commerce. The FTC's penalty offenses framework for bait-and-switch applies when:
- A service is advertised with specific capabilities ("20x", "5-hour windows")
- The advertised service is not consistently available as described
- Consumers pay based on the advertised description
While the FTC's Click-to-Cancel rule was vacated by the Eighth Circuit in July 2025, the FTC retains enforcement authority under Section 5 and ROSCA for deceptive subscription practices. Penalties can reach $53,088 per violation.
4. Anthropic's Terms of Service Defense - And Why It's Insufficient
Anthropic's Consumer Terms (anthropic.com/legal/consumer-terms) state:
"We may sometimes add or remove features, increase or decrease capacity limits..." "We reserve the right to modify, suspend, or discontinue the Services... at any time without notice to you."
This defense has significant limitations:
- Unconscionability: A take-it-or-leave-it clause that allows unlimited service degradation while maintaining full pricing may be unconscionable under California law, particularly for a $200/month consumer subscription.
- Marketing overrides ToS: When marketing materials make specific quantified promises ("20x", "900+ messages", "5-hour windows"), those promises create binding obligations that cannot be entirely disclaimed by buried ToS clauses. See Weinstat v. Dentsply International, 180 Cal.App.4th 1213 (2010).
- Implied covenant of good faith: Even if Anthropic can modify limits, they must do so in good faith. Silently reducing effective capacity by 3x while maintaining pricing and continuing to advertise "20x" and "5-hour windows" arguably violates the implied covenant.
- The inconsistency is the problem: If Anthropic uniformly reduced limits post-holiday, users could evaluate and decide whether to continue subscribing. The intermittent nature of the degradation (normal one session, 3x faster the next) prevents users from making informed purchasing decisions.
---
Corroborating Community Reports
This is not an isolated experience. Multiple GitHub issues and community reports document the same pattern:
| Issue | Title | Upvotes | Key Finding |
|-------|-------|---------|-------------|
| #16157 | Instantly hitting usage limits with Max subscription | 490+ | Never hit limits in 3 months, now exhausted in 2 hours |
| #17084 | Opus 4.5 limits significantly reduced since Jan 2026 | 237+ | "Most restrictive since Opus 4.5 launch" |
| #16868 | Credit consumption rate increased 3-5x after Jan 1 | - | "CRITICAL: Max 5x plan credits depleting 3-5x faster" |
| #20767 | Pro Quota reduced to ~20 prompts per 5 hours | - | "Significantly lower than pre-holiday standard levels" |
| #19673 | "You've hit your limit" at 84% usage | - | Rate limiting before reaching advertised capacity |
| #16270 | Usage limits bugged after double limits expired | - | Opened Jan 4, day 4 after holiday bonus ended |
Media coverage: The Register reported a community-estimated ~60% reduction in token usage limits, with developers' criticism allegedly being censored in Anthropic's Discord.
---
Anthropic's Response - And Why It's Inadequate
Anthropic's official position (via The Register, Jan 5, 2026):
"The concerns appear to be largely a response to the resumption of normal limits."
Our data directly contradicts this:
- Sessions 1, 2, and 4 show "normal limits" (~10%/hr) - these ARE post-holiday
- Sessions 3, 5, 6, and 7 show 3-6x consumption - these are ALSO post-holiday
- Both rate patterns occur on the same account, same day
If limits simply "returned to normal," they would return to a consistent normal rate. The 10x variance between sessions is either:
- A bug in quota accounting (which Anthropic should acknowledge and fix)
- An intentional server-side change (which should be disclosed per ToS transparency obligations)
---
Evidence Collection Tool: Quota Tracking Dashboard
To systematically collect and verify the evidence presented in this report, we developed an open-source Quota Tracking Dashboard as part of the claude-thinking-audit project. This tool is available for any Claude subscriber to independently reproduce our findings.
What It Does
The dashboard is a dedicated QUOTA tab added to a mitmproxy-based monitoring UI that passively intercepts Claude API responses and extracts rate limit data from Anthropic's own HTTP headers. It provides:
- Real-Time Status Cards — Current 5-hour and 7-day quota utilization, live burn rate calculation, and estimated time to 100% exhaustion with deviation warnings when burn rate exceeds expected levels
- Burn Rate Chart — ASCII sparkline visualization of quota progression over selectable time ranges (1h, 6h, 24h, 7d), clearly showing the sawtooth pattern of quota cycles and resets
- Session History Table — Automatic detection of quota reset events to identify individual usage sessions, with per-session burn rate calculation and color-coded status flags (green=normal, red=anomalously fast)
- Token-to-Quota Correlation — Statistical analysis of how many tokens correspond to 1% of quota, revealing the non-deterministic nature of Anthropic's quota accounting
- Systemd Log Viewer — Ingests rate limit entries from journalctl into SQLite for fast querying, showing a scrollable table of timestamped quota snapshots directly from the service logs
- One-Click Evidence Export — Generates a downloadable JSON report containing all of the above data in a structured format suitable for filing bug reports, including quota progression history, session analysis, correlation statistics, and raw log entries
Why This Matters
Before this tool, users had no way to verify their quota usage claims. Anthropic's own /context command in Claude Code has a known bug (showing 0% when usage is non-zero), and there is no user-facing quota dashboard. Users reporting issues on GitHub could only say "it feels faster" — now they can provide timestamped, header-sourced evidence.
The evidence in this report was generated using the Export Evidence feature of this tool. Every data point traces back to an x-ratelimit-* header in an actual API response stored in our SQLite database.
Sources
- The Register: Claude devs complain about surprise usage limits - https://www.theregister.com/2026/01/05/claude_devs_usage_limits/
- GitHub Issue #17084: Opus 4.5 limits reduced - https://github.com/anthropics/claude-code/issues/17084
- GitHub Issue #16157: Instantly hitting limits with Max - https://github.com/anthropics/claude-code/issues/16157
- GitHub Issue #20767: Pro quota reduced to ~20 prompts - https://github.com/anthropics/claude-code/issues/20767
- GitHub Issue #16868: Credit consumption 3-5x increase - https://github.com/anthropics/claude-code/issues/16868
- GitHub Issue #16270: Usage limits bugged after holiday - https://github.com/anthropics/claude-code/issues/16270
- Anthropic Consumer Terms of Service - https://www.anthropic.com/legal/consumer-terms
- What is the Max plan? Claude Help Center - https://support.claude.com/en/articles/11049741-what-is-the-max-plan
- Holiday 2025 Usage Promotion - https://support.claude.com/en/articles/13163666-holiday-2025-usage-promotion
- FTC Penalty Offenses: Bait and Switch - https://www.ftc.gov/enforcement/penalty-offenses/bait-switch
- California UCL 17200 - https://codes.findlaw.com/ca/business-and-professions-code/bpc-sect-17200/
- dev.ua: Claude Code users reported limit drop - https://dev.ua/en/news/korystuvachi-claude-code-zaiavyly-pro-padinnia-limitiv-anthropic-poslalysia-na-kinets-sviatkovoho-bonusu-1767694917
- Anthropic: Updates to Consumer Terms - https://www.anthropic.com/news/updates-to-our-consumer-terms
- Claude AI Plans 2026 Guide - https://www.glbgpt.com/hub/claude-ai-plans-2026/
- Byteiota: Anthropic blocks Claude Max in OpenCode - https://byteiota.com/anthropic-blocks-claude-max-in-opencode-devs-cancel-200-month-plans/
What Should Happen?
Requested Resolution
Immediate
- Acknowledge the inconsistency in quota accounting
- Publish the formula used to calculate quota consumption per request
- Fix the variance that causes 10x differences in effective burn rates
Structural
- Transparency: Display actual quota consumption per request in Claude Code's UI (Anthropic already sends x-ratelimit-* headers - surface them to users)
- Notification: Notify users BEFORE service level changes, not after complaints
- Consistency: Ensure quota accounting is deterministic - same tokens should always cost the same quota percentage
Compensatory
- Credit affected users for periods where quota burned at anomalous rates
- Update marketing materials to reflect actual achievable usage levels, or restore advertised levels
---
Error Messages/Logs
Steps to Reproduce
How to Reproduce
Any Claude subscriber can deploy this tool:
- Install mitmproxy with the custom rate-limit extraction addon
- Route Claude Code traffic through the local proxy
- Open the Quota Tracking Dashboard in a browser
- Use Claude Code normally — the tool passively records every API response
- Click "Export Evidence Report" to generate a structured JSON report
---
Appendix: Data Collection Methodology
Tools Used
- mitmproxy (v11.x) intercepting HTTPS traffic to api.anthropic.com
- Custom addon extracting rate limit headers from every API response
- SQLite database storing 5,396+ samples with timestamps, token counts, and quota utilization values
- claude-thinking-audit quota tracking dashboard for visualization and analysis (https://github.com/user/claude-thinking-audit)
Data Integrity
- All data is captured passively from Anthropic's own HTTP response headers
- No sampling, estimation, or interpolation - every API response is recorded
- Timestamps are from system clock, synchronized via NTP
- Database is available for independent verification
Reproducibility
The monitoring tools are open source and can be deployed by any user to independently verify these findings. The anomalous burn rate pattern should be reproducible by any Max subscriber during the affected time periods.
---
Claude Model
Not sure / Multiple models
Is this a regression?
Yes, this worked in a previous version
Last Working Version
_No response_
Claude Code Version
2.1.29
Platform
Anthropic API
Operating System
Ubuntu/Debian Linux
Terminal/Shell
Xterm
Additional Information
<img width="1836" height="782" alt="Image" src="https://github.com/user-attachments/assets/9f6d966d-2f22-44ac-800e-aa4d4a0c8e7b" />
14 Comments
@argosdevo-svg : I applaud the analysis and the detailed effort. If I had not been previously demotivated/demoralized by Anthropic's lack of responsiveness, I would have done something similar. Please note that, as of this writing, the links to
claude-thinking-audittool are broken in your initial post: https://github.com/user/claude-thinking-audit should be https://github.com/argosdevo-svg/claude-thinking-audit.Another factor to consider: was the 5-minute cache warm when the requests were made? My own experimentation shows that far less quota is consumed if I issue a new prompt within 5 minutes of the end Claude's previous work cycle (final assistant message in agentic workflow). While this is an inconvenient pattern of use, it has made a huge difference in the number of user prompts that I can issue in a given 5-hour period. Basically, treat Claude Code like a real-time strategy game which needs constant attention and response. Only take longer than 5 minutes to write a user message after a compaction or at the start of a new session, when the cache is already busted/cold and there are fewer tokens in the window. That said, I do completely agree that Anthropic is violating its own advertised limits, regardless of how Claude Code is being used.
Awesome observation and thank you for pointing out the link was broken!
To test this hypothesis with my dataset, I plan to:
1) Compute “time since last assistant message” for each request.
2) Bucket requests into <5 min vs ≥5 min.
3) Compare median tokens‑per‑1% quota and burn rates across buckets.
4) Treat compaction events and new sessions as cold‑start boundaries.
Even if a cache effect exists, it would still need to be documented and reflected in the advertised “5‑hour window” and “20x usage” claims. If caching materially changes quota cost, the behavior should be deterministic and disclosed.
Cache-Warmth Hypothesis - Tested
Thanks @emcd for the observation about cache warmth affecting quota burn. It's a reasonable hypothesis given Anthropic's documented 5-minute cache TTL.
I ran the analysis on my dataset to check if prompt timing correlates with quota cost:
Results from 3,429 quota samples + 8,444 audit samples:
Cold/Warm quota delta ratio 1.08x
Correlation (gap_minutes vs quota_delta) 0.024
Cache hit rate <5 min gap 97.7%
Cache hit rate >30 min gap 93.8%
What this suggests for my specific usage pattern:
The cache appears to stay warm even during longer gaps - hit rates remain 93%+ across all time buckets. The quota cost difference between "warm" and "cold" requests is negligible (1.08x).
This doesn't mean the cache effect doesn't exist in general - it may well matter for different usage patterns (longer idle periods, session restarts, compactions). But for my dataset, it doesn't appear to explain the 6x variance I observed.
@emcd's core point stands though: even if caching were the explanation, Anthropic should document how it affects the advertised "20x usage" and "5-hour window" claims. Users shouldn't need to reverse-engineer optimal prompting cadence.
27 minutes and im at 33% of my max subscription 5h quota with only one session running, this is getting ridiculous.
It seems the model got a bit more juice in it, and they try to balance the cost with less context window quota and higher burn rates.
Fraud report template is included in the audit tool.
<img width="1706" height="861" alt="Image" src="https://github.com/user-attachments/assets/ddbab603-c411-403b-88c2-6a794b25ec7b" />
Here is what Claude himself thinks of it:
Why is this happening?
Adding another data point to the billing inconsistency. After Claude Code v2.1.51 update, my 1M context sessions began consuming Extra Usage without notice. Feb 23 (v2.1.50): 644 Opus 1M calls, 85M cache_read tokens — no Extra Usage charged. Feb 25 (v2.1.51+): 392 calls, 80M tokens — $48.79 Extra Usage charged. Same workload, different billing. Full evidence in #28927.
Any news on this bug?
I want to add a note here;
ephemeral_5m_input_tokensin the jsonl files would state the tokens used for responses.ephemeral_1h_input_tokenswould be zeros all over. Through testing, the cache would live more than 5 minutes; cache was hit after more than 5 minutes of inactivity.Today I rechecked. Now
ephemeral_1h_input_tokensis actually not always zero.Must have been a bug fixed from Jan 27th onwards, not documented publicly, v2.1.20;
<details>
<summary>
<body>
<table>
<thead>
<tr>
<th>Period</th>
<th>5m!=0 occ</th>
<th>1h!=0 occ</th>
<th>5m=0 occ</th>
<th>1h=0 occ</th>
<th>5m!=0 files</th>
<th>1h!=0 files</th>
<th>5m=0 files</th>
<th>1h=0 files</th>
</tr>
</thead>
<tbody>
<tr>
<td>W49</td>
<td>49</td>
<td>0</td>
<td>0</td>
<td>49</td>
<td>1</td>
<td>0</td>
<td>0</td>
<td>1</td>
</tr>
<tr>
<td>W51</td>
<td>3495</td>
<td>0</td>
<td>3</td>
<td>3498</td>
<td>19</td>
<td>0</td>
<td>3</td>
<td>19</td>
</tr>
<tr>
<td>W52</td>
<td>3283</td>
<td>0</td>
<td>2</td>
<td>3285</td>
<td>36</td>
<td>0</td>
<td>2</td>
<td>36</td>
</tr>
<tr>
<td>W1</td>
<td>4109</td>
<td>0</td>
<td>48</td>
<td>4157</td>
<td>38</td>
<td>0</td>
<td>28</td>
<td>56</td>
</tr>
<tr>
<td>W2</td>
<td>1104</td>
<td>0</td>
<td>0</td>
<td>1104</td>
<td>8</td>
<td>0</td>
<td>0</td>
<td>8</td>
</tr>
<tr>
<td>W3</td>
<td>2234</td>
<td>0</td>
<td>1</td>
<td>2235</td>
<td>13</td>
<td>0</td>
<td>1</td>
<td>13</td>
</tr>
<tr>
<td>W4</td>
<td>28946</td>
<td>0</td>
<td>87</td>
<td>29033</td>
<td>147</td>
<td>0</td>
<td>36</td>
<td>152</td>
</tr>
<tr>
<td>2025-01-26</td>
<td>8080</td>
<td>0</td>
<td>31</td>
<td>8111</td>
<td>29</td>
<td>0</td>
<td>12</td>
<td>31</td>
</tr>
<tr>
<td>2025-01-27</td>
<td>3427</td>
<td>1487</td>
<td>1505</td>
<td>3445</td>
<td>42</td>
<td>12</td>
<td>21</td>
<td>43</td>
</tr>
<tr>
<td>2025-01-28</td>
<td>2941</td>
<td>336</td>
<td>349</td>
<td>2954</td>
<td>48</td>
<td>12</td>
<td>18</td>
<td>48</td>
</tr>
<tr>
<td>2025-01-29</td>
<td>3754</td>
<td>6011</td>
<td>6053</td>
<td>3796</td>
<td>37</td>
<td>41</td>
<td>51</td>
<td>48</td>
</tr>
<tr>
<td>2025-01-30</td>
<td>1</td>
<td>1575</td>
<td>1581</td>
<td>7</td>
<td>1</td>
<td>17</td>
<td>18</td>
<td>4</td>
</tr>
<tr>
<td>2025-01-31</td>
<td>34</td>
<td>4525</td>
<td>4527</td>
<td>36</td>
<td>1</td>
<td>30</td>
<td>30</td>
<td>2</td>
</tr>
<tr>
<td>2025-02-01</td>
<td>0</td>
<td>13</td>
<td>13</td>
<td>0</td>
<td>0</td>
<td>1</td>
<td>1</td>
<td>0</td>
</tr>
<tr>
<td>W6</td>
<td>443</td>
<td>8906</td>
<td>8938</td>
<td>475</td>
<td>6</td>
<td>84</td>
<td>84</td>
<td>23</td>
</tr>
<tr>
<td>W7</td>
<td>99</td>
<td>14353</td>
<td>14384</td>
<td>130</td>
<td>4</td>
<td>119</td>
<td>120</td>
<td>22</td>
</tr>
<tr>
<td>W8</td>
<td>0</td>
<td>14843</td>
<td>14875</td>
<td>32</td>
<td>0</td>
<td>144</td>
<td>144</td>
<td>20</td>
</tr>
<tr>
<td>W9</td>
<td>35</td>
<td>11219</td>
<td>11246</td>
<td>62</td>
<td>2</td>
<td>123</td>
<td>125</td>
<td>13</td>
</tr>
<tr>
<td>W10</td>
<td>7771</td>
<td>20850</td>
<td>20921</td>
<td>7842</td>
<td>54</td>
<td>231</td>
<td>234</td>
<td>80</td>
</tr>
<tr>
<td>W11</td>
<td>6037</td>
<td>18839</td>
<td>18899</td>
<td>6097</td>
<td>79</td>
<td>163</td>
<td>165</td>
<td>90</td>
</tr>
<tr>
<td>W12</td>
<td>6</td>
<td>290</td>
<td>292</td>
<td>8</td>
<td>1</td>
<td>2</td>
<td>2</td>
<td>2</td>
</tr>
</tbody>
</table>
</body>
</html>
</summary>
<body>
<pre><code>#!/usr/bin/env bash
set -euo pipefail
DIR="${1:-.}"
GLOB='[0-9a-f]-[0-9a-f]-[0-9a-f]-[0-9a-f]-[0-9a-f]*.jsonl'
cd "$DIR"
<span class="comment"># Step 1: Find all matching files</span>
echo "Finding files..." >&2
mapfile -t files < <(rg --glob "$GLOB" -l 'ephemeral_5m_input_tokens' 2> /dev/null)
echo "Found ${#files[@]} files" >&2
<span class="comment"># Step 2: Build file->date map</span>
echo "Reading file dates..." >&2
declare -A file_date
for f in "${files[@]}"; do
file_date["$f"]=$(date -d @"$(stat --format='%Y' -- "$f")" +%Y-%m-%d)
done
<span class="comment"># Step 3: Count matches per file with single rg passes (4 total, not 4 per file)</span>
echo "Counting 5m!=0..." >&2
declare -A count_5m_nz
while IFS=: read -r f c; do
count_5m_nz["$f"]=$c
done < <(rg -c 'ephemeral_5m_input_tokens":\s*[1-9]' --glob "$GLOB" 2> /dev/null || true)
echo "Counting 5m=0..." >&2
declare -A count_5m_z
while IFS=: read -r f c; do
count_5m_z["$f"]=$c
done < <(rg -c 'ephemeral_5m_input_tokens":\s*0[,}]' --glob "$GLOB" 2> /dev/null || true)
echo "Counting 1h!=0..." >&2
declare -A count_1h_nz
while IFS=: read -r f c; do
count_1h_nz["$f"]=$c
done < <(rg -c 'ephemeral_1h_input_tokens":\s*[1-9]' --glob "$GLOB" 2> /dev/null || true)
echo "Counting 1h=0..." >&2
declare -A count_1h_z
while IFS=: read -r f c; do
count_1h_z["$f"]=$c
done < <(rg -c 'ephemeral_1h_input_tokens":\s*0[,}]' --glob "$GLOB" 2> /dev/null || true)
<span class="comment"># Step 4: Aggregate by date</span>
echo "Aggregating..." >&2
declare -A s5nz s5z s1nz s1z f5nz f5z f1nz f1z
for f in "${files[@]}"; do
d="${file_date[$f]}"
v="${count_5m_nz[$f]:-0}"
s5nz[$d]=$(( ${s5nz[$d]:-0} + v ))
(( v > 0 )) && f5nz[$d]=$(( ${f5nz[$d]:-0} + 1 ))
v="${count_5m_z[$f]:-0}"
s5z[$d]=$(( ${s5z[$d]:-0} + v ))
(( v > 0 )) && f5z[$d]=$(( ${f5z[$d]:-0} + 1 ))
v="${count_1h_nz[$f]:-0}"
s1nz[$d]=$(( ${s1nz[$d]:-0} + v ))
(( v > 0 )) && f1nz[$d]=$(( ${f1nz[$d]:-0} + 1 ))
v="${count_1h_z[$f]:-0}"
s1z[$d]=$(( ${s1z[$d]:-0} + v ))
(( v > 0 )) && f1z[$d]=$(( ${f1z[$d]:-0} + 1 ))
done
echo "| Date | 5m!=0 occ | 1h!=0 occ | 5m=0 occ | 1h=0 occ | 5m!=0 files | 1h!=0 files | 5m=0 files | 1h=0 files |"
echo "|------|----------:|----------:|---------:|---------:|------------:|------------:|-----------:|-----------:|"
<span class="comment"># Step 5: Print sorted</span>
echo "" >&2
for d in "${!s5nz[@]}"; do
printf "| %s | %d | %d | %d | %d | %d | %d | %d | %d |\n" \
"$d" \
"${s5nz[$d]:-0}" "${s1nz[$d]:-0}" \
"${s5z[$d]:-0}" "${s1z[$d]:-0}" \
"${f5nz[$d]:-0}" "${f1nz[$d]:-0}" \
"${f5z[$d]:-0}" "${f1z[$d]:-0}"
done | sort
</code></pre>
</body>
</details>
I've done extensive analysis on this exact problem. Using ccusage_go (open-source Claude Code usage tracker), I found that Cache Read tokens consumed 97.7% of my session costs — API actual cost was $1.47, total billed cost was $64.98 (a 44x markup). Cache also degrades instruction following in long sessions, which I documented with per-turn JSONL analysis.
Full write-up with data, community issue references, and Claude Code's own self-analysis report:
https://blog.sd.idv.tw/en/posts/2026-03-25_claude-code-cache-trap/
Tool: https://github.com/SDpower/ccusage_go
The undisclosed quota changes continue. Max 20 ($200/mo), v2.1.89, April 1 2026: 100% in ~70 min after reset. No transparency from Anthropic on what changed.
Support acknowledged "resolved incidents" on Mar 31-Apr 1 but offered no explanation or fix — just "enable extra usage" (pay more).
Full report: #41788
Related: #38335, #38239, #40790, #41055, #41663, #40903
I've been experiencing the same issue on Max 20 ($200/mo) — rate limit 100% exhausted in ~70 minutes.
After setting up a monitoring proxy using the official
ANTHROPIC_BASE_URLenv var, I identified two cache bugs as the root cause (#40524, #34629) and measured the impact: cache read ratio dropped to 4.3%, meaning ~20x token inflation per turn. After applying workarounds it stabilized at 89-99%.Full analysis with per-request measured data, safe workarounds, and community references (including cc-cache-fix): https://github.com/ArkNill/claude-code-cache-analysis
Update (April 2): v2.1.90 has significantly improved cache efficiency — benchmark shows 95-99% cache read in stable sessions (both npm and standalone installations).
If you're still affected:
claude update(ornpm install -g @anthropic-ai/claude-code)"DISABLE_AUTOUPDATER": "1"to~/.claude/settings.jsonenv section--resume(still broken)Note: server-side quota issues (org-level pool sharing, accounting mismatches) remain unresolved — the above fixes the client-side cache drain only.
Benchmark data: https://github.com/ArkNill/claude-code-cache-analysis
April 3 update: v2.1.91 fixes the cache regression that caused the worst drain. If you are still hitting limits after updating, there are additional unfixed mechanisms: a 200K tool result budget cap, a client-side false rate limiter, and silent context stripping — all confirmed via proxy testing. Anthropic acknowledged peak-hour tightening on X (Lydia Hallie) but stated "none were over-charging you." Measured data and analysis: claude-code-cache-analysis
Same issue here. Max 20x subscriber.
Since the May 27–28 server-side change that merged Claude Design's previously separate quota pool into the shared Claude.ai + Claude Code allowance, my Cowork usage is visibly eating into my main quota.
This week (June 1–4, 2026) I'm running multiple parallel Cowork sessions (Claude Design agents) and quota burn rate feels 2–3x what it was two weeks ago. The extra capacity I used to have via Cowork's separate budget is gone.
I subscribed to Max 20x specifically because Cowork had its own meter — that's what made parallel agent workflows viable on this plan. The silent merge undermines that value proposition.
Asking for one of:
This is a material change to the Max plan that wasn't disclosed at purchase. Please give us back the separate Cowork budget. Thanks for transparent communication going forward.