[BUG] Inconsistent and Undisclosed Quota Accounting Changes in Claude Max Plan. Legal liability claim ready.

Status Open
Reported on v2.1.29
Maintainer reply ✓ Yes — emcd
Activity 14 comments · opened Feb 1, 2026
💡 Likely answer: A maintainer (emcd, contributor) responded on this thread — see the highlighted reply below.

Preflight Checklist

  • [x] I have searched existing issues and this hasn't been reported yet
  • [x] This is a single bug report (please file separate reports for different bugs)
  • [x] I am using the latest version of Claude Code

What's Wrong?

Bug Report: Inconsistent and Undisclosed Quota Accounting Changes in Claude Max Plan

Filed: 2026-02-01
Plan: Claude Max 20x ($200/month)
Affected Period: January 30 - February 1, 2026
Severity: Critical - Service degradation without disclosure

---

Executive Summary

Instrumented monitoring of Anthropic's rate limit headers reveals inconsistent quota consumption rates that cannot be explained by the "holiday bonus expiration" cited by Anthropic. The same user, same plan, same workload type shows burn rates varying from 5.6%/hour to 59.9%/hour within the same 48-hour period. This 10x variance constitutes either a bug in quota accounting or an undisclosed server-side change to rate limiting behavior.

---

Evidence

Methodology

All data was collected by intercepting HTTP responses from api.anthropic.com/v1/messages via a local mitmproxy instance. Rate limit values are extracted directly from Anthropic's own response headers:

x-ratelimit-5h-utilization: 0.81
x-ratelimit-7d-utilization: 0.23
x-ratelimit-5h-status: allowed

Data is stored in SQLite with 5,396+ samples spanning January 30 - February 1, 2026. No sampling bias - every API response is recorded.

Session Analysis (Quota Reset to Reset)

| # | Start | End | Start% | End% | Duration | Rate (%/hr) | Status |
|---|-------|-----|--------|------|----------|-------------|--------|
| 1 | Jan 30 18:41 | Jan 31 02:37 | 9% | 83% | 7.9h | 9.3%/hr | Normal |
| 2 | Jan 31 10:02 | Jan 31 20:24 | 0% | 100% | 10.4h | 9.6%/hr | Normal |
| 3 | Jan 31 20:28 | Jan 31 22:01 | 0% | 86% | 1.5h | 56.0%/hr | 2.8x fast |
| 4 | Jan 31 22:18 | Feb 01 15:00 | 0% | 94% | 16.7h | 5.6%/hr | Normal |
| 5 | Feb 01 15:00 | Feb 01 16:40 | 0% | 100% | 1.7h | 59.9%/hr | 3.0x fast |
| 6 | Feb 01 16:43 | Feb 01 18:30 | 0% | 100% | 1.8h | 56.1%/hr | 2.8x fast |
| 7 | Feb 01 20:01 | ongoing | 0% | 36% | 0.9h | 40.0%/hr | 2.0x fast |

Key Observation

Sessions 1, 2, and 4 all occur AFTER the holiday bonus expiration (Dec 31) and show normal ~10%/hr rates. If the holiday bonus expiration were the sole explanation, ALL sessions should show the same rate. Instead:

  • Normal sessions: 5.6 - 9.6%/hr (consistent with advertised 5-hour window)
  • Anomalous sessions: 40.0 - 59.9%/hr (3-6x the expected rate)

This variance occurs on the same account, same plan, same day, ruling out:

  • Holiday bonus as the explanation (normal sessions exist post-holiday)
  • User behavior differences (same operator, same workload type)
  • Plan differences (same Max 20x account throughout)

Token-to-Quota Correlation

From 722 quota-increasing request pairs:

| Metric | Value |
|--------|-------|
| Median tokens per 1% quota | 2,517 |
| Mean tokens per 1% quota | 39,152 |
| Min (worst efficiency) | 12,300 tokens per 1% |
| Max (best efficiency) | 18,531,900 tokens per 1% |

The 1,500x spread between min and max tokens-per-percent is not explainable by cache behavior differences alone. This suggests the cost-per-token in quota terms is not deterministic.

---

Anthropic's Advertised Promises vs. Reality

What Anthropic Advertises

From claude.com/pricing and Claude Help Center:

"Max plan: 20x more usage than the Pro plan" "About 900+ messages per 5-hour window" "Rolling 5-hour windows rather than daily caps"

From Anthropic's official announcement:

"Introducing a new Max plan for Claude. It's flexible, with options for 5x or 20x more usage compared to our Pro plan."

What Users Actually Experience

During anomalous sessions (3, 5, 6, 7), the 5-hour quota window is exhausted in 1.5-1.8 hours. This means:

  • Advertised: 5-hour window -> 900+ messages
  • Actual: ~1.7-hour window -> significantly fewer messages
  • Effective multiplier: Not 20x Pro, but approximately 6-7x Pro during anomalous periods

Users are paying $200/month for a 20x multiplier and intermittently receiving what amounts to a 6-7x multiplier with no disclosure, notification, or compensation.

---

Legal Analysis

1. Breach of Express Warranty (UCC 2-313)

Anthropic's marketing materials constitute express warranties under the Uniform Commercial Code:

  • "20x more usage than Pro" is a specific, quantified promise
  • "900+ messages per 5-hour window" is a specific performance guarantee
  • "Rolling 5-hour windows" implies consistent, predictable behavior

When users intermittently receive 1.7-hour windows instead of 5-hour windows, the express warranty is breached.

2. California Unfair Competition Law (Bus. & Prof. Code 17200)

Anthropic is headquartered in San Francisco, California. California's UCL prohibits:

  • Unfair business practices: Reducing service levels without notice or compensation while continuing to charge full price is "substantially injurious to consumers" (unfair prong)
  • Fraudulent practices: Advertising "20x usage" and "5-hour windows" while delivering 1.7-hour windows during anomalous periods is conduct "likely to deceive" consumers (fraudulent prong)
  • Unlawful practices: If the above also violates FTC Act Section 5, it's actionable under the UCL's unlawful prong as well

UCL imposes strict liability - Anthropic's intent is irrelevant. The fact that the practice is unfair or deceptive is sufficient.

3. FTC Act Section 5 - Deceptive Practices

The FTC prohibits deceptive acts or practices in commerce. The FTC's penalty offenses framework for bait-and-switch applies when:

  • A service is advertised with specific capabilities ("20x", "5-hour windows")
  • The advertised service is not consistently available as described
  • Consumers pay based on the advertised description

While the FTC's Click-to-Cancel rule was vacated by the Eighth Circuit in July 2025, the FTC retains enforcement authority under Section 5 and ROSCA for deceptive subscription practices. Penalties can reach $53,088 per violation.

4. Anthropic's Terms of Service Defense - And Why It's Insufficient

Anthropic's Consumer Terms (anthropic.com/legal/consumer-terms) state:

"We may sometimes add or remove features, increase or decrease capacity limits..." "We reserve the right to modify, suspend, or discontinue the Services... at any time without notice to you."

This defense has significant limitations:

  1. Unconscionability: A take-it-or-leave-it clause that allows unlimited service degradation while maintaining full pricing may be unconscionable under California law, particularly for a $200/month consumer subscription.
  1. Marketing overrides ToS: When marketing materials make specific quantified promises ("20x", "900+ messages", "5-hour windows"), those promises create binding obligations that cannot be entirely disclaimed by buried ToS clauses. See Weinstat v. Dentsply International, 180 Cal.App.4th 1213 (2010).
  1. Implied covenant of good faith: Even if Anthropic can modify limits, they must do so in good faith. Silently reducing effective capacity by 3x while maintaining pricing and continuing to advertise "20x" and "5-hour windows" arguably violates the implied covenant.
  1. The inconsistency is the problem: If Anthropic uniformly reduced limits post-holiday, users could evaluate and decide whether to continue subscribing. The intermittent nature of the degradation (normal one session, 3x faster the next) prevents users from making informed purchasing decisions.

---

Corroborating Community Reports

This is not an isolated experience. Multiple GitHub issues and community reports document the same pattern:

| Issue | Title | Upvotes | Key Finding |
|-------|-------|---------|-------------|
| #16157 | Instantly hitting usage limits with Max subscription | 490+ | Never hit limits in 3 months, now exhausted in 2 hours |
| #17084 | Opus 4.5 limits significantly reduced since Jan 2026 | 237+ | "Most restrictive since Opus 4.5 launch" |
| #16868 | Credit consumption rate increased 3-5x after Jan 1 | - | "CRITICAL: Max 5x plan credits depleting 3-5x faster" |
| #20767 | Pro Quota reduced to ~20 prompts per 5 hours | - | "Significantly lower than pre-holiday standard levels" |
| #19673 | "You've hit your limit" at 84% usage | - | Rate limiting before reaching advertised capacity |
| #16270 | Usage limits bugged after double limits expired | - | Opened Jan 4, day 4 after holiday bonus ended |

Media coverage: The Register reported a community-estimated ~60% reduction in token usage limits, with developers' criticism allegedly being censored in Anthropic's Discord.

---

Anthropic's Response - And Why It's Inadequate

Anthropic's official position (via The Register, Jan 5, 2026):

"The concerns appear to be largely a response to the resumption of normal limits."

Our data directly contradicts this:

  1. Sessions 1, 2, and 4 show "normal limits" (~10%/hr) - these ARE post-holiday
  2. Sessions 3, 5, 6, and 7 show 3-6x consumption - these are ALSO post-holiday
  3. Both rate patterns occur on the same account, same day

If limits simply "returned to normal," they would return to a consistent normal rate. The 10x variance between sessions is either:

  • A bug in quota accounting (which Anthropic should acknowledge and fix)
  • An intentional server-side change (which should be disclosed per ToS transparency obligations)

---

Evidence Collection Tool: Quota Tracking Dashboard

To systematically collect and verify the evidence presented in this report, we developed an open-source Quota Tracking Dashboard as part of the claude-thinking-audit project. This tool is available for any Claude subscriber to independently reproduce our findings.

What It Does

The dashboard is a dedicated QUOTA tab added to a mitmproxy-based monitoring UI that passively intercepts Claude API responses and extracts rate limit data from Anthropic's own HTTP headers. It provides:

  1. Real-Time Status Cards — Current 5-hour and 7-day quota utilization, live burn rate calculation, and estimated time to 100% exhaustion with deviation warnings when burn rate exceeds expected levels
  1. Burn Rate Chart — ASCII sparkline visualization of quota progression over selectable time ranges (1h, 6h, 24h, 7d), clearly showing the sawtooth pattern of quota cycles and resets
  1. Session History Table — Automatic detection of quota reset events to identify individual usage sessions, with per-session burn rate calculation and color-coded status flags (green=normal, red=anomalously fast)
  1. Token-to-Quota Correlation — Statistical analysis of how many tokens correspond to 1% of quota, revealing the non-deterministic nature of Anthropic's quota accounting
  1. Systemd Log Viewer — Ingests rate limit entries from journalctl into SQLite for fast querying, showing a scrollable table of timestamped quota snapshots directly from the service logs
  1. One-Click Evidence Export — Generates a downloadable JSON report containing all of the above data in a structured format suitable for filing bug reports, including quota progression history, session analysis, correlation statistics, and raw log entries

Why This Matters

Before this tool, users had no way to verify their quota usage claims. Anthropic's own /context command in Claude Code has a known bug (showing 0% when usage is non-zero), and there is no user-facing quota dashboard. Users reporting issues on GitHub could only say "it feels faster" — now they can provide timestamped, header-sourced evidence.

The evidence in this report was generated using the Export Evidence feature of this tool. Every data point traces back to an x-ratelimit-* header in an actual API response stored in our SQLite database.

Sources

What Should Happen?

Requested Resolution

Immediate

  1. Acknowledge the inconsistency in quota accounting
  2. Publish the formula used to calculate quota consumption per request
  3. Fix the variance that causes 10x differences in effective burn rates

Structural

  1. Transparency: Display actual quota consumption per request in Claude Code's UI (Anthropic already sends x-ratelimit-* headers - surface them to users)
  2. Notification: Notify users BEFORE service level changes, not after complaints
  3. Consistency: Ensure quota accounting is deterministic - same tokens should always cost the same quota percentage

Compensatory

  1. Credit affected users for periods where quota burned at anomalous rates
  2. Update marketing materials to reflect actual achievable usage levels, or restore advertised levels

---

Error Messages/Logs

Steps to Reproduce

How to Reproduce

Any Claude subscriber can deploy this tool:

  1. Install mitmproxy with the custom rate-limit extraction addon
  2. Route Claude Code traffic through the local proxy
  3. Open the Quota Tracking Dashboard in a browser
  4. Use Claude Code normally — the tool passively records every API response
  5. Click "Export Evidence Report" to generate a structured JSON report

---

Appendix: Data Collection Methodology

Tools Used

  • mitmproxy (v11.x) intercepting HTTPS traffic to api.anthropic.com
  • Custom addon extracting rate limit headers from every API response
  • SQLite database storing 5,396+ samples with timestamps, token counts, and quota utilization values
  • claude-thinking-audit quota tracking dashboard for visualization and analysis (https://github.com/user/claude-thinking-audit)

Data Integrity

  • All data is captured passively from Anthropic's own HTTP response headers
  • No sampling, estimation, or interpolation - every API response is recorded
  • Timestamps are from system clock, synchronized via NTP
  • Database is available for independent verification

Reproducibility

The monitoring tools are open source and can be deployed by any user to independently verify these findings. The anomalous burn rate pattern should be reproducible by any Max subscriber during the affected time periods.

---

Claude Model

Not sure / Multiple models

Is this a regression?

Yes, this worked in a previous version

Last Working Version

_No response_

Claude Code Version

2.1.29

Platform

Anthropic API

Operating System

Ubuntu/Debian Linux

Terminal/Shell

Xterm

Additional Information

<img width="1836" height="782" alt="Image" src="https://github.com/user-attachments/assets/9f6d966d-2f22-44ac-800e-aa4d4a0c8e7b" />

View original on GitHub ↗

14 Comments

emcd contributor · 6 months ago

@argosdevo-svg : I applaud the analysis and the detailed effort. If I had not been previously demotivated/demoralized by Anthropic's lack of responsiveness, I would have done something similar. Please note that, as of this writing, the links to claude-thinking-audit tool are broken in your initial post: https://github.com/user/claude-thinking-audit should be https://github.com/argosdevo-svg/claude-thinking-audit.

Another factor to consider: was the 5-minute cache warm when the requests were made? My own experimentation shows that far less quota is consumed if I issue a new prompt within 5 minutes of the end Claude's previous work cycle (final assistant message in agentic workflow). While this is an inconvenient pattern of use, it has made a huge difference in the number of user prompts that I can issue in a given 5-hour period. Basically, treat Claude Code like a real-time strategy game which needs constant attention and response. Only take longer than 5 minutes to write a user message after a compaction or at the start of a new session, when the cache is already busted/cold and there are fewer tokens in the window. That said, I do completely agree that Anthropic is violating its own advertised limits, regardless of how Claude Code is being used.

ghost · 6 months ago

Awesome observation and thank you for pointing out the link was broken!

To test this hypothesis with my dataset, I plan to:

1) Compute “time since last assistant message” for each request.
2) Bucket requests into <5 min vs ≥5 min.
3) Compare median tokens‑per‑1% quota and burn rates across buckets.
4) Treat compaction events and new sessions as cold‑start boundaries.

Even if a cache effect exists, it would still need to be documented and reflected in the advertised “5‑hour window” and “20x usage” claims. If caching materially changes quota cost, the behavior should be deterministic and disclosed.

ghost · 6 months ago

Cache-Warmth Hypothesis - Tested

Thanks @emcd for the observation about cache warmth affecting quota burn. It's a reasonable hypothesis given Anthropic's documented 5-minute cache TTL.

I ran the analysis on my dataset to check if prompt timing correlates with quota cost:

Results from 3,429 quota samples + 8,444 audit samples:

Cold/Warm quota delta ratio 1.08x

Correlation (gap_minutes vs quota_delta) 0.024

Cache hit rate <5 min gap 97.7%

Cache hit rate >30 min gap 93.8%

What this suggests for my specific usage pattern:

The cache appears to stay warm even during longer gaps - hit rates remain 93%+ across all time buckets. The quota cost difference between "warm" and "cold" requests is negligible (1.08x).

This doesn't mean the cache effect doesn't exist in general - it may well matter for different usage patterns (longer idle periods, session restarts, compactions). But for my dataset, it doesn't appear to explain the 6x variance I observed.

@emcd's core point stands though: even if caching were the explanation, Anthropic should document how it affects the advertised "20x usage" and "5-hour window" claims. Users shouldn't need to reverse-engineer optimal prompting cadence.

ghost · 6 months ago

27 minutes and im at 33% of my max subscription 5h quota with only one session running, this is getting ridiculous.

It seems the model got a bit more juice in it, and they try to balance the cost with less context window quota and higher burn rates.

Fraud report template is included in the audit tool.

<img width="1706" height="861" alt="Image" src="https://github.com/user-attachments/assets/ddbab603-c411-403b-88c2-6a794b25ec7b" />

privacyguy123 · 6 months ago
│ 2026-02-03 │ - sonnet-4-5            │       127 │       493 │       304,565 │   3,996,488 │     4,301,673 │       $2.35 │

Here is what Claude himself thinks of it:

Key findings:
Non-deterministic quota accounting. The same token consumption burns wildly different percentages of quota - ranging from 5.6%/hour to 59.9%/hour within 48 hours on the same Max 20x plan.
"Token-to-Quota Correlation" analysis reveals Anthropic's quota system is inconsistent - same tokens don't always cost the same percentage.
Your numbers validate the bug. With only 127 input + 493 output + 304K cache creation today, hitting 68% utilization is absurdly high. Cache reads (4M) explicitly don't count toward limits.
Multiple users affected. The bug reporter used mitmproxy to intercept x-ratelimit-5h-utilization headers directly from Anthropic's API responses, proving the inconsistency is server-side.
The issue suggests Anthropic either has a quota accounting bug or made undisclosed changes to rate limiting after their "holiday bonus" expired. Either way, your quota is burning 5-10x faster than your actual token consumption justifies.

Why is this happening?

paulseongminpark · 6 months ago

Adding another data point to the billing inconsistency. After Claude Code v2.1.51 update, my 1M context sessions began consuming Extra Usage without notice. Feb 23 (v2.1.50): 644 Opus 1M calls, 85M cache_read tokens — no Extra Usage charged. Feb 25 (v2.1.51+): 392 calls, 80M tokens — $48.79 Extra Usage charged. Same workload, different billing. Full evidence in #28927.

mik3fly · 5 months ago

Any news on this bug?

NubeBuster · 5 months ago

I want to add a note here;

ephemeral_5m_input_tokens in the jsonl files would state the tokens used for responses. ephemeral_1h_input_tokens would be zeros all over. Through testing, the cache would live more than 5 minutes; cache was hit after more than 5 minutes of inactivity.

Today I rechecked. Now ephemeral_1h_input_tokens is actually not always zero.

Must have been a bug fixed from Jan 27th onwards, not documented publicly, v2.1.20;

<details>
<summary>

<body>
<table>
<thead>
<tr>
<th>Period</th>
<th>5m!=0 occ</th>
<th>1h!=0 occ</th>
<th>5m=0 occ</th>
<th>1h=0 occ</th>
<th>5m!=0 files</th>
<th>1h!=0 files</th>
<th>5m=0 files</th>
<th>1h=0 files</th>
</tr>
</thead>
<tbody>
<tr>
<td>W49</td>
<td>49</td>
<td>0</td>
<td>0</td>
<td>49</td>
<td>1</td>
<td>0</td>
<td>0</td>
<td>1</td>
</tr>
<tr>
<td>W51</td>
<td>3495</td>
<td>0</td>
<td>3</td>
<td>3498</td>
<td>19</td>
<td>0</td>
<td>3</td>
<td>19</td>
</tr>
<tr>
<td>W52</td>
<td>3283</td>
<td>0</td>
<td>2</td>
<td>3285</td>
<td>36</td>
<td>0</td>
<td>2</td>
<td>36</td>
</tr>
<tr>
<td>W1</td>
<td>4109</td>
<td>0</td>
<td>48</td>
<td>4157</td>
<td>38</td>
<td>0</td>
<td>28</td>
<td>56</td>
</tr>
<tr>
<td>W2</td>
<td>1104</td>
<td>0</td>
<td>0</td>
<td>1104</td>
<td>8</td>
<td>0</td>
<td>0</td>
<td>8</td>
</tr>
<tr>
<td>W3</td>
<td>2234</td>
<td>0</td>
<td>1</td>
<td>2235</td>
<td>13</td>
<td>0</td>
<td>1</td>
<td>13</td>
</tr>
<tr>
<td>W4</td>
<td>28946</td>
<td>0</td>
<td>87</td>
<td>29033</td>
<td>147</td>
<td>0</td>
<td>36</td>
<td>152</td>
</tr>
<tr>
<td>2025-01-26</td>
<td>8080</td>
<td>0</td>
<td>31</td>
<td>8111</td>
<td>29</td>
<td>0</td>
<td>12</td>
<td>31</td>
</tr>
<tr>
<td>2025-01-27</td>
<td>3427</td>
<td>1487</td>
<td>1505</td>
<td>3445</td>
<td>42</td>
<td>12</td>
<td>21</td>
<td>43</td>
</tr>
<tr>
<td>2025-01-28</td>
<td>2941</td>
<td>336</td>
<td>349</td>
<td>2954</td>
<td>48</td>
<td>12</td>
<td>18</td>
<td>48</td>
</tr>
<tr>
<td>2025-01-29</td>
<td>3754</td>
<td>6011</td>
<td>6053</td>
<td>3796</td>
<td>37</td>
<td>41</td>
<td>51</td>
<td>48</td>
</tr>
<tr>
<td>2025-01-30</td>
<td>1</td>
<td>1575</td>
<td>1581</td>
<td>7</td>
<td>1</td>
<td>17</td>
<td>18</td>
<td>4</td>
</tr>
<tr>
<td>2025-01-31</td>
<td>34</td>
<td>4525</td>
<td>4527</td>
<td>36</td>
<td>1</td>
<td>30</td>
<td>30</td>
<td>2</td>
</tr>
<tr>
<td>2025-02-01</td>
<td>0</td>
<td>13</td>
<td>13</td>
<td>0</td>
<td>0</td>
<td>1</td>
<td>1</td>
<td>0</td>
</tr>
<tr>
<td>W6</td>
<td>443</td>
<td>8906</td>
<td>8938</td>
<td>475</td>
<td>6</td>
<td>84</td>
<td>84</td>
<td>23</td>
</tr>
<tr>
<td>W7</td>
<td>99</td>
<td>14353</td>
<td>14384</td>
<td>130</td>
<td>4</td>
<td>119</td>
<td>120</td>
<td>22</td>
</tr>
<tr>
<td>W8</td>
<td>0</td>
<td>14843</td>
<td>14875</td>
<td>32</td>
<td>0</td>
<td>144</td>
<td>144</td>
<td>20</td>
</tr>
<tr>
<td>W9</td>
<td>35</td>
<td>11219</td>
<td>11246</td>
<td>62</td>
<td>2</td>
<td>123</td>
<td>125</td>
<td>13</td>
</tr>
<tr>
<td>W10</td>
<td>7771</td>
<td>20850</td>
<td>20921</td>
<td>7842</td>
<td>54</td>
<td>231</td>
<td>234</td>
<td>80</td>
</tr>
<tr>
<td>W11</td>
<td>6037</td>
<td>18839</td>
<td>18899</td>
<td>6097</td>
<td>79</td>
<td>163</td>
<td>165</td>
<td>90</td>
</tr>
<tr>
<td>W12</td>
<td>6</td>
<td>290</td>
<td>292</td>
<td>8</td>
<td>1</td>
<td>2</td>
<td>2</td>
<td>2</td>
</tr>
</tbody>
</table>
</body>
</html>

</summary>

<body>
<pre><code>#!/usr/bin/env bash
set -euo pipefail
DIR=&quot;${1:-.}&quot;
GLOB=&#39;[0-9a-f]-[0-9a-f]-[0-9a-f]-[0-9a-f]-[0-9a-f]*.jsonl&#39;
cd &quot;$DIR&quot;

<span class="comment"># Step 1: Find all matching files</span>
echo &quot;Finding files...&quot; &gt;&amp;2
mapfile -t files &lt; &lt;(rg --glob &quot;$GLOB&quot; -l &#39;ephemeral_5m_input_tokens&#39; 2&gt; /dev/null)
echo &quot;Found ${#files[@]} files&quot; &gt;&amp;2

<span class="comment"># Step 2: Build file-&gt;date map</span>
echo &quot;Reading file dates...&quot; &gt;&amp;2
declare -A file_date
for f in &quot;${files[@]}&quot;; do
file_date[&quot;$f&quot;]=$(date -d @&quot;$(stat --format=&#39;%Y&#39; -- &quot;$f&quot;)&quot; +%Y-%m-%d)
done

<span class="comment"># Step 3: Count matches per file with single rg passes (4 total, not 4 per file)</span>
echo &quot;Counting 5m!=0...&quot; &gt;&amp;2
declare -A count_5m_nz
while IFS=: read -r f c; do
count_5m_nz[&quot;$f&quot;]=$c
done &lt; &lt;(rg -c &#39;ephemeral_5m_input_tokens&quot;:\s*[1-9]&#39; --glob &quot;$GLOB&quot; 2&gt; /dev/null || true)

echo &quot;Counting 5m=0...&quot; &gt;&amp;2
declare -A count_5m_z
while IFS=: read -r f c; do
count_5m_z[&quot;$f&quot;]=$c
done &lt; &lt;(rg -c &#39;ephemeral_5m_input_tokens&quot;:\s*0[,}]&#39; --glob &quot;$GLOB&quot; 2&gt; /dev/null || true)

echo &quot;Counting 1h!=0...&quot; &gt;&amp;2
declare -A count_1h_nz
while IFS=: read -r f c; do
count_1h_nz[&quot;$f&quot;]=$c
done &lt; &lt;(rg -c &#39;ephemeral_1h_input_tokens&quot;:\s*[1-9]&#39; --glob &quot;$GLOB&quot; 2&gt; /dev/null || true)

echo &quot;Counting 1h=0...&quot; &gt;&amp;2
declare -A count_1h_z
while IFS=: read -r f c; do
count_1h_z[&quot;$f&quot;]=$c
done &lt; &lt;(rg -c &#39;ephemeral_1h_input_tokens&quot;:\s*0[,}]&#39; --glob &quot;$GLOB&quot; 2&gt; /dev/null || true)

<span class="comment"># Step 4: Aggregate by date</span>
echo &quot;Aggregating...&quot; &gt;&amp;2
declare -A s5nz s5z s1nz s1z f5nz f5z f1nz f1z
for f in &quot;${files[@]}&quot;; do
d=&quot;${file_date[$f]}&quot;
v=&quot;${count_5m_nz[$f]:-0}&quot;
s5nz[$d]=$(( ${s5nz[$d]:-0} + v ))
(( v &gt; 0 )) &amp;&amp; f5nz[$d]=$(( ${f5nz[$d]:-0} + 1 ))
v=&quot;${count_5m_z[$f]:-0}&quot;
s5z[$d]=$(( ${s5z[$d]:-0} + v ))
(( v &gt; 0 )) &amp;&amp; f5z[$d]=$(( ${f5z[$d]:-0} + 1 ))
v=&quot;${count_1h_nz[$f]:-0}&quot;
s1nz[$d]=$(( ${s1nz[$d]:-0} + v ))
(( v &gt; 0 )) &amp;&amp; f1nz[$d]=$(( ${f1nz[$d]:-0} + 1 ))
v=&quot;${count_1h_z[$f]:-0}&quot;
s1z[$d]=$(( ${s1z[$d]:-0} + v ))
(( v &gt; 0 )) &amp;&amp; f1z[$d]=$(( ${f1z[$d]:-0} + 1 ))
done

echo &quot;| Date | 5m!=0 occ | 1h!=0 occ | 5m=0 occ | 1h=0 occ | 5m!=0 files | 1h!=0 files | 5m=0 files | 1h=0 files |&quot;
echo &quot;|------|----------:|----------:|---------:|---------:|------------:|------------:|-----------:|-----------:|&quot;

<span class="comment"># Step 5: Print sorted</span>
echo &quot;&quot; &gt;&amp;2
for d in &quot;${!s5nz[@]}&quot;; do
printf &quot;| %s | %d | %d | %d | %d | %d | %d | %d | %d |\n&quot; \
&quot;$d&quot; \
&quot;${s5nz[$d]:-0}&quot; &quot;${s1nz[$d]:-0}&quot; \
&quot;${s5z[$d]:-0}&quot; &quot;${s1z[$d]:-0}&quot; \
&quot;${f5nz[$d]:-0}&quot; &quot;${f1nz[$d]:-0}&quot; \
&quot;${f5z[$d]:-0}&quot; &quot;${f1z[$d]:-0}&quot;
done | sort
</code></pre>
</body>

</details>

SDpower · 5 months ago

I've done extensive analysis on this exact problem. Using ccusage_go (open-source Claude Code usage tracker), I found that Cache Read tokens consumed 97.7% of my session costs — API actual cost was $1.47, total billed cost was $64.98 (a 44x markup). Cache also degrades instruction following in long sessions, which I documented with per-turn JSONL analysis.
Full write-up with data, community issue references, and Claude Code's own self-analysis report:
https://blog.sd.idv.tw/en/posts/2026-03-25_claude-code-cache-trap/
Tool: https://github.com/SDpower/ccusage_go

ArkNill · 5 months ago

The undisclosed quota changes continue. Max 20 ($200/mo), v2.1.89, April 1 2026: 100% in ~70 min after reset. No transparency from Anthropic on what changed.

Support acknowledged "resolved incidents" on Mar 31-Apr 1 but offered no explanation or fix — just "enable extra usage" (pay more).

Full report: #41788
Related: #38335, #38239, #40790, #41055, #41663, #40903

ArkNill · 5 months ago

I've been experiencing the same issue on Max 20 ($200/mo) — rate limit 100% exhausted in ~70 minutes.

After setting up a monitoring proxy using the official ANTHROPIC_BASE_URL env var, I identified two cache bugs as the root cause (#40524, #34629) and measured the impact: cache read ratio dropped to 4.3%, meaning ~20x token inflation per turn. After applying workarounds it stabilized at 89-99%.

Full analysis with per-request measured data, safe workarounds, and community references (including cc-cache-fix): https://github.com/ArkNill/claude-code-cache-analysis

ArkNill · 5 months ago

Update (April 2): v2.1.90 has significantly improved cache efficiency — benchmark shows 95-99% cache read in stable sessions (both npm and standalone installations).

If you're still affected:

  1. Update: claude update (or npm install -g @anthropic-ai/claude-code)
  2. Pin the version: add "DISABLE_AUTOUPDATER": "1" to ~/.claude/settings.json env section
  3. Avoid --resume (still broken)

Note: server-side quota issues (org-level pool sharing, accounting mismatches) remain unresolved — the above fixes the client-side cache drain only.

Benchmark data: https://github.com/ArkNill/claude-code-cache-analysis

ArkNill · 5 months ago

April 3 update: v2.1.91 fixes the cache regression that caused the worst drain. If you are still hitting limits after updating, there are additional unfixed mechanisms: a 200K tool result budget cap, a client-side false rate limiter, and silent context stripping — all confirmed via proxy testing. Anthropic acknowledged peak-hour tightening on X (Lydia Hallie) but stated "none were over-charging you." Measured data and analysis: claude-code-cache-analysis

Chenjanyuan · 2 months ago

Same issue here. Max 20x subscriber.

Since the May 27–28 server-side change that merged Claude Design's previously separate quota pool into the shared Claude.ai + Claude Code allowance, my Cowork usage is visibly eating into my main quota.

This week (June 1–4, 2026) I'm running multiple parallel Cowork sessions (Claude Design agents) and quota burn rate feels 2–3x what it was two weeks ago. The extra capacity I used to have via Cowork's separate budget is gone.

I subscribed to Max 20x specifically because Cowork had its own meter — that's what made parallel agent workflows viable on this plan. The silent merge undermines that value proposition.

Asking for one of:

  1. Restore Claude Design (Cowork) as a separately metered budget for Max 20x users, OR
  2. 2. Officially announce the merge with impact analysis for heavy Cowork users, OR
  3. 3. Increase Max 20x weekly limits proportionally to compensate.

This is a material change to the Max plan that wasn't disclosed at purchase. Please give us back the separate Cowork budget. Thanks for transparent communication going forward.