[BUG] Prompt cache for resumed session history intermittently fails to hit on `resume`, independent of idle time or TTL
Preflight Checklist
- [x] I have searched existing issues and this hasn't been reported yet
- [x] This is a single bug report (please file separate reports for different bugs)
- [x] I am using the latest version of Claude Code
What's Wrong?
Bug Report: Claude Agent SDK — Prompt cache for resumed session history intermittently fails to hit on resume, independent of idle time or TTL
Summary
When using the Claude Agent SDK's resume option (equivalent to CLI --resume <sessionId>) to continue a session across separate query() invocations, the cached conversation-history portion of the prompt intermittently fails to hit cache and gets fully rewritten as a new cache_creation — even when the resume happens seconds after the prior turn ended, well inside the prompt-cache TTL, and even when a separate resume of the same session minutes or tens of minutes later hits cache successfully. The static system-prompt/tool-schema portion of the same request consistently hits cache in every case, including the cases where the conversation-history portion misses — indicating the failure is isolated to the cache breakpoint covering prior turns, not the request as a whole.
Environment
- Claude Agent SDK (Node/TypeScript), package
@anthropic-ai/claude-agent-sdk-win32-x64, version0.3.261 - OS: Windows 11
- Model:
claude-fable-5-1[1m](1M-context variant) - Account has 1-hour extended prompt caching enabled
- Session resumed via the SDK's
resume: <sessionId>option on each newquery()call, with only the new incremental user message sent as prompt input — full conversation history is expected to be reconstructed by the CLI from the on-disk session file under~/.claude/projects/.../<sessionId>.jsonl
Actual behavior
The same "resume shortly after last turn" scenario intermittently and unpredictably falls back to writing the entire conversation history as a fresh cache_creation, at roughly 10x the price of a cache read, with no discernible trigger tied to elapsed time, message content, or turn structure.
Reproduction data (same session, times UTC)
| Resume time | Elapsed since prior turn ended | Result |
|---|---|---|
| 18:05:53 | 55 min | cache hit (98%) |
| 18:09:37 | 3.4 min | cache miss (6%) |
| 18:13:45 | 3.0 min | cache hit (97%) |
| 18:27:15 | 12.4 min | cache miss (5%) |
| 18:42:47 | ~0 (immediate) | cache hit (99%) |
| 18:43:23 | ~0 (immediate) | cache hit (100%) |
| 19:06:32 | 23 min | cache miss (6%) |
Key observations from this data:
- Elapsed time does not predict the outcome. A 55-minute-old resume hit cache cleanly, while a 3.4-minute-old resume immediately after it missed. This rules out any simple TTL-expiry explanation (extended 1-hour caching was active throughout).
- On every miss,
cache_read_input_tokenscollapses to the same small fixed value (~6,459 tokens) — consistent with only the static system-prompt/tool-schema portion of the request surviving, while the entire conversation-history portion (in one case >170,000 tokens) is rewritten from scratch. - This happens on the single most common turn-completion path — an ordinary next-user-turn resume after Claude's previous reply fully completed, not any unusual state (no aborted queries, no error recovery, no multi-session merge involved in these events).
- Four additional independent reset events were captured in a separate long-running session with the same signature: full cache rewrite of the historical portion on a
resumecall, with gaps between the prior turn and the resume of 6.14, 20.59, 6.29, and 3.67 minutes respectively — again showing no consistent relationship to elapsed time.
Additional notes
- The intermittency (same code path, same session, minutes apart, alternating hit/miss) suggests this is not deterministically tied to any specific message content or turn shape — I compared the JSON of the assistant messages immediately preceding several hit and several miss events (including plain single-block text replies and multi-step tool-use turns) and found no consistent structural difference between the two groups.
- This may be related to interaction between prompt caching and the 1M-context (
[1m]) beta variant specifically, since all captured examples used a 1M-context model; it has not been verified against the same conversation on a standard 200K-context model. - Financial impact is significant for any long-running session that relies on
resumeacross many short-livedquery()calls, since each unpredictable miss re-bills the entire accumulated conversation history at full cache-write price instead of the ~10x cheaper cache-read price.
Ask
Please investigate whether the resume/--resume code path has a known race or eviction condition affecting the conversation-history cache breakpoint specifically (as opposed to the system-prompt/tool-schema breakpoint, which was never observed to fail). If this is already addressed in a later CLI release, confirming the fix version would be appreciated so we can validate against these reproduction steps.
What Should Happen?
Expected behavior
A resume of an existing session, made shortly after the prior turn ended, should hit the previously cached conversation-history prefix (cache_read_input_tokens close to the full prior context size, cache_creation_input_tokens covering only the newly added turn) — consistent with what already happens for the majority of resumes in the same session.
Error Messages/Logs
No error is thrown — the request completes successfully with stop_reason: end_turn. The only observable symptom is in the response's usage block, where cache_read_input_tokens drops to a small fixed value and cache_creation_input_tokens covers the full prior conversation instead of just the new turn.
### Reproduction data (same session, times UTC)
| Resume time | Elapsed since prior turn ended | Result |
|---|---|---|
| 18:05:53 | 55 min | cache **hit** (98%) |
| 18:09:37 | 3.4 min | cache **miss** (6%) |
| 18:13:45 | 3.0 min | cache **hit** (97%) |
| 18:27:15 | 12.4 min | cache **miss** (5%) |
| 18:42:47 | ~0 (immediate) | cache **hit** (99%) |
| 18:43:23 | ~0 (immediate) | cache **hit** (100%) |
| 19:06:32 | 23 min | cache **miss** (6%) |
Steps to Reproduce
Steps to reproduce
- Start a session and drive several turns of normal conversation (mix of plain text replies and tool-use turns), so a large prompt-cache checkpoint builds up over the conversation history.
- Let the query/process for a turn close normally (turn ends with
stop_reason: end_turn, no pending tool calls). - Start a new
query()call withresume: <same sessionId>, sending only the next user message as prompt input (not replaying history — relying on--resumeto load it). - Inspect the
usageblock of the resulting response:cache_read_input_tokensvscache_creation_input_tokens. - Repeat step 2–4 across many turns in the same long-running session and log the result of each resume.
Claude Model
Not sure / Multiple models
Is this a regression?
I don't know
Last Working Version
_No response_
Claude Code Version
2.1.265 (@ time of submission)
Platform
Anthropic API
Operating System
Windows
Terminal/Shell
Other
Additional Information
_No response_