[FEATURE] Per-model prompt-cache TTL — extend promptCacheTtl to a per-model map

Status Open
Maintainer reply None cached
Activity 0 comments · opened Aug 25, 2026

Preflight Checklist

  • [x] I have searched existing requests and this feature hasn't been requested yet
  • [x] This is a single feature request (not multiple features)

Problem Statement

We run Claude Code for an enterprise team on Amazon Bedrock (token billing, self-hosted gateway).
Several developers exhausted their monthly budget in ~10 days. A representative session's cost split:

| Component | Tokens | Cost | Share |
|-------------|--------|-------|-------|
| Input | 59.6k | $0.30 | 5% |
| Output | 65.7k | $1.64 | 27% |
| Cache read | 3.1M | $1.55 | 26% |
| Cache write | 408.4k | $2.55 | 42% |
| Total | | $6.04 | |

Caching itself works well (each write is read back ~7.6x on average). The problem is TTL fit:
interactive sessions routinely have >5-minute think-time gaps, so the 5m cache (the default on
Bedrock) expires and unchanged context is rewritten at 1.25x the input rate.

promptCacheTtl / subagentPromptCacheTtl (v2.1.242+) split TTL by request kind, which helps.
But the right TTL also depends on the model, because the write premium is priced per model:
on Opus 4.8, a 5m write is $6.25/MTok and a 1h write is $10/MTok against a $0.50/MTok read.
The break-even between "pay 2x once" and "pay 1.25x on every expiry" lands differently for an
expensive model carrying long interactive sessions than for a cheap model doing short-lived
background work — and the main/subagent split can't express that when the same bucket runs
different models (mid-session /model switches, background tasks on cheaper models, or a gateway
routing different user groups to different models).

A single global value forces a compromise: 1h everywhere overpays on short-lived work; 5m
everywhere burns full-context rewrites on interactive sessions.

Proposed Solution

Let the existing settings accept a per-model map, keeping the current string form as shorthand:

{
"promptCacheTtl": {
"default": "1h",
"claude-haiku-4-5": "5m"
},
"subagentPromptCacheTtl": "5m"
}

Longer term we'd prefer the platform to infer TTL from workload shape (short-lived subagent
requests at 5m, interactive sessions at 1h) — #56307 asked for that and was closed not-planned,
so per-model configuration is the practical next step operators can act on.

Related: #74075 (shipped the configurable TTL), #47425 (intermediate TTL tiers — complementary),
#62217 / #60316 (superseded by the shipped settings).

Alternative Solutions

_No response_

Priority

High - Significant impact on productivity

Feature Category

Performance and speed

Use Case Example

_No response_

Additional Context

_No response_

View original on GitHub ↗