Token repetition loop on Opus 4.8 at >400k context — did not occur on Opus 4.7 under the same workload
On claude-opus-4-8, our long-running agentic sessions repeatedly collapsed into a token repetition loop: the fixed token court repeated up to ~32k times in a single message, replacing an intended tool call. I went through our raw transcripts (.jsonl) and extracted occurrence data: 37 burst events across 3 sessions, all at ≥434k input tokens — and zero occurrences on claude-opus-4-7 (17,525 turns) or claude-fable-5 (2,253 turns) under the same workload.
Related: #65823, #68740 (same symptom). Filing separately because this report includes cross-model control data; happy to have it merged into #65823 if maintainers prefer.
Environment
- Claude Code (desktop app, local agent mode), macOS (Darwin 25.x), v2.1.197 at the time of all bursts
- Model:
claude-opus-4-8with large context (up to ~1M) - Japanese-language, tool-heavy agentic coding sessions against a large monorepo (hundreds of tool calls per session)
Occurrence summary (from raw transcripts)
| Session | Burst events | Max repeats in one message | Total repeats | Context size at bursts (input incl. cache) |
|---|---|---|---|---|
| A (Jul 1–3) | 7 | 31,876 | ~188,800 | 434k – 763k |
| B (Jul 3–4, continuation of A) | 25 | 4,721 | ~10,800 | 558k – 715k |
| C (long-lived, Jun 19–Jul 10) | 5 | 51 | ~150 | 872k – 938k |
Key observations
- Every burst occurred at ≥ ~434k input tokens — but context size alone is not sufficient. I bucketed all assistant turns across our 16 largest transcripts by context size (exposure-adjusted, so this is a per-turn rate, not just "bursts happen late in sessions"):
| Context (input incl. cache) | Opus 4.8 turns | bursts | rate |
|---|---|---|---|
| < 400k | 11,057 | 0 | 0.00% |
| 400k–500k | 3,307 | 2 | 0.06% |
| 500k–700k | 6,596 | 23 | 0.35% |
| 700k–1M | 8,346 | 12 | 0.14% |
Below 400k: zero bursts in 11k turns. Above 400k it fires probabilistically at ~0.1–0.4% per turn. Independent first onsets in the three affected sessions were at 434k / 558k / 872k. (Caveat: the 37 bursts include lock-in cascades, so they are not fully independent events.)
- Model correlation — this looks like a regression in Opus 4.8, not an inherent large-context failure. In the same environment and workload:
claude-opus-4-7produced zero bursts across 17,525 turns, including 1,087 turns at 900k–1M context;claude-fable-5produced zero across 2,253 turns (111 of them interleaved in the same contaminated context as session B). All 37 bursts were onclaude-opus-4-8. - The onset zone is unreachable on the standard 200k window. All our occurrences (and the ranges reported in #65823) sit beyond 200k context.
- Trigger shape: every burst begins immediately after a short conversational preamble that announces an upcoming action and ends right where a
tool_useblock should start (e.g. "…I'll fetch the new digest and re-run.\n\n" →court court court…). The repetition appears to replace the tool call. This matches the workaround described in #68740 ("execute tool calls without preamble text"). - Lock-in / escalation: in session A, once the first burst occurred, the next 6 consecutive turns each ran away to ~31.3k–31.9k repeats — i.e. pinned at what looks like the 32k output-token cap, every turn, until we killed the session. Total ~189k wasted output tokens; the transcript grew to 24MB and the session could no longer be opened in the UI (we had to do manual transcript surgery to recover the work). This suggests the contaminated history itself sharply raises recurrence probability.
- Partial self-recovery exists: session B shows many bursts of exactly ~60 repeats after which the model recovers within the same message and proceeds — so there seem to be two regimes (transient blip vs. full runaway to the output cap).
- Session duration itself did not separate burst from non-burst turns in our data — context size did (see table in 1). Session C ran for three weeks and only burst once its context exceeded 872k.
- Recovery (observed): abandoning the contaminated session and starting a fresh one (with a handoff note) ended the cascade every time we did it, and the same task then proceeded normally. It is not permanent immunity, though: session B, itself a fresh continuation of session A, burst again once its own context regrew past ~558k.
- The exact repeated unit is
courtas a standalone paragraph —court\n\ncourt\n\ncourt…(blank-line separated), not space-separated. Onset is always immediately after a paragraph break (…。\n\nin Japanese prose). stop_reasonseparates the regimes: all five ~31k runaways in session A ended withstop_reason: "max_tokens"; every transient burst that recovered ended withstop_reason: "tool_use"— i.e. the intended tool call was eventually emitted in the same turn. (Some mid-size bursts in session B showstop_reason: nullin our transcript — possibly mid-stream captures; the raw values are in the table below.)- The model notices mid-burst and still cannot exit the loop. In sessions B and C, the model interrupted its own repetition with a brief apology for the extraneous output — in one case even explicitly stating it would emit tool calls with no preamble from then on — and then immediately resumed the repetition. Self-awareness of the corruption does not break the attractor.
Server-side lookup keys — in case it is easier to pull the inference traces internally than to work from our transcripts, here are the request/message IDs of all 37 burst turns:
<details>
<summary>All 37 burst turns: timestamp (UTC), session, repeats, stop_reason, context tokens, message id, request id</summary>
| Timestamp (UTC) | Session | Repeats | stop_reason | Context | Message ID | Request ID |
|---|---|---|---|---|---|---|
| 2026-07-03T14:52:07 | A | 20 | tool_use | 433,939 | msg_01ErQq5mLyiHXAXJUuUNw165 | req_011CcfK361dbcTqgRmx6L7Jh |
| 2026-07-03T15:03:19 | A | 31,342 | max_tokens | 439,580 | msg_01SbvvMZwNQmiy1GM2wmAQfh | req_011CcfKKAPWJAQfNkPFf5AV3 |
| 2026-07-03T15:11:47 | A | 31,784 | max_tokens | 503,644 | msg_013RMjwJXHPiD55zxBYKD2qc | req_011CcfKy1GdbrSxKkE2rgKnv |
| 2026-07-03T15:20:32 | A | 31,876 | max_tokens | 568,592 | msg_01AGfFceEU49VA51kKgxokqV | req_011CcfLdtEoHV5VMjFtPVcNN |
| 2026-07-03T15:30:00 | A | 31,668 | max_tokens | 633,757 | msg_01MCHenPMKht881xDCTUYSmj | req_011CcfMMbAJm2PBvBDXkn1ZC |
| 2026-07-03T15:39:14 | A | 31,786 | max_tokens | 698,118 | msg_01E6BnuJ7BAgzZMx66YzYQDH | req_011CcfN4JFiL4Kbh1Z57ykAd |
| 2026-07-03T15:49:07 | A | 30,326 | max_tokens | 762,909 | msg_01XUtSpSXxHPbJHHJDsUWPzA | req_011CcfNne5jqxvu1m6mrKNTc |
| 2026-07-04T08:56:40 | B | 30 | tool_use | 557,761 | msg_0165MyGtMSjX1FyJZ246Qt9M | req_011CcgjkJJFwAyzsjKBNtRNz |
| 2026-07-04T08:58:31 | B | 4,721 | None | 561,056 | msg_01PAPgZ2ENacczpN45k3XFF3 | req_011CcgjqEKn5LYeuLJc6AeHi |
| 2026-07-04T09:08:00 | B | 65 | tool_use | 582,464 | msg_01PF91nKU7Eee7r8wMKTrGsu | req_011CcgkfSeH7xfSkLYsydn3u |
| 2026-07-04T09:08:44 | B | 261 | None | 583,404 | msg_0158NmNJdq4XAe9gCJFSm7md | req_011CcgkifvZuWCokL3boJz1p |
| 2026-07-04T09:09:02 | B | 8 | tool_use | 584,200 | msg_01NH6V8sYhmLiasCLTmMWfpb | req_011CcgkkGp59ZKyV8wiwNyiR |
| 2026-07-04T09:10:14 | B | 3,689 | None | 585,197 | msg_01E9i7w1C9RMvPmUmzKNigE8 | req_011CcgkmiFoYwe94PszmDpJe |
| 2026-07-04T09:10:44 | B | 61 | tool_use | 592,914 | msg_01Cc1d9nzWVWjvRJ5fWqXdNd | req_011CcgksttRsTr9wvf4o6v6L |
| 2026-07-04T09:22:22 | B | 60 | None | 595,733 | msg_01MWyVBhN2aFrG95tNg3bSgH | req_011CcgmkrXUT2z8cK72WiCo9 |
| 2026-07-04T09:35:45 | B | 60 | None | 611,385 | msg_014xVbF439jx3zqz5voDV3EH | req_011CcgnmJMMZYyoTwKNjZ6ug |
| 2026-07-04T13:12:33 | B | 793 | None | 620,290 | msg_01LdQSxKzQ7QHDh5V8FDAHRv | req_011Cch5FHnGeD9HzawoLr7os |
| 2026-07-04T13:25:19 | B | 151 | None | 642,375 | msg_01GaUWQTx2GWpUrvfUquBhjq | req_011Cch6HsM3KgM7hqHqxLgPk |
| 2026-07-04T13:32:25 | B | 60 | None | 646,080 | msg_01YHcGB1ME53zzFQNeJZ8rrt | req_011Cch6qShc5hPFg6q8DVfEu |
| 2026-07-04T14:28:25 | B | 60 | end_turn | 650,316 | msg_018YEqpXaQLpjhEGe5kXR3wt | req_011CchB6v1oWUi2ksDpui4wL |
| 2026-07-04T14:48:25 | B | 60 | None | 669,061 | msg_018aeZXgPnZLcFMD6zhs2Jrj | req_011CchCci3Ry4g3CHJqfwbCM |
| 2026-07-04T14:51:40 | B | 60 | tool_use | 680,950 | msg_01HiysSDLmHzpnNG5dAjMqcY | req_011CchCqezjgffjVP9JD5Xdt |
| 2026-07-04T15:17:40 | C | 11 | tool_use | 872,446 | msg_01EwCTFdvVw7mLUGx3R2ZaQC | req_011CchEpJ1dZmYCFwv85NvUL |
| 2026-07-04T15:32:09 | B | 60 | tool_use | 691,782 | msg_018FQtANwxhkW9tsfcADFmR5 | req_011CchFxJENCggytKn77GoS2 |
| 2026-07-04T15:32:59 | B | 60 | tool_use | 693,608 | msg_01TRLDWbLm8LK4EDu9i8cb7m | req_011CchG1V7Hs9JjEUgJggXN2 |
| 2026-07-04T15:33:43 | B | 60 | tool_use | 695,578 | msg_01TrnSKHdY5TxzTTGtvxYyxZ | req_011CchG3o6qCxaLAfq6rJ1iT |
| 2026-07-04T15:34:28 | B | 60 | tool_use | 699,285 | msg_01WPMk7mi41yxAzbkSnSnP5k | req_011CchG6hbWzTJaxiU9bQN38 |
| 2026-07-04T15:34:42 | B | 60 | tool_use | 702,420 | msg_016K9iR8fozekzkiHLcs1ofp | req_011CchGA44f7u2kFC2VMrT6f |
| 2026-07-04T15:35:42 | B | 60 | tool_use | 704,862 | msg_01KsnDKzFpm9qP1ZmFTxhWCY | req_011CchGC6PcKc3tCNNhWdtRi |
| 2026-07-04T15:36:48 | B | 60 | tool_use | 708,851 | msg_01ESSgnbZ44Rk4WFuauxxE1X | req_011CchGHBBEX4Finiy7EhnBR |
| 2026-07-04T15:42:32 | B | 142 | None | 712,758 | msg_01HCqXmRJbsJWwJj4eEA6kva | req_011CchGkN7of15pUi6qt59gp |
| 2026-07-04T15:42:41 | B | 61 | tool_use | 713,524 | msg_01Xxd8o3z2VmMbZMqmD9wgAC | req_011CchGmeBiteyzjUUquVZeu |
| 2026-07-04T15:43:04 | B | 60 | None | 714,492 | msg_01D2LPJFmHAXcU8FXgdCBToF | req_011CchGnrhKvtqapXvRzKi4E |
| 2026-07-05T02:34:32 | C | 13 | tool_use | 913,346 | msg_01FofSAdnsh2ScamTeqqppTB | req_011Cci8URsQHEgbpBAiXY6Fh |
| 2026-07-05T02:53:17 | C | 37 | tool_use | 920,618 | msg_01RVzavWAXG6QMmm9oXvs1bU | req_011Cci9uHnJSdSLaCFzLuoLd |
| 2026-07-08T11:14:31 | C | 41 | tool_use | 928,960 | msg_011CcpVYjoErCXPtUPj7rPZa | req_011CcpVYZ1biXPuDD3bBSNpz |
| 2026-07-08T11:15:27 | C | 51 | None | 937,600 | msg_011CcpVbR5wCpQ57T6Y5c5Ek | req_011CcpVbErnEH7mbhN3TDvab |
</details>
Repro suggestion based on the above: drive a Japanese-language agentic session on claude-opus-4-8 past ~450k input tokens with high-frequency tool calls, and have the model consistently write a 1–3 sentence announcement ending with "\n\n" immediately before each tool call. In our data that combination is what fires, probabilistically.
I have the full raw transcripts (including the 189k-repeat session) and can share sanitized excerpts — event sequences around burst onset with usage counters — if that helps the investigation.