[BUG] Transcript loader silently discards all but one conversation branch on resume, losing delivered assistant messages

Status Open
Reported on v2.1.98
Maintainer reply None cached
Activity 1 comment · opened Aug 21, 2026

Preflight Checklist

  • [x] I have searched existing issues and this hasn't been reported yet
  • [x] This is a single bug report (please file separate reports for different bugs)
  • [x] I am using the latest version of Claude Code

What's Wrong?

RELATED BUT DISTINCT — please read this before closing as a duplicate.

There is a cluster of reports about lost conversation history. This is not one of
them, and the difference is mechanical rather than a matter of degree.

  • #21617, #83094: a parentUuid referencing a uuid that exists NOWHERE in the file.

A dangling pointer; everything before the break is orphaned. In the bug below
nothing dangles — the parent exists, and it has TWO children. The loader detects
this (J.size > 1) and deliberately discards one of them.

  • #83050: a render bug in the agents view, where claude --resume is the WORKAROUND.

Here resume is where the loss occurs.

  • #76232: EnterWorktree physically relocates the transcript file and mints a new

session. No file is moved here.

  • #86075: large tool results replaced by placeholders. Same family (the harness

losing content while reporting success), different layer.

What is new in this report: the exact code path that discards the branch, a version
boundary for when it started, and the fact that the codebase already contains both
the flag (keepAllLeaves) and the counter (J.size) needed to fix or detect it.

When a session's parentUuid chain forks, resuming that session silently discards
every branch but one. Delivered assistant messages vanish from the UI transcript and
from the model's reconstructed history. No warning, no telemetry, no error. The
content is still intact on disk.

The user-visible symptom is worse than "a message is missing": the model has no
record of work it did, so if you ask about it, it reconstructs. In my case Claude
told me a second window must have been editing my files, because from where it was
standing that was the only explanation for changes it had no memory of making.

Across a 214-transcript archive I found 22 orphaned branches holding delivered
assistant messages, in 14 sessions, ~81,000 characters, the earliest 24 July 2026.

WHAT PRODUCES THE FORK

Two records end up sharing a parent:

T+0ms assistant the delivered reply
T+220ms system stop_hook_summary

The system record's ancestry runs back through an api_error record to a checkpoint
written before the error, so after an API error the harness keeps writing system
records from a stale parent. Both records are now children of the same uuid.

19 of my 22 cases have an api_error near the fork. The other 3 have
stop_hook_summary records and no api_error at all, so the error is sufficient but
not necessary.

insertMessageChain takes the parent as a caller-supplied argument
(let u = n ?? null) and never validates it against the current chain head, so a
stale value is written without complaint.

WHY THE BRANCH IS THEN LOST — the actual defect

The transcript loader collects every leaf, walks each back to its nearest
user/assistant ancestor, and then:

if (!t?.keepAllLeaves && J.size > 1) { // J = the set of leaves
let Ce = Z && J.has(Z) ? Z : B; // B = last non-sidechain record IN FILE ORDER
...
J.clear(); J.add(Ie.uuid); // discard the rest
}
return fe(J);

J.size > 1 means the loader KNOWS there is more than one live branch. It keeps the
one under B, and B is assigned while streaming the file as
if (!J.isSidechain) B = J.uuid — whichever record appears last. So a
stop_hook_summary written 220ms after an assistant reply wins the conversation
purely on file position.

There is telemetry for tengu_transcript_parent_cycle, tengu_chain_parent_cycle,
tengu_phantom_parent_write, tengu_chain_self_reference_write,
tengu_relink_walk_broken and tengu_resume_unchained_transcript — six adjacent
conditions — and none at all for "N branches of conversation were just discarded".

What Should Happen?

Any of these three would be enough, and they are not mutually exclusive.

  1. NOTHING SHOULD BE DISCARDED SILENTLY. At minimum, emit a telemetry event when

J.size > 1 collapses. It is one line next to code that already computes the
number. Right now this failure is invisible to both the user and to you.

  1. keepAllLeaves ALREADY EXISTS AND IS NOT USED ON THIS PATH. The code that

enumerates sessions for the resume picker passes keepAllLeaves: true. The code
that loads a session for resume does not. The same file preserves every branch
when listing and collapses to one when loading. If a session has two live
branches, resume should keep the conversation, and prefer the branch containing
user/assistant messages over one containing only system records.

  1. THE FORK SHOULD NOT BE CREATED. insertMessageChain trusts its caller's parent

uuid. Validating that the supplied parent is the current head — or that a fork is
intentional — stops this at the write rather than tidying it up afterwards.

Concretely, for the shape above: the assistant reply and the stop_hook_summary
should not both be children of the same uuid; and if they somehow are, resume should
not choose the system record's branch and throw away the reply.

Error Messages/Logs

There is no error message. That is the substance of the report — the failure is
completely silent in both directions, and nothing in the product can currently
observe it.

What can be shown instead:

1. THE FORK, as it appears in a session .jsonl. Two records, same parentUuid,
   220ms apart, the later one winning (uuids redacted, structure verbatim):

   {"type":"assistant","uuid":"<A>","parentUuid":"<P>",
    "timestamp":"...T23:44:08.029Z", ...}          <- the delivered reply
   {"type":"system","subtype":"stop_hook_summary","uuid":"<S>","parentUuid":"<P>",
    "timestamp":"...T23:44:08.249Z", ...}          <- written 220ms later

   The next user message parents to <S>. <A>'s branch is now unreachable from the
   last record and is dropped on the next resume.

2. THE BEHAVIOUR, DOCUMENTED IN YOUR OWN WARNING STRING. Present in the shipped
   binary, aimed at a different audience:

   "Conversation reconstruction walks parentUuid from the last record, so unlinked
    records are dropped — the file's producer must chain records (parentUuid null on
    the first, the previous record's uuid on each subsequent one)."

3. NOTHING ELSE FLAGS IT. `isSidechain` and `isMeta` are false on every record
   involved. Reachability from the last record is the only signal that anything
   happened.

Steps to Reproduce

The organic path needs a real mid-turn API error, so it is not deterministic. A
constructed transcript reproduces the loader behaviour on its own, and is what I
would start from.

A — MINIMAL CONSTRUCTED CASE (deterministic)

  1. Create a .jsonl in a project's session directory with six records. Two of them —

the assistant reply and a stop_hook_summary — share the parent a1, with the
system record written afterwards:

{"uuid":"u1","parentUuid":null,"type":"user","timestamp":"2026-08-21T12:00:00.000Z","message":{"role":"user","content":[{"type":"text","text":"hello"}]}}
{"uuid":"a1","parentUuid":"u1","type":"assistant","timestamp":"2026-08-21T12:00:01.000Z","message":{"role":"assistant","content":[{"type":"text","text":"anchor"}]}}
{"uuid":"b1","parentUuid":"a1","type":"assistant","timestamp":"2026-08-21T12:00:02.000Z","message":{"role":"assistant","content":[{"type":"text","text":"THE REPLY THAT DISAPPEARS"}]}}
{"uuid":"s1","parentUuid":"a1","type":"system","subtype":"stop_hook_summary","timestamp":"2026-08-21T12:00:02.220Z"}
{"uuid":"u2","parentUuid":"s1","type":"user","timestamp":"2026-08-21T12:00:30.000Z","message":{"role":"user","content":[{"type":"text","text":"next"}]}}
{"uuid":"a2","parentUuid":"u2","type":"assistant","timestamp":"2026-08-21T12:00:31.000Z","message":{"role":"assistant","content":[{"type":"text","text":"carrying on"}]}}

  1. Resume that session.
  1. OBSERVED: "THE REPLY THAT DISAPPEARS" is not in the transcript and the model has

no knowledge of it. The record is still in the file.
EXPECTED: it is present, or at minimum something says it was dropped.

  1. Verify independently without resuming — walk parentUuid from the last record

upward and collect what you reach. b1 is not on that walk. That walk is the
reconstruction your own warning string describes.

B — THE ORGANIC PATH, for context on how it actually arises

  1. Run a long interactive session with at least one Stop hook configured.
  2. Have a turn hit an API error mid-flight (api_error record appears).
  3. Let Claude complete a substantial reply after the error.
  4. Send another message, end the session, resume it.
  5. The reply from step 3 is gone from the transcript and from the model's history.

C — HOW I FOUND THE HISTORICAL CASES

A branch is orphaned iff it is unreachable by walking parentUuid from the last
record. Two caveats for anyone writing the same check:

  • EXCLUDE everything before the last compact_boundary. Compaction legitimately

abandons that branch, and a scanner without this exclusion reports normal
behaviour as corruption.

  • COUNT ONLY branches containing an assistant message. In my archive 601 of 623

orphaned branches held nothing but stray single records; 22 held delivered work.
The raw number is alarming and meaningless.

Claude Model

Not sure / Multiple models

Is this a regression?

Yes, this worked in a previous version

Last Working Version

2.1.98 — 1,044 user turns with 865 stop_hook_summary records and zero orphaned branches. First version showing it: 2.1.219. The api_error record type does not exist before 2.1.219: version api_error records first seen 2.1.98 0 never 2.1.219 4 2026-07-24 2.1.220 84 2026-07-31 My earliest orphaned branch is 2026-07-24 — the same day the record type first appears. 2.1.219 also introduced turn_duration, away_summary, local_command, informational and model_refusal_fallback, none of which exist under 2.1.98. More system record types written around a turn boundary means more candidates for the parent slot. Caveat, stated honestly: I have no data between 2.1.98 and 2.1.215, so the introduction point cannot be narrowed further from my archive. 2.1.215–2.1.218 are also clean but total only 196 turns, which is too small to be evidence of anything.

Claude Code Version

2.1.238 (Claude Code) — VS Code extension, which is where every measurement came from Also observed on 2.1.219, 2.1.220, 2.1.223, 2.1.226, 2.1.228, 2.1.233 and 2.1.235.

Platform

Anthropic API

Operating System

Windows

Terminal/Shell

VS Code integrated terminal

Additional Information

RATE, and it is not just "it happens sometimes"

Orphaned branches per user turn, corrected for turns behind a compaction boundary
(those can never show an orphan, so they do not belong in the denominator):

orphans per 1k visible turns
before 12 5.0
after 7 68.6

n=19 across the whole archive, so please read that as a direction and not a
statistic. I am not claiming significance.

COMPACTION DESTROYS THE EVIDENCE — this matters if you go looking yourselves

Records behind a compact_boundary are unreachable by design, so previously
detected orphans become indistinguishable from ordinary compaction. In my archive a
single compaction dropped the detected count from 22 to 19 with nothing fixed and
nothing newly lost. Any historical count of this bug is a decaying lower bound, and
the older the history the cleaner it will look for reasons that have nothing to do
with how often the bug fires. 3,013 of my 5,516 archived turns are already blind.

IT IS REPAIRABLE OFFLINE, WHICH ALSO CONFIRMS THE DIAGNOSIS

Repointing the winning sibling's parentUuid to the tail of the orphaned branch —
one field, one record, per branch — puts the branch back on the walk. Verified on a
real transcript:

records on the resume walk the lost 2,720-char message
original 3,290 of 4,126 not on it
repaired 3,441 of 4,126 on it

Zero records added, zero removed, zero altered in any other way. It works precisely
because the loader insists on a single leaf: give it one leaf and there is nothing
left to discard. I mention it because a fix that only has to make J.size == 1 true
is a small fix.

WHY THIS IS WORSE THAN A MISSING MESSAGE

The model is reconstructed from the same walk as the UI, so it does not know work is
missing — it knows it has no record of work that visibly happened. Mine concluded
another process must have edited the files. That is a confident, wrong, entirely
harness-caused account of the user's own project, and there is currently nothing
anywhere that would let either party notice.

View original on GitHub ↗

This issue has 1 comment on GitHub. Read the full discussion on GitHub ↗