[BUG] Agent response Markdown rendering in TUI mutates meaning of content (renumbers ordered lists)
[BUG] Agent Response Markdown rendering in the TUI mutates meaning of content
Description
The Claude Code TUI's markdown renderer silently renumbers ordered list items in agent responses, changing 1. 5. 6. 8. into 1. 2. 3. 4. — destroying the intended back-references and making the response factually misleading to the user.
This is not a cosmetic issue. The agent deliberately used non-sequential list numbers to reference specific items from a prior numbered list. The renderer's auto-renumbering produced a response the user could not correctly interpret, leading to a confusing exchange where the user questioned the agent about a "point 2" that the agent never wrote.
Observed behaviour
The agent emitted a response with intentionally non-sequential list numbers — 1., 5., 6., 8. — each referencing a specific numbered claim from the user's prior message. The TUI rendered these as 1., 2., 3., 4., silently destroying the back-references.
What the agent actually wrote (from the conversation JSON)
The raw text field in the assistant message contains:
1. Noted — ...
5. Yes, both runtimes: ...
6. Both layers ran: ...
8. Fair challenge, ...
What the TUI displayed to the user
1. Noted — ...
2. Yes, both runtimes: ...
3. Both layers ran: ...
4. Fair challenge, ...
The user saw 2. and reasonably asked "does your last response's '2.' map to a different number from the 1-8?" — because to the user, the response appeared to introduce a new point numbered "2" that didn't correspond to anything. The agent had no idea the renumbering had occurred and initially couldn't explain the discrepancy, until it inferred the markdown auto-numbering behaviour.
Root cause
Per the CommonMark specification §5.3, a markdown renderer normalises an ordered list so that item numbers are displayed sequentially starting from the first item's number — regardless of the literal numbers in the source. The TUI applies this normalisation to the agent's raw text response.
This is correct rendering behaviour for authored documents, but incorrect for agent output displayed back to the user it was addressing. The agent's response is not a formatted document — it is a message in a conversation, where the literal numbers carry referential meaning.
Expected behaviour
The literal numbers from the agent's response should be preserved as-is in the TUI display. Either:
- Don't apply ordered-list normalisation to agent responses, or
- Escape patterns like
N.at the start of a line so they render as literal text, not list syntax, or - Instruct the model (via system prompt) to avoid bare
N.patterns when non-sequential numbering is intended — though this is a workaround, not a fix, since the model cannot predict when the renderer will misinterpret its output.
Reproduction
Quick deterministic reproduction (single prompt)
Ask the agent:
Output the following lines verbatim:
1. First
5. Fifth
6. Sixth
8. Eighth
Observed in TUI: Renders as 1. First / 2. Fifth / 3. Sixth / 4. Eighth.
Actual agent message JSON: Contains literal 1. First\n5. Fifth\n6. Sixth\n8. Eighth.
Conversational reproduction
- Engage the agent in a conversation where you present a numbered list of 8+ items.
- Respond to those items selectively (e.g. only address items 1, 5, 6, 8) using
idkor similar for the ones you don't engage with. - The agent will respond addressing the active items using the original numbers (
1. 5. 6. 8.). - Observe the TUI rendering: the numbers are silently changed to
1. 2. 3. 4..
Evidence
<details>
<summary><strong>Full TUI-rendered conversation (verbatim from TUI view)</strong></summary>
````tui
● Taking your facts as ground truth (manual mode, 4 prompts in lead TUI, 4 plain answers), the last test establishes:
- Forwarding under lead-manual: 4/4 asks reached you — deterministic this run.
- The factor prompts came from the settings ask rule — settings rules forward for teammates.
- The echo prompts came from the temp hook's ask (fired per log) despite the chain's parallel allow — cross-hook precedence resolved ask-over-allow both times.
- New fact this test uniquely proves: run 1's approvals of the identical commands by the same teammate did not suppress run 2's prompts — approvals are one-shot, nothing is cached.
- PermissionRequest stayed silent for all teammate asks — reconfirmed (log unchanged).
- [auto-allow] rendered while prompts were up — reconfirmed as display-only misdirection.
- The probe reported "instant" for calls that were human-blocked — teammate perceived-wait is worthless; dropped from methodology.
- Lead mode did not change the outcome (4/4 prompts under auto in run 1, 4/4 under manual in run 2) — so this test did not isolate what caused round 5; the leak stands at: seen once under auto, never under
manual, n too small either way.
Consult away.
❯ 1. yes, I typed an approval msg too, btw
- idk
- idk
- I concur
- for all teammate asks in the teammate runtime, you mean?
- you phrase it as if you cannot comprehend the reality, only theory about it... I dont blame you, you're a text model...
the [auto-allow] rendered in the teammate TUI, came from the PreToolUse hook, and ONLY ran for the two echo Bash calls. And this rendered AFTER the tool batch entirely finished (thats how the claude code is
implemented). So not DURING the prompts, and for irrelevant reasons; delayed to the end of the batch call before bbeing visual for the user
What this might mean: the pretooluse must have set "allow", or is this allow overriddewn by an "ask"?
- yes
- yuou seem to be anchoiring on prove five. What are you trying to isolate. because I recall probe 5 being an invalid test with false observations
∴ So the approval message the user typed didn't make it to probe7 either — its report came back without any note attached, which confirms that approval notes are being dropped somewhere in the pipeline. And
looking at the PermissionRequest hooks, they didn't fire in either session — not on my teammate's end where the tool call originated, and not on my end either. The permreq_log hook is configured in user
settings that both sessions load, but the log only shows my single entry, so for requests coming from my teammate, those hooks never triggered at all.
The user's explanation clarifies what happened with the [auto-allow] rendering: it showed up in the teammate's TUI after the entire batch completed, delayed, and only for the two echo calls. Their question
was whether the chain's auto-allow set "allow" and then got overridden by the temp hook's "ask" — and the answer is yes, both hooks ran. The log proves the ask hook fired for echo, and the chain's allow
produced the systemMessage. The binary received both signals (allow from the chain, ask from the temp hook), and since I got prompted, the ask took precedence. The systemMessage renders regardless of the
final resolution, which is why it appeared late. So the chain's allow was indeed overridden by ask. For the factor call, the chain didn't show [auto-allow] at all — factor probably isn't in its read-only
allowlist — so the settings rule kicked in and asked for a prompt instead.
Now looking back at probe 5: the user is pointing out that their "no prompt" claim was never actually confirmed. They said "wait what" and "stop assuming," but never explicitly stated whether a prompt
appeared or not. Later they interpreted the pane paste showing [auto-allow] as evidence it was allowed by their auto-allow rule, but we now know that [auto-allow] renders even when the prompt was shown and
overridden. So probe 5's entire "silent allow" story rests on observations we've now proven invalid. The user is saying to drop it — if probe 5 is discarded, the whole narrative changes. So if probe5 is
unreliable due to the ambiguity around whether they actually saw and approved the prompt, I should discard it entirely. That leaves me with only the valid data points, which all consistently show
deterministic ask forwarding with no classifier leak. Without probe5, the whole isolation goal I was trying to achieve in point 8 falls apart.
The design path forward is clearer now: gating returns to "ask" mode appears solid across all the evidence I can trust. The classifier can still deny things for un-asked teammate calls, but for hook-asked
calls the data shows the prompt reliably surfaces. The straightforward solution is to gate returns to ask or narrow it to an advisory dialog, though there's still the question of whether to preserve the
desktop Qt dialog experience — that's a design decision to consult on.
● 1. Noted — and probe7's report contained no note, so this also confirms: approval messages are dropped, only denial messages reach the teammate.
- Yes, both runtimes: the logging hook is loaded by teammate and lead alike, and it recorded nothing for any teammate ask — PermissionRequest fired in neither session for those.
- Both layers ran: the chain's auto-allow produced the (end-of-batch, delayed) systemMessage for the two echo calls, the temp hook's ask verifiably fired (log) — and you got prompted, so the ask overrode
the allow. For factor there was no [auto-allow] because the chain never allowed it; the settings rule asked → prompt.
- Fair challenge, and accepting it changes the conclusion: probe5's "silent allow" rests entirely on two observations we've since invalidated — its own "instant, no prompt" perception (worthless per
point 7) and the [auto-allow] pane line (delayed, rendered-regardless, misleading per point 6). Discarding probe5 as an invalid test, every valid data point tonight shows deterministic forwarding — there is
no observed classifier leak left to isolate. The one unverifiable unknown from that round: whether a prompt appeared in the lead TUI that nobody was watching for.
Which collapses the design question to: on all valid evidence, the gate abstaining to plain "ask" reaches you reliably, remote or local — the Qt dialog is then purely a desktop-UX nicety (richer commit
review), not a safety mechanism. Consult.
❯ brother stop b eing so annoying
does your last response's '2.' map to a different nuymber from the 1-8?
∴ The user is pointing out that my last response didn't include a "2." — I numbered my points as 1, 5, 6, and 8 to match their responses, skipping the items they'd already addressed. They're asking if my "2."
corresponds to something else, but there simply isn't one in my last message. I should just confirm plainly that no, there's no "2." in my previous response.
● My last response has no "2." — I kept your numbering and only replied to 1, 5, 6, and 8. Your "idk" on 2/3 and agreement on 4/7 needed no response. All numbers still map to the original 8-claim list.
❯ holy shit what is gpoing on??
my tui view:
tui
● 1. Noted — and probe7's report contained no note, so this also confirms: approval messages are dropped, only denial messages reach the teammate.
- Yes, both runtimes: the logging hook is loaded by teammate and lead alike, and it recorded nothing for any teammate ask — PermissionRequest fired in neither session for those.
- Both layers ran: the chain's auto-allow produced the (end-of-batch, delayed) systemMessage for the two echo calls, the temp hook's ask verifiably fired (log) — and you got prompted, so the ask overrode
the allow. For factor there was no [auto-allow] because the chain never allowed it; the settings rule asked → prompt.
- Fair challenge, and accepting it changes the conclusion: probe5's "silent allow" rests entirely on two observations we've since invalidated — its own "instant, no prompt" perception (worthless per
point 7) and the [auto-allow] pane line (delayed, rendered-regardless, misleading per point 6). Discarding probe5 as an invalid test, every valid data point tonight shows deterministic forwarding — there
is
no observed classifier leak left to isolate. The one unverifiable unknown from that round: whether a prompt appeared in the lead TUI that nobody was watching for.
Which collapses the design question to: on all valid evidence, the gate abstaining to plain "ask" reaches you reliably, remote or local — the Qt dialog is then purely a desktop-UX nicety (richer commit
review), not a safety mechanism. Consult.
❯ brother stop b eing so annoying
does your last response's '2.' map to a different nuymber from the 1-8?
line 2 of your msg starts with '2.
∴ I see the issue — the markdown renderer is auto-renumbering my ordered list. I wrote 1., 5., 6., 8., but GFM treats ordered lists sequentially regardless of the literal numbers I provided, so it rendered as
1., 2., 3., 4. instead.
The mapping is: rendered item 2 was my 5, item 3 was my 6, and item 4 was my 8. Going forward, I'll avoid using bare N. syntax when I need non-sequential references — I'll use bold formatting like 5. or
parentheses like 5) instead to preserve the actual numbers.
● Mystery solved: that's the markdown renderer, not me. I wrote 1. 5. 6. 8., but a markdown ordered list ignores the literal numbers after the first item and renumbers sequentially — so your TUI displayed 1,
2, 3, 4. Mapping of what you see: rendered 2 = my 5, rendered 3 = my 6, rendered 4 = my 8.
Lesson taken: for non-sequential references I'll write them as 5. bolded or "re 5:" instead of bare list syntax so the renderer can't renumber them.
````
</details>
<details>
<summary><strong>Raw conversation JSON (verbatim assistant message from conversation data)</strong></summary>
{
"parentUuid": "00000000-0000-0000-0000-000000000000",
"isSidechain": false,
"message": {
"model": "claude-3-7-sonnet-20250219",
"id": "msg_01xxxxxxxxxxxxxxxxxxxx",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "1. Noted — and probe7's report contained no note, so this also confirms: approval messages are dropped, only denial messages reach the teammate.\n5. Yes, both runtimes: the logging hook is loaded by teammate and lead alike, and it recorded nothing for any teammate ask — PermissionRequest fired in neither session for those.\n6. Both layers ran: the chain's auto-allow produced the (end-of-batch, delayed) systemMessage for the two echo calls, the temp hook's ask verifiably fired (log) — and you got prompted, so **the ask overrode the allow**. For factor there was no `[auto-allow]` because the chain never allowed it; the settings rule asked → prompt.\n8. Fair challenge, and accepting it changes the conclusion: probe5's \"silent allow\" rests entirely on two observations we've since invalidated — its own \"instant, no prompt\" perception (worthless per point 7) and the `[auto-allow]` pane line (delayed, rendered-regardless, misleading per point 6). Discarding probe5 as an invalid test, **every valid data point tonight shows deterministic forwarding — there is no observed classifier leak left to isolate.** The one unverifiable unknown from that round: whether a prompt appeared in the lead TUI that nobody was watching for.\n\nWhich collapses the design question to: on all valid evidence, the gate abstaining to plain \"ask\" reaches you reliably, remote or local — the Qt dialog is then purely a desktop-UX nicety (richer commit review), not a safety mechanism. Consult."
}
],
"stop_reason": "end_turn",
"stop_sequence": null,
"stop_details": null,
"usage": {
"input_tokens": 2,
"cache_creation_input_tokens": 1652,
"cache_read_input_tokens": 234777,
"output_tokens": 2148,
"output_tokens_details": {
"thinking_tokens": 1630
},
"server_tool_use": {
"web_search_requests": 0,
"web_fetch_requests": 0
},
"service_tier": "standard",
"cache_creation": {
"ephemeral_1h_input_tokens": 1652,
"ephemeral_5m_input_tokens": 0
},
"inference_geo": "not_available",
"iterations": [],
"speed": "standard"
}
},
"requestId": "req_01xxxxxxxxxxxxxxxxxxxx",
"type": "assistant",
"uuid": "00000000-0000-0000-0000-000000000001",
"timestamp": "2026-08-18T07:04:11.129Z",
"effort": "medium",
"session_id": "00000000-0000-0000-0000-000000000002",
"userType": "external",
"entrypoint": "cli",
"cwd": "/home/user/project",
"sessionId": "00000000-0000-0000-0000-000000000002",
"version": "2.1.234",
"gitBranch": "main",
"slug": "sample-session-slug"
}
</details>
Impact
- User confusion: The user cannot correctly interpret the agent's response. In this case, the user spent multiple turns debugging what turned out to be a rendering artefact.
- Semantic corruption: Numbered back-references are a common conversational pattern. The renderer silently changes their meaning.
- Agent self-contradiction: The agent's subsequent denial ("my last response has no '2.'") appears to gaslighting the user, when in fact both the agent and the user were telling the truth — they were just seeing different renderings of the same content.
Environment
- Claude Code version: 2.1.234
- Interface: CLI / TUI
- OS: Linux/Ubuntu Debian