Backgrounded process tree is group-killed ~106s after spawn (bg-pty host-dead watchdog / daemon auth_required window)
Versions observed: 2.1.220 (incident), 2.1.221 (current). macOS arm64 (Darwin 25.5.0).
Summary
A process tree containing a backgrounded task inside an interactive Claude Code
session was terminated three times at 105–107s (±1s) from spawn by an
external, group-directed signal sequence (catchable SIGTERM, then an escalation
the tree did not survive). The tree included a nested SDK-spawned claude, its
MCP servers, and the user's own CLI process (sequant run) — all sharing one
process group. Sessions that exit cleanly are never killed at any lifetime
(specimens: 131.7s, 316s, 26m54s clean runs).
Timeline (2026-07-29, UTC)
| Event | Time |
| -------------------------------------------------------------------- | ---------------------------- |
| Background daemon auth: proactive refresh failed → auth_required | 06:04:22.6 |
| Victim 1 spawn → death | 06:20:10 → 06:21:57 (106.2s) |
| Victim 2 spawn → death | 06:34:20 → 06:36:05 (105.4s) |
| Victim 3 spawn → death | 06:37:52 → 06:39:36 (104.6s) |
| Daemon idle-exits (idle 5s with no clients) | 09:48:23 |
All three deaths fall inside the daemon's unauthenticated window; none have
occurred in the days since the daemon exited. No deaths before the auth
failure either, across a multi-day record.
Evidence
- Victim lifetimes from
mDNSResponderper-PID stamps: 107.08s / 106.23s /
105.41s — spread 1.7s across three deaths, consistent with a fixed timer or
retry ladder anchored at spawn.
- The killed run's own captured stdout survives:
```
[01:34:18] ▸ #846 spec
! Received SIGTERM, shutting down gracefully...
✓ Aborted 1 active phase
```
…then nothing — cleanup truncated mid-sequence. So the first signal is
catchable SIGTERM, followed by a kill the handler did not survive.
- All MCP servers in the tree exited within a 60ms window at each death —
a single group-directed signal, not a cascade.
- The interactive session that spawned the backgrounded task survived each
time (it is in a different process group).
Candidate mechanisms (both in Claude Code, neither confirmed)
[bg-pty]host-dead watchdog. Extracted from the 2.1.220 binary: after
xhp = 30 failed socket connects to a bg-pty host whose PID is still alive,
the client runs process.kill(-t, "SIGKILL") — a process-group kill.
Backoff ladder [50,100,250,500,1000,2000]ms. Telemetry names:
tengu_bg_ptyhost_crash, tengu_bg_adopt_sock_unlinked,
tengu_bg_revival_guard. Matches the "alive but not accepting"
discriminator exactly; the ladder alone sums to 51.9s, so ~105s implies
~1.8s extra per attempt (unconfirmed).
- The background daemon during its
auth_requiredwindow.
CLAUDE_BG_BACKEND="daemon"; bg-pty hosts are spawned by a transient
daemon as a spare pool. The daemon logged
auth: headless daemon cannot complete OAuth 17 minutes before the first
death, and ~/.claude/daemon-auth-status.json recorded auth_required at
the same millisecond. Every death falls inside that window.
These may be one mechanism (an unauthenticated daemon wedging its bg-pty hosts,
whose clients then declare them dead and group-kill). The victims' post-turn
hang (Stop-hook→SessionEnd gaps of 58–96s vs a 0.36s baseline) appears to be
the same defect, not a separate one.
Why this is a bug and not policy
The group it kills contains the user's own ancestor processes — the
backgrounded command the user launched, not just Claude-internal helpers. A
watchdog for a wedged PTY host should not be able to SIGKILL the user's build,
test run, or (here) a 30-minute orchestration process as collateral.
Ruled out by direct measurement
Idle-output timeouts, lifetime caps, CLAUDE_STREAM_IDLE_TIMEOUT_MS, shellTMOUT/ulimit, launchd jobs, jetsam/memory pressure, third-party tooling
(entire CLI), and the victim application signalling itself (its only
cross-process kill path is CLI-gated and self/parent-guarded). Full matrix indocs/incidents/856/README.md of sequant-io/sequant
(sequant-io/sequant#856).
Update 2026-08-03: the daemon no longer spawns at all, blocking our repro
On 2.1.221 with healthy auth (claude auth status → loggedIn: true),
backgrounding a task from an interactive session spawns no daemon — no~/.claude/daemon.log activity since the 2026-07-29 idle-exit, no/tmp/cc-daemon-<uid>/ socket dir, no bg-pty-host processes. The
backgrounded task runs as a plain child of the interactive claude in its own
process group.
We had attributed the daemon's earlier absence to an auth_required cooldown;
healthy auth disproves that. Whatever gates CLAUDE_BG_BACKEND="daemon"
activation (rollout flag? version change in 2.1.221?) is now the blocker for
reproducing either candidate mechanism — and also means post-turn hangs on this
machine (63s/281s/345s observed since) are currently harmless: with no daemon,
nothing kills them.
Repro status
Not reproducible on demand: the vulnerable window requires a live daemon inauth_required state, and (per the update above) the daemon currently does
not spawn at all on this machine. Prepared instrumentation, verified by
positive control against group-directed signals, is committed insequant-io/sequant under docs/incidents/856/tools/:
signal-canary.c— SA_SIGINFO handler capturingsiginfo_t.si_pid(the
sender) for the catchable SIGTERM. Used because on a SIP-enabled machine the
dtrace proc:::signal-send provider is withheld entirely.
verify-capture.sh— positive-control test of the whole rig (5/5, no sudo).induce-bgpty-hang.sh— SIGSTOPs a bg-pty host to induce the exact
"alive but not accepting" state the watchdog keys on.
Negative control: with the daemon down, a canary-instrumented backgrounded task
ran 200s (≈2× the threshold) untouched — backgrounding alone is not sufficient.
Ask
- Confirm whether the bg-pty host-dead watchdog can group-kill a tree
containing the client's own ancestors, and bound its blast radius.
- Confirm what the daemon does to live bg sessions when it enters
auth_required — the observed window correlation is exact.
- If a fixed ~105s timer exists in either path, name it so field reports can
match on it.
- What gates
CLAUDE_BG_BACKEND="daemon"activation? It no longer engages on
this machine as of 2.1.221, which blocks our repro — if it was disabled or
changed in 2.1.221, that would be useful to know (and may itself be the fix).