Summary
When a session spawns multiple delegated sub-chats via create_session (relationship: "currentSession"), the session-management tool surface (list_sessions, get_session_context) reports status, title, and transcript "freshness" that do not reliably reflect the actual state of the sub-chats. An orchestrating agent (or human) monitoring delegated work cannot trust these signals and is forced to fall back to ground-truth verification (reading files directly, running builds) to determine real progress.
Environment
GitHub Copilot CLI — the create_session, list_sessions, and get_session_context tool surface used to spawn and monitor delegated sub-agent work. Reproduced with sub-chats created via create_session with relationship: "currentSession", each attaching as a separate ?chat=<uuid> thread under one session (e.g. agent-host-session://copilotcli/<session-guid>?chat=<chat-uuid>).
Repro steps
- In a session, call
create_session 3+ times with relationship: "currentSession", each with a distinct, few-minute-long autonomous coding task.
- While those run, poll
list_sessions({ withChanges: true }) and get_session_context on each ?chat= UUID at various intervals.
- Compare the reported
status / state / title / modifiedAt against actual completion (verified via direct file inspection or build output) at the same wall-clock moments.
Issue 1 — Aggregate session status field is wrong/stale
- Called
list_sessions({ withChanges: true }) on a session with 6 active sub-chats.
- Result:
"status": "inputNeeded" — implying at least one chat is blocked waiting for a reply.
- Drilled into all 6 individual chats via
get_session_context (detail: "full" and "digest"). None had actually asked a question or paused for input — 5 showed "state": "complete" with clean final reports, 1 showed "state": "inProgress" mid-tool-execution.
- Re-polled the same session ~15 minutes later:
status was still "inputNeeded", with modifiedAt having advanced (e.g. from 2026-09-18T16:41:16.511Z to 2026-09-18T16:56:16.347Z) — i.e. something updated the timestamp, but no new content appeared and the label never resolved to something accurate.
- Ground truth (direct file read +
dotnet build) confirmed the "still in progress" chat's target file was in fact already complete and correct — the tool's own status reporting lagged reality by at least that 15-minute window, and did not appear to self-correct.
Expected: status should reflect real-time chat state — inputNeeded only when a chat is genuinely blocked on a prompt/question, idle/complete otherwise.
Actual: the label was flatly wrong for an extended period and did not self-correct on subsequent polls.
Issue 2 — Per-chat transcript can also appear stale under get_session_context
- On the chat still reporting
"state": "inProgress", two separate get_session_context calls, several minutes apart, returned byte-identical transcript content (same last tool call: a dependency scan) while state remained "inProgress" both times.
- There is no way to distinguish, from the tool output alone, "still actually running, transcript just hasn't been persisted yet" from "actually finished/stalled, but state was never flipped."
- Only resolved by bypassing the tool entirely: reading the target file directly and running the actual build, which showed the work was done and correct.
Issue 3 — Session/chat title does not update to reflect current activity
- A hub session's title remained fixed to the description of the very first task given to it, even after 8 additional, unrelated-by-name sub-chats were appended to the same session via
relationship: "currentSession".
- Expected: title either reflects the most recent/active chat's task, or is otherwise kept in sync with what's actually running.
- Actual: title is frozen at creation time and never revisited, making the session list UI/API misleading about what a session is currently doing.
Contributing/related quirk (same tooling surface, included for context)
Sub-chats created via create_session each appear to get an isolated SQL-backed todo/database context — when a sub-chat's prompt instructed it to "update the SQL todos table," it found an empty table (not the parent's authoritative table) and correctly declined to fabricate rows, self-reporting the discrepancy. This means workflows relying on session-management tools to reflect or propagate cross-chat/cross-session state cannot trust automatic state propagation — it must be manually reconciled by the orchestrating session via direct verification, which is what surfaced issues 1-3 above.
Impact
An orchestrating agent (or a human operator) monitoring delegated sub-session work via list_sessions/get_session_context cannot trust status, title, or transcript "freshness" as signals of real progress or blockage. The only reliable signal found was bypassing these tools and checking ground truth directly (source files on disk, actual build/test output). This significantly undermines the usefulness of the delegated/background sub-agent workflow for any multi-chat orchestration.
Summary
When a session spawns multiple delegated sub-chats via
create_session(relationship: "currentSession"), the session-management tool surface (list_sessions,get_session_context) reportsstatus,title, and transcript "freshness" that do not reliably reflect the actual state of the sub-chats. An orchestrating agent (or human) monitoring delegated work cannot trust these signals and is forced to fall back to ground-truth verification (reading files directly, running builds) to determine real progress.Environment
GitHub Copilot CLI — the
create_session,list_sessions, andget_session_contexttool surface used to spawn and monitor delegated sub-agent work. Reproduced with sub-chats created viacreate_sessionwithrelationship: "currentSession", each attaching as a separate?chat=<uuid>thread under one session (e.g.agent-host-session://copilotcli/<session-guid>?chat=<chat-uuid>).Repro steps
create_session3+ times withrelationship: "currentSession", each with a distinct, few-minute-long autonomous coding task.list_sessions({ withChanges: true })andget_session_contexton each?chat=UUID at various intervals.status/state/title/modifiedAtagainst actual completion (verified via direct file inspection or build output) at the same wall-clock moments.Issue 1 — Aggregate session
statusfield is wrong/stalelist_sessions({ withChanges: true })on a session with 6 active sub-chats."status": "inputNeeded"— implying at least one chat is blocked waiting for a reply.get_session_context(detail: "full"and"digest"). None had actually asked a question or paused for input — 5 showed"state": "complete"with clean final reports, 1 showed"state": "inProgress"mid-tool-execution.statuswas still"inputNeeded", withmodifiedAthaving advanced (e.g. from2026-09-18T16:41:16.511Zto2026-09-18T16:56:16.347Z) — i.e. something updated the timestamp, but no new content appeared and the label never resolved to something accurate.dotnet build) confirmed the "still in progress" chat's target file was in fact already complete and correct — the tool's own status reporting lagged reality by at least that 15-minute window, and did not appear to self-correct.Expected:
statusshould reflect real-time chat state —inputNeededonly when a chat is genuinely blocked on a prompt/question,idle/completeotherwise.Actual: the label was flatly wrong for an extended period and did not self-correct on subsequent polls.
Issue 2 — Per-chat transcript can also appear stale under
get_session_context"state": "inProgress", two separateget_session_contextcalls, several minutes apart, returned byte-identical transcript content (same last tool call: a dependency scan) whilestateremained"inProgress"both times.Issue 3 — Session/chat
titledoes not update to reflect current activityrelationship: "currentSession".Contributing/related quirk (same tooling surface, included for context)
Sub-chats created via
create_sessioneach appear to get an isolated SQL-backed todo/database context — when a sub-chat's prompt instructed it to "update the SQL todos table," it found an empty table (not the parent's authoritative table) and correctly declined to fabricate rows, self-reporting the discrepancy. This means workflows relying on session-management tools to reflect or propagate cross-chat/cross-session state cannot trust automatic state propagation — it must be manually reconciled by the orchestrating session via direct verification, which is what surfaced issues 1-3 above.Impact
An orchestrating agent (or a human operator) monitoring delegated sub-session work via
list_sessions/get_session_contextcannot truststatus,title, or transcript "freshness" as signals of real progress or blockage. The only reliable signal found was bypassing these tools and checking ground truth directly (source files on disk, actual build/test output). This significantly undermines the usefulness of the delegated/background sub-agent workflow for any multi-chat orchestration.