Skip to content

Session/chat metadata (status, title, transcript) is stale/unreliable relative to actual sub-session state #4904

Description

@pfeurean

Summary

When a session spawns multiple delegated sub-chats via create_session (relationship: "currentSession"), the session-management tool surface (list_sessions, get_session_context) reports status, title, and transcript "freshness" that do not reliably reflect the actual state of the sub-chats. An orchestrating agent (or human) monitoring delegated work cannot trust these signals and is forced to fall back to ground-truth verification (reading files directly, running builds) to determine real progress.

Environment

GitHub Copilot CLI — the create_session, list_sessions, and get_session_context tool surface used to spawn and monitor delegated sub-agent work. Reproduced with sub-chats created via create_session with relationship: "currentSession", each attaching as a separate ?chat=<uuid> thread under one session (e.g. agent-host-session://copilotcli/<session-guid>?chat=<chat-uuid>).

Repro steps

  1. In a session, call create_session 3+ times with relationship: "currentSession", each with a distinct, few-minute-long autonomous coding task.
  2. While those run, poll list_sessions({ withChanges: true }) and get_session_context on each ?chat= UUID at various intervals.
  3. Compare the reported status / state / title / modifiedAt against actual completion (verified via direct file inspection or build output) at the same wall-clock moments.

Issue 1 — Aggregate session status field is wrong/stale

  • Called list_sessions({ withChanges: true }) on a session with 6 active sub-chats.
  • Result: "status": "inputNeeded" — implying at least one chat is blocked waiting for a reply.
  • Drilled into all 6 individual chats via get_session_context (detail: "full" and "digest"). None had actually asked a question or paused for input — 5 showed "state": "complete" with clean final reports, 1 showed "state": "inProgress" mid-tool-execution.
  • Re-polled the same session ~15 minutes later: status was still "inputNeeded", with modifiedAt having advanced (e.g. from 2026-09-18T16:41:16.511Z to 2026-09-18T16:56:16.347Z) — i.e. something updated the timestamp, but no new content appeared and the label never resolved to something accurate.
  • Ground truth (direct file read + dotnet build) confirmed the "still in progress" chat's target file was in fact already complete and correct — the tool's own status reporting lagged reality by at least that 15-minute window, and did not appear to self-correct.

Expected: status should reflect real-time chat state — inputNeeded only when a chat is genuinely blocked on a prompt/question, idle/complete otherwise.

Actual: the label was flatly wrong for an extended period and did not self-correct on subsequent polls.

Issue 2 — Per-chat transcript can also appear stale under get_session_context

  • On the chat still reporting "state": "inProgress", two separate get_session_context calls, several minutes apart, returned byte-identical transcript content (same last tool call: a dependency scan) while state remained "inProgress" both times.
  • There is no way to distinguish, from the tool output alone, "still actually running, transcript just hasn't been persisted yet" from "actually finished/stalled, but state was never flipped."
  • Only resolved by bypassing the tool entirely: reading the target file directly and running the actual build, which showed the work was done and correct.

Issue 3 — Session/chat title does not update to reflect current activity

  • A hub session's title remained fixed to the description of the very first task given to it, even after 8 additional, unrelated-by-name sub-chats were appended to the same session via relationship: "currentSession".
  • Expected: title either reflects the most recent/active chat's task, or is otherwise kept in sync with what's actually running.
  • Actual: title is frozen at creation time and never revisited, making the session list UI/API misleading about what a session is currently doing.

Contributing/related quirk (same tooling surface, included for context)

Sub-chats created via create_session each appear to get an isolated SQL-backed todo/database context — when a sub-chat's prompt instructed it to "update the SQL todos table," it found an empty table (not the parent's authoritative table) and correctly declined to fabricate rows, self-reporting the discrepancy. This means workflows relying on session-management tools to reflect or propagate cross-chat/cross-session state cannot trust automatic state propagation — it must be manually reconciled by the orchestrating session via direct verification, which is what surfaced issues 1-3 above.

Impact

An orchestrating agent (or a human operator) monitoring delegated sub-session work via list_sessions/get_session_context cannot trust status, title, or transcript "freshness" as signals of real progress or blockage. The only reliable signal found was bypassing these tools and checking ground truth directly (source files on disk, actual build/test output). This significantly undermines the usefulness of the delegated/background sub-agent workflow for any multi-chat orchestration.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions