Cut fixed per-request costs out of the reshard control plane - #808
Open
yinlin09 wants to merge 1 commit into
Open
Cut fixed per-request costs out of the reshard control plane#808yinlin09 wants to merge 1 commit into
yinlin09 wants to merge 1 commit into
Conversation
Every Stage-3 coordination paid four avoidable fixed costs on its framed RPCs (coordinate, GET_METADATA, receiver arm): 1. The 4-byte length prefix and the body went out as two separate send() calls with Nagle enabled, on requests and responses alike, exposing every hop to the delayed-ACK stall (tens of ms on small RPCs). 2. The destination controller was asked for every registered unit's full pool manifest on every request, although work units register once per engine lifetime. 3. Client sockets carried no keepalive, so a black-holed peer was only detected at the full receive timeout. 4. The framed server never reaped its per-connection threads: one std::thread handle and stack per request, held until shutdown. Changes: - framed_rpc: single-buffer framing on both directions; TCP_NODELAY on client and accepted sockets; SO_KEEPALIVE plus TCP_USER_TIMEOUT bounded by the call's I/O timeout on client sockets; the accept loop joins finished connection threads. - reshard_coordinator: destination metadata is cached per controller address. Staleness (engine replacement) surfaces as a plan-build or receiver-arm failure, both side-effect-free beyond the abandoned claim, and is repaired by invalidate + fresh query + one replay; a failure after the receiver ack is never replayed. Validation: reshard package builds and reshard_service_test passes in the glibc-2.36 container, including the new RemoteMetadataCachedAndRefreshedOnStaleFailure test (cache hit on the second request, exactly one refetch after a fingerprint-mismatch replay).
copybara-service Bot
pushed a commit
that referenced
this pull request
Aug 29, 2026
Every Stage-3 coordination paid four avoidable fixed costs on its framed RPCs (coordinate, GET_METADATA, receiver arm): 1. The 4-byte length prefix and the body went out as two separate send() calls with Nagle enabled, on requests and responses alike, exposing every hop to the delayed-ACK stall (tens of ms on small RPCs). 2. The destination controller was asked for every registered unit's full pool manifest on every request, although work units register once per engine lifetime. 3. Client sockets carried no keepalive, so a black-holed peer was only detected at the full receive timeout. 4. The framed server never reaped its per-connection threads: one std::thread handle and stack per request, held until shutdown. Changes: - framed_rpc: single-buffer framing on both directions; TCP_NODELAY on client and accepted sockets; SO_KEEPALIVE plus TCP_USER_TIMEOUT bounded by the call's I/O timeout on client sockets; the accept loop joins finished connection threads. - reshard_coordinator: destination metadata is cached per controller address. Staleness (engine replacement) surfaces as a plan-build or receiver-arm failure, both side-effect-free beyond the abandoned claim, and is repaired by invalidate + fresh query + one replay; a failure after the receiver ack is never replayed. Validation: reshard package builds and reshard_service_test passes in the glibc-2.36 container, including the new RemoteMetadataCachedAndRefreshedOnStaleFailure test (cache hit on the second request, exactly one refetch after a fingerprint-mismatch replay). GitHub: #808 PiperOrigin-RevId: 973205820
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every Stage-3 coordination pays fixed costs on its framed RPCs (coordinate, GET_METADATA, receiver arm):
send()calls with Nagle enabled, on requests and responses alike, exposing every hop to the delayed-ACK stall.std::threadhandle and stack per request, held until shutdown.Changes: single-buffer framing both directions plus TCP_NODELAY on client and accepted sockets; SO_KEEPALIVE and TCP_USER_TIMEOUT bounded by the call's I/O timeout; the accept loop joins finished connection threads; and the destination metadata is cached per controller address — staleness (engine replacement) surfaces as a plan-build or receiver-arm failure, both side-effect-free beyond the abandoned claim, and is repaired by invalidate + fresh query + one replay. A failure after the receiver ack is never replayed.
Validation
reshard_service_testpasses in the glibc-2.36 build container, including the newRemoteMetadataCachedAndRefreshedOnStaleFailuretest (cache hit on the second request; exactly one refetch after a fingerprint-mismatch replay).Benchmarked on tpu7x 1P1D (Qwen3.5-397B, prefill PCP8 → decode DP8, vllm-torchtpu with background Stage-3 submission, 888 transfers per leg, zero transfer failures either leg): plan build including the peer-metadata path drops from p50 2.1 ms / max 5.1 ms to p50 1.5 ms / max 3.0 ms; end-to-end serving metrics are unchanged. The coordinate call's floor is dominated by sender dispatch (p50 ~26 ms under 4-way concurrent coordination — the reply currently awaits all 8 sender-worker acks), which this PR deliberately does not touch: returning the coordinate reply at receiver-arm and letting sender dispatch complete asynchronously is the follow-up that would remove most of the remaining floor.