http2: count a lost ping exactly once, and never on a failed connection - #262
Open
ali-sayyah wants to merge 1 commit into
Open
http2: count a lost ping exactly once, and never on a failed connection#262ali-sayyah wants to merge 1 commit into
ali-sayyah wants to merge 1 commit into
Conversation
ali-sayyah
force-pushed
the
http2-close-healthcheck
branch
from
August 17, 2026 19:24
75cdb10 to
5d6dfed
Compare
The health-check timer is re-armed after every ReadFrame return,
including the final erroring one, and no close path stops it: one
ReadIdleTimeout after every connection close a post-mortem health
check pings the dead connection, fails immediately, and reports a
spurious CountError("conn_close_lost_ping"). A genuinely lost ping is
counted twice: the real detection, then the post-close echo.
CL 198040 shipped the health check with the timer stopped when the
read loop exits; CL 572378, an unrelated test-infrastructure
conversion, dropped the stop with no behavioral rationale (v0.22.0
has it, v0.23.0 does not). Restore the stop, publish the read loop's
terminal exit under cc.mu before the failure is counted, and make
lost-ping classification a single-claim close: the eligibility check
and the claim share one critical section in closeForLostPing, so
overlapping health checks, or a ping racing the read loop's terminal
exit, can never produce a second count. The CountError callback runs
outside the lock. A ping that fails on a connection that is live at
claim time is still counted and still closes the connection,
preserving the write-blocked-ping detection of CL 354389.
This fixes the legacy transport, used on Go versions before 1.27 and
with the http2legacy build tag (go build -tags=http2legacy). The
default Go 1.27 path uses net/http's copy of this code, which has the
same defect and is fixed separately; that change closes the issue.
Updates golang/go#80920
ali-sayyah
force-pushed
the
http2-close-healthcheck
branch
from
August 17, 2026 20:07
5d6dfed to
8261ddc
Compare
ali-sayyah
added a commit
to ali-sayyah/go
that referenced
this pull request
Aug 17, 2026
…d connection
The HTTP/2 health-check timer is re-armed after every ReadFrame
return, including the final erroring one, and no close path stops it:
one SendPingTimeout after every connection close a post-mortem health
check pings the dead connection, fails immediately, and reports a
spurious CountError("conn_close_lost_ping"). A genuinely lost ping is
counted twice: the real detection, then the post-close echo.
Stop the timer when the read loop exits, publish the read loop's
terminal exit under cc.mu before the failure is counted, and make
lost-ping classification a single-claim close: the eligibility check
and the claim share one critical section in closeForLostPing, so
overlapping health checks, or a ping racing the read loop's terminal
exit, can never produce a second count. The CountError callback runs
outside the lock. A ping that fails on a connection that is live at
claim time is still counted and still closes the connection,
preserving write-blocked-ping detection.
The equivalent x/net change (golang/net#262) fixes the legacy
transport used on Go versions before 1.27 and with the http2legacy
build tag.
Fixes golang#80920
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The health-check timer is re-armed after every ReadFrame return,
including the final erroring one, and no close path stops it: one
ReadIdleTimeout after every connection close a post-mortem health
check pings the dead connection, fails immediately, and reports a
spurious CountError("conn_close_lost_ping"). A genuinely lost ping is
counted twice: the real detection, then the post-close echo.
CL 198040 shipped the health check with the timer stopped when the
read loop exits; CL 572378, an unrelated test-infrastructure
conversion, dropped the stop with no behavioral rationale (v0.22.0
has it, v0.23.0 does not). Restore the stop, publish the read loop's
terminal exit under cc.mu before the failure is counted, and make
lost-ping classification a single-claim close: the eligibility check
and the claim share one critical section in closeForLostPing, so
overlapping health checks, or a ping racing the read loop's terminal
exit, can never produce a second count. The CountError callback runs
outside the lock. A ping that fails on a connection that is live at
claim time is still counted and still closes the connection,
preserving the write-blocked-ping detection of CL 354389.
This fixes the legacy transport, used on Go versions before 1.27 and
with the http2legacy build tag (go build -tags=http2legacy). The
default Go 1.27 path uses net/http's copy of this code, which has the
same defect and is fixed separately; that change closes the issue.
Updates golang/go#80920