Skip to content

gh-156333: rebuild the proactor self-pipe on EOF instead of busy-looping - #156343

Open
aidaodedjl wants to merge 4 commits into
python:mainfrom
aidaodedjl:gh-156333-proactor-self-pipe-eof
Open

aidaodedjl wants to merge 4 commits into
python:mainfrom
aidaodedjl:gh-156333-proactor-self-pipe-eof

Conversation

@aidaodedjl

@aidaodedjl aidaodedjl commented Aug 25, 2026 •

Copy link
Copy Markdown

Summary

On Windows, BaseProactorEventLoop wakes itself through a self-pipe — a loopback TCP socketpair created by socket.socketpair(). If that connection reaches a clean EOF while the loop is running — for example the OS tears the idle loopback connection down across a power/session state change — _loop_self_reading re-armed recv() on the dead socket, which completed immediately and rescheduled the callback forever: one core pinned at 100% CPU, no exception raised, nothing logged, the process never recovers. Reproduced deterministically on main (3.16.0a0, self-built): 346,702 re-arms of _loop_self_reading during a 3-second idle sleep (a graceful loop._csock.shutdown(socket.SHUT_WR) models the OS teardown).

At EOF f.result() returns b'', which is not an exception, so control fell through to the else branch and armed a read that could never block again.

Fix: when the recv result is empty, rebuild the socketpair instead of re-arming on the dead one:

  • allocate the replacement pair first, so a failure inside the rebuild leaves the previous state untouched and the error propagates instead of being swallowed by the catch-all handler;
  • re-register signal.set_wakeup_fd() on the new socket before closing the old sockets, mirroring the ordering used by close();
  • arm the next read on the new socket, so cross-thread wakeups (call_soon_threadsafe) keep working after the rebuild.

Verification (real Windows machine, self-built 3.16.0a0 from the commit this PR is based on)

  • The issue's reproducer: 346,702 re-arms / 1.42s CPU during a 3s sleep before → 0 re-arms / 0.00s CPU after; the loop still completes cross-thread wakeups after the rebuild.
  • Full test_asyncio suite: 35/35 files, 2,625 tests, 0 failures (both new tests included).
  • Both new regression tests were verified to fail on unpatched main (git checkout of the pristine proactor_events.py, tests re-run: mock test FAIL, functional test FAIL), then pass with the patch.

Out of scope / follow-ups

  • The same OS teardown can also surface as ConnectionResetError from the pending recv instead of a clean EOF; that pre-existing path is not handled here.
  • BaseSelectorEventLoop._read_from_self (used by WindowsSelectorEventLoop) has the same EOF busy-loop shape: measured 582,692 _read_from_self calls during a 3s idle sleep on the same machine/reproducer shape. Will be reported separately.

@python-cla-bot

python-cla-bot Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

All commit authors signed the Contributor License Agreement.

CLA signed

@aidaodedjl
aidaodedjl force-pushed the gh-156333-proactor-self-pipe-eof branch from 2e1719c to 3ce66f2 Compare September 4, 2026 05:23
…y-looping

When the self-pipe socketpair of a BaseProactorEventLoop reaches a clean
EOF (e.g. the OS tears the loopback connection down across a power or
session state change on Windows), _loop_self_reading re-armed recv() on
the dead socket, which completed immediately and rescheduled the callback
forever, pinning one core at 100% CPU with nothing logged.

Detect the EOF via the empty recv result and rebuild the socketpair
instead: allocate the replacement first (so a failure leaves the previous
state untouched), re-register signal.set_wakeup_fd on the new socket
before closing the old one (mirroring close()), then arm the next read on
the new socket so cross-thread wakeups keep working.
@aidaodedjl
aidaodedjl force-pushed the gh-156333-proactor-self-pipe-eof branch from 3ce66f2 to c2ae83f Compare September 4, 2026 14:52
@xksk-doer

Copy link
Copy Markdown

Heads-up: there is a sibling failure mode of the same line that this PR does not cover.

When the pending recv() on the self-pipe is aborted rather than reaching EOF — f.result() raises ConnectionResetError because finish_socket_func() maps ERROR_OPERATION_ABORTED (WinError 995) / ERROR_NETNAME_DELETED (1236) to it — control still falls into except BaseException in _loop_self_reading(), which reports once and returns without re-arming. The loop then keeps running=True while its wakeup channel is permanently dead: every later call_soon_threadsafe() / run_coroutine_threadsafe() from another thread enqueues a callback that never runs. Compared with the EOF case this is quieter (no CPU burn) and therefore harder to notice, and it applies to any exception on that path, not only 995.

Deterministic 15-line repro (CPython 3.11.15 and 3.14.7, 3/3 runs each) plus production evidence (8/8 stalls coinciding with the self-pipe error to the second; zero WinError 10054 / WSAECONNRESET in the same window) are in #156333.

Would _rebuild_self_pipe() be safe to call from that branch too — or is re-arming there the preferred minimal fix? Happy to test a patch on Windows (3.11.15 + 3.14.7) if that helps.

…nstead of losing wakeups

A failed read on the self-pipe (for example an aborted overlapped
operation on Windows, reported as ConnectionResetError with winerror
995 or 1236) left the loop running with no read armed: every later
call_soon_threadsafe() or run_coroutine_threadsafe() from another
thread enqueued a callback that nothing would ever wake the loop to
run.  Report the error, then rebuild the pipe and arm a fresh read.

A future that is no longer the current one still means the loop is
closing (pythongh-39010) and stops without touching the sockets.  A recovery
that fails in turn is reported and leaves the field unset so a later
run_forever() can arm a read again.
@aidaodedjl

Copy link
Copy Markdown
Author

Reproduced independently before changing the patch: stock 3.13.15 and 3.10.11 on Windows 10, CancelIoEx on the pending read (winerror 995) and a foreign close of the read socket (1236) both leave the loop running with call_soon_threadsafe never delivered. No-sabotage baselines are clean on both.

Recovery for that path is on the branch now: the except BaseException branch rebuilds the pipe and arms a fresh read when the failed future is still the current one. A stale future keeps the gh-39010 meaning (loop closing, sockets untouched). A recovery that fails in turn is reported once and leaves the read unarmed so a later run_forever() starts clean. The new mock tests fail on the unfixed code, and the patched interpreters pass their own test_asyncio proactor and windows suites, including test_read_self_pipe_restart.

Since you offered: the branch is gh-156333-proactor-self-pipe-eof on my fork, and a 3.11.15/3.14.7 run against it would be useful. One caveat on the production numbers: 8/8 timestamp correlation plus zero 10054 fits a dead wakeup channel but doesn't rule out every other stall cause on its own, so a re-check against a build with the fix would close that.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants