gh-156333: rebuild the proactor self-pipe on EOF instead of busy-looping - #156343
aidaodedjl wants to merge 4 commits into
Conversation
d2702a4 to
2e1719c
Compare
2e1719c to
3ce66f2
Compare
…y-looping When the self-pipe socketpair of a BaseProactorEventLoop reaches a clean EOF (e.g. the OS tears the loopback connection down across a power or session state change on Windows), _loop_self_reading re-armed recv() on the dead socket, which completed immediately and rescheduled the callback forever, pinning one core at 100% CPU with nothing logged. Detect the EOF via the empty recv result and rebuild the socketpair instead: allocate the replacement first (so a failure leaves the previous state untouched), re-register signal.set_wakeup_fd on the new socket before closing the old one (mirroring close()), then arm the next read on the new socket so cross-thread wakeups keep working.
3ce66f2 to
c2ae83f
Compare
|
Heads-up: there is a sibling failure mode of the same line that this PR does not cover. When the pending Deterministic 15-line repro (CPython 3.11.15 and 3.14.7, 3/3 runs each) plus production evidence (8/8 stalls coinciding with the self-pipe error to the second; zero Would |
…nstead of losing wakeups A failed read on the self-pipe (for example an aborted overlapped operation on Windows, reported as ConnectionResetError with winerror 995 or 1236) left the loop running with no read armed: every later call_soon_threadsafe() or run_coroutine_threadsafe() from another thread enqueued a callback that nothing would ever wake the loop to run. Report the error, then rebuild the pipe and arm a fresh read. A future that is no longer the current one still means the loop is closing (pythongh-39010) and stops without touching the sockets. A recovery that fails in turn is reported and leaves the field unset so a later run_forever() can arm a read again.
…tor-self-pipe-eof
|
Reproduced independently before changing the patch: stock 3.13.15 and 3.10.11 on Windows 10, CancelIoEx on the pending read (winerror 995) and a foreign close of the read socket (1236) both leave the loop running with call_soon_threadsafe never delivered. No-sabotage baselines are clean on both. Recovery for that path is on the branch now: the except BaseException branch rebuilds the pipe and arms a fresh read when the failed future is still the current one. A stale future keeps the gh-39010 meaning (loop closing, sockets untouched). A recovery that fails in turn is reported once and leaves the read unarmed so a later run_forever() starts clean. The new mock tests fail on the unfixed code, and the patched interpreters pass their own test_asyncio proactor and windows suites, including test_read_self_pipe_restart. Since you offered: the branch is gh-156333-proactor-self-pipe-eof on my fork, and a 3.11.15/3.14.7 run against it would be useful. One caveat on the production numbers: 8/8 timestamp correlation plus zero 10054 fits a dead wakeup channel but doesn't rule out every other stall cause on its own, so a re-check against a build with the fix would close that. |
Summary
On Windows,
BaseProactorEventLoopwakes itself through a self-pipe — a loopback TCP socketpair created bysocket.socketpair(). If that connection reaches a clean EOF while the loop is running — for example the OS tears the idle loopback connection down across a power/session state change —_loop_self_readingre-armedrecv()on the dead socket, which completed immediately and rescheduled the callback forever: one core pinned at 100% CPU, no exception raised, nothing logged, the process never recovers. Reproduced deterministically on main (3.16.0a0, self-built): 346,702 re-arms of_loop_self_readingduring a 3-second idle sleep (a gracefulloop._csock.shutdown(socket.SHUT_WR)models the OS teardown).At EOF
f.result()returnsb'', which is not an exception, so control fell through to theelsebranch and armed a read that could never block again.Fix: when the recv result is empty, rebuild the socketpair instead of re-arming on the dead one:
signal.set_wakeup_fd()on the new socket before closing the old sockets, mirroring the ordering used byclose();call_soon_threadsafe) keep working after the rebuild.Verification (real Windows machine, self-built 3.16.0a0 from the commit this PR is based on)
test_asynciosuite: 35/35 files, 2,625 tests, 0 failures (both new tests included).main(git checkoutof the pristineproactor_events.py, tests re-run: mock test FAIL, functional test FAIL), then pass with the patch.Out of scope / follow-ups
ConnectionResetErrorfrom the pending recv instead of a clean EOF; that pre-existing path is not handled here.BaseSelectorEventLoop._read_from_self(used byWindowsSelectorEventLoop) has the same EOF busy-loop shape: measured 582,692_read_from_selfcalls during a 3s idle sleep on the same machine/reproducer shape. Will be reported separately.