Repository navigation
stdio_server busy-loops at 100% CPU on Windows after the host suspends and resumes #3411
Description
Activity
- addedv2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)v1Affects the v1.x maintenance lineAffects the v1.x maintenance line
on Aug 28, 2026 Poked at where a spin could even live in this path, since the SDK-side surface is small:
stdio_server()reads stdin throughanyio.wrap_file(...), which does plain blocking reads on a worker thread — the stdin handle is never registered with the event loop at all. That fits your thread profile exactly: the reader thread is parked inReadFile(0 CPU), and in an otherwise idle FastMCP stdio server the only IOCP-registered handles are the proactor's own internals (its self-waking socket pair etc.). So the spin is almost certainly insideProactorEventLoop/IocpProactor, below the SDK.That has an awkward consequence for your mitigation 2: "detect the dead pipe and exit" already works today for the case where the pipe really dies — the reader thread sees EOF, the reader task finishes, and the server shuts down cleanly. In your hang stdin isn't dead (the worker thread is still blocked on it), so an SDK-level dead-pipe check would never fire.
Which makes your selector-policy experiment the most valuable next data point, because it swaps out exactly the suspect component. If
WindowsSelectorEventLoopPolicyavoids the spin, the SDK has a real lever: a stdio server doesn't need the proactor unless user code spawns subprocesses, so the stdio run path could use a selector loop on Windows — or at minimum document that as the workaround. Changing the default is a maintainer call though, since it silently breaksasyncio.create_subprocess_*inside tool handlers.For capture:
py-spy dump --native --pid <pid>, three or four samples a few seconds apart. If the main thread keeps showingIocpProactor._poll/proactor_eventsframes with nothing of yours above them, that pins the loop itself, and the native frames should show which handle keeps completing.If the maintainers settle on a direction here (selector loop for stdio, docs, or something else), happy to put the PR together — though I can't drive suspend/resume myself, so verification would lean on your repro.
Thanks — that narrows it usefully, and I checked the
anyioside against the
installed source rather than taking it on trust. You're right:# anyio 4.14.2, _core/_fileio.py:125 async def readline(self) -> AnyStr: return await to_thread.run_sync(self._fp.readline, limiter=self._limiter)
AsyncFileroutes every read throughto_thread.run_sync, so the stdin handle
is never registered with the loop at all. And I take your point that mitigation 2
is a dead end for this particular hang: stdin isn't dead here, so a dead-pipe
check would never fire.Baseline capture (healthy, idle server)
py-spydoes work against the uv-managed CPython on Windows, so here is the
idle baseline with native frames:Thread 12292 (idle): "MainThread" NtRemoveIoCompletion (ntdll.dll) GetQueuedCompletionStatus (KERNELBASE.dll) PyInit__overlapped (_overlapped.pyd) ... _poll (asyncio\windows_events.py:778) select (asyncio\windows_events.py:446) _run_once (asyncio\base_events.py:2022) run_forever (asyncio\base_events.py:683) run_until_complete (asyncio\base_events.py:712) run (asyncio\runners.py:119) run (anyio\_backends\_asyncio.py:2481) run (anyio\_core\_eventloop.py:83) run (mcp\server\fastmcp\server.py:299) main (run.py:7) Thread 19416 (idle): "AnyIO worker thread" NtReadFile (ntdll.dll) ReadFile (KERNELBASE.dll) read (ucrtbase.dll) ... run (anyio\_backends\_asyncio.py:1033)So the worker thread is parked in
ReadFileexactly as you described, and the
main thread sits inGetQueuedCompletionStatusunderIocpProactor._poll.One caveat worth flagging before the spin dump arrives: a healthy idle
server already shows_poll (windows_events.py:778)/selecton the main
thread. So the Python frames alone will not distinguish the hang from normal
idle — the discriminators are py-spy's thread state (idlevsactive+gil) and
the CPU time. I mention it so the eventual spin dump doesn't get read as
"same frames, therefore nothing to see".Capture is armed
I have a watchdog running that samples CPU of the server processes; anything
holding >80% of one core for 15s gets fourpy-spy dump --nativesamples three
seconds apart written to a file, and is then killed. So the next suspend/resume
that reproduces this should produce the dumps without me having to be at the
machine. I've verified the capture path end to end against a deliberately
spinning decoy process.On ordering
I'm deliberately not applying the selector-policy workaround yet. Doing so
would remove my ability to capture the proactor in the spinning state, and that
dump seems like the more perishable of the two data points. My plan is:- Wait for a reproduction, collect the native dumps (unattended, as above).
- Then switch this process to
WindowsSelectorEventLoopPolicyvia a
sitecustomize.pyin the tool's venv and run for a while across several
suspend/resume cycles to see whether the spin recurs.
Say the word if you'd rather have the selector result first and are content to
treat the spin dump as optional — happy to flip the order.Caveats on my side
This is a single machine, one Windows build, one Python (uv-managed CPython
3.13.14 on Windows 10 19045). I haven't reproduced it on demand — it has
happened on natural suspend/resume cycles, so "n" is small and my reproduction
latency is however long until the next one. If there's a faster way to force the
condition than actually sleeping the machine, I'm glad to try it.And yes — happy to test a patch or a PR branch against the real repro whenever
you have something you'd like exercised.Reproduced, captured unattended, and then the selector experiment came back with
a result I did not expect: it does not fix the spin — but it makes the cause
visible. I think the two runs together point at one root cause for both loops.It does not need a long sleep. A ~40 second sleep was enough:
21:44:10 Kernel-Power 42 entering sleep 21:44:48 Kernel-Power 131 resume 21:45:26 watchdog: pid 26360 at 91% of one core 21:45:36 watchdog: pid 4832 at 160% of one core1. Proactor loop (stock behaviour)
Four
py-spy dump --nativesamples per process, three seconds apart. All 8
samples across both processes:Thread 12292 (active): "MainThread" NtRemoveIoCompletion (ntdll.dll) GetQueuedCompletionStatus (KERNELBASE.dll) PyInit__overlapped (_overlapped.pyd) ... _poll (asyncio\windows_events.py:778) select (asyncio\windows_events.py:446) _run_once (asyncio\base_events.py:2022) run_forever (asyncio\base_events.py:683) Thread 19416 (idle): "AnyIO worker thread" NtReadFile / ReadFile ...MainThreadis alwaysactive/active+gil; both worker threads areidlein
every sample, parked inReadFile. Nothing of the SDK or of application code
sits above the loop. Note that these are the same frames a healthy idle server
shows — I have a pre-sleep baseline with an identical stack whereMainThread
reads(idle). The discriminator is thread state and CPU, not the frames.2. Selector loop — the spin survives
I then forced
WindowsSelectorEventLoopPolicyfor this interpreter. Smoke test
was clean (initializeOK, 22 tools, clean exit on stdin close). After two more
suspend/resume cycles the spin came back, at 94% and 162% of a core.But now the samples rotate through this:
sample 1: _run (asyncio\events.py:89) _run_once (asyncio\base_events.py:2060) sample 2: _select (selectors.py:305) select (selectors.py:314) _run_once (asyncio\base_events.py:2022) sample 3: _read_from_self (asyncio\selector_events.py:132) _run (asyncio\events.py:89) _run_once (asyncio\base_events.py:2060) sample 4: _add_callback (asyncio\base_events.py:1968) _process_events (asyncio\selector_events.py:754) _run_once (asyncio\base_events.py:2023)That is the loop's self-pipe firing continuously:
select()returns
immediately,_process_eventsqueues_read_from_self,_runexecutes it,
repeat.3. What ties the two together
Both loops build their self-wakeup out of the same thing:
# selector_events.py:118 and proactor_events.py:783 - identical def _make_self_pipe(self): # A self-socket, really. :-) self._ssock, self._csock = socket.socketpair()
On Windows
socket.socketpair()is emulated over a loopback TCP connection. My
hypothesis is that this connection does not survive S3, and comes back with the
peer closed. Neither loop treats a zero-length read on the self-pipe as "this
pipe is broken":- selector: a closed socket is permanently readable, so
select()returns
immediately forever._read_from_selfdoesrecv()→b''→break, and
the next iteration starts over. - proactor:
_loop_self_readingre-arms with
f.add_done_callback(self._loop_self_reading)wheneverf.result()does not
raise. Arecvon a closed socket completes successfully withb'', so it
re-arms immediately and completes immediately, forever — which is exactly the
GetQueuedCompletionStatusreturning with no delay that run 1 shows.
Note the asymmetry that makes this quiet: on the proactor path an error would
be caught and reported viacall_exception_handlerwithout re-arming. It is the
clean-EOF case that loops silently.4. Consequence for the proposed lever
This is the awkward part: if the mechanism above is right, then "use a selector
loop for stdio on Windows" would not fix this. I switched loops and kept the
spin. The problem does not look specific to IOCP; it looks like the self-pipe
socketpair, which both loops share.That also suggests this is more a CPython asyncio issue than an SDK one — a
self-pipe that reads EOF is broken by definition and could be rebuilt or raise,
rather than being polled forever. Whether the SDK wants to carry a workaround in
the meantime (supervise the loop, or detect a spinning self-pipe and exit so the
host can respawn) is your call. Given how it presents — no error, no log line,
just a pinned core and a server that has silently stopped answering — even
detect-and-exit would be a real improvement over the current failure mode.5. Caveats
- One machine, one Windows build (10.0.19045), uv-managed CPython 3.13.14.
- That the socketpair specifically dies across S3 is inference. I have the
before/after stacks and the code paths, but I have not directly inspected the
socket state after resume. If you want that, tell me what to capture and I
will get it on the next cycle. - I would not read anything into the
160%figure: only one thread is ever
active in the dumps, and one thread cannot exceed one core, so that is most
likely CPU accounting across the resume boundary. - Both server instances spin every time, including the one the host is actively
connected to.
6. Impact
Killing the spinning process frees the core but does not restore service — the
host reports the server as disconnected and does not respawn it, so it needs a
restart of the host application.I have full dumps for all four runs (proactor and selector, four samples each)
plus the healthy baseline, and a reproduction that takes about a minute. Happy
to attach them, run something more targeted, or test a patch.- selector: a closed socket is permanently readable, so
Two things from a fresh incident today, one of which corrects my previous
comment.1. Retraction: the spin does not need a suspend/resume cycle
Everything I posted so far framed this as a post-S3 failure. That framing is
wrong. Today's spin hit a machine that never slept:07:49:54 boot 07:58:45 hass server starts, initialize OK, 22 tools 09:02 .. 11:37:52 watchdog samples 4 processes every 5 min - all 0% ~11:40 (see 4 below) 11:42:53 RUNAWAY pid 26736 - 100% of one core over 15s 11:43:04 RUNAWAY pid 26672 - 173% of one core over 15s- no
Kernel-Power42 / 107 / 131 in the System log for the whole session
(the provider is alive — the boot-time 172 is there) powercfg /lastwake→Wake History Count - 0- uptime continuous from 07:49:54; the host's own processes date from 07:58
- for completeness: this box is S3-capable, and its S3 idle timeout is 2 hours,
its display timeout 5 minutes. What looked like sleep was the monitors
switching off.
So whatever kills the self-pipe, S3 is one way to get there and not the only
one. If you were scoping a fix or a test around resume specifically, that scope
is too narrow.2. The proactor self-pipe loop, observed rather than inferred
My previous comment derived the proactor half of the mechanism from reading
_loop_self_reading; the actual samples only ever showed
GetQueuedCompletionStatus/_poll, frames a healthy idle server shows too.This run was on a stock proactor (I had removed the forced selector policy), and
8 samples across the two processes rotate through this:sample: _loop_self_reading (asyncio\proactor_events.py:802) <- re-arm recv (asyncio\windows_events.py:488) sample: _loop_self_reading (asyncio\proactor_events.py:802) recv (asyncio\windows_events.py:494) _register (asyncio\windows_events.py:724) __init__ (asyncio\windows_events.py:56) sample: _loop_self_reading (asyncio\proactor_events.py:816) <- re-schedule sample: _poll (asyncio\windows_events.py:812) set_result (asyncio\windows_events.py:93) call_soon (asyncio\base_events.py:833) sample: _run_once (asyncio\base_events.py:2044)Those two line numbers are the whole cycle, and they are the two ends of it:
802: f = self._proactor.recv(self._ssock, 4096) # re-arm ... 808: except BaseException as exc: 809: self.call_exception_handler({...}) # not taken 814: else: 816: f.add_done_callback(self._loop_self_reading) # re-schedule
Re-arm → register overlapped → immediate completion →
set_result→
call_soon→ callback → re-arm, with theelsebranch at 814 confirming the
asymmetry: nothing raised, socall_exception_handleris never reached and the
callback re-schedules itself. Same cycle the selector run showed as
_select/_process_events/_read_from_self, now on the IOCP path with
_loop_self_readingnamed in the stack. Nothing of the SDK or of application
code appears above the loop in any sample.3. What a dead socketpair looks like in the TCP table
Since you may want to check this from your side, here is the signature, measured
on a synthetic pair rather than on a live spin:shutdown(SHUT_WR)on one end while both fds stay open →recv()on the
survivor returnsb''. That is the clean-EOF case that re-arms silently
instead of going throughcall_exception_handler.- both rows then disappear from the TCP table within a couple of seconds,
even though neither fd was closed. What lingers is a bare
0.0.0.0:<port> Boundremnant.
Practical consequence: for a spin that has been running for minutes, do not
expect to find a half-closed pair. Expect the loopback rows to be absent
while the loop still polls them. I have extended my watchdog to capture the
process's TCP sockets at detection time, next to the py-spy samples, so the next
incident should say directly whether the socketpair rows are still there. If
there is something else you would rather have captured in that window, say so
and I will add it.4. Candidate trigger — reported, not claimed
The spin began somewhere in 11:37:52 → ~11:42:40. The only notable system event
in that window is a Hyper-V virtual switch NIC being torn down and rebuilt
(ROOT\VMS_MP\0001deleted → configured → started, plus a network profile
re-identification) at 11:40:28–34. A TCP/IP-stack-level event would explain a
dead loopback socketpair as well as S3 does.But three byte-for-byte identical rebuilds earlier the same day (09:20:48,
09:36:58, 09:47:59) did not spin these processes, which were alive and at 0%
throughout. So it is not sufficient on its own, and I am not putting it forward
as the cause — either it needs a coincidence I have not identified, or the
trigger is elsewhere and unlogged.5. Failure mode, restated because today made it sharper
The host's MCP log for this server contains nothing at all between
07:58:49(startup,tools/listanswered) and11:43:13, which is when my
watchdog killed the process and the host noticed:09:43:13Z [hass] [info] Server transport closed 09:43:13Z [hass] [error] Server disconnected.Three and a half hours of healthy service, then a pinned core and silence. No
error, no warning, no log line on either side, and the host does not respawn —
it takes a restart of the host application. Both server instances spun again,
including the one actively connected.Environment as before: Windows 10.0.19045, uv-managed CPython 3.13.14, one
machine. Full dumps for today's two processes available on request.- no
Thanks for the dumps — the line numbers were enough to pin this down.
Root cause: CPython, not the SDK.
- Tracked upstream as asyncio: ProactorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF python/cpython#156333 (
ProactorEventLoop) and asyncio: SelectorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF python/cpython#156344 (SelectorEventLoop), with fix PRs gh-156333: rebuild the proactor self-pipe on EOF instead of busy-looping python/cpython#156343 / gh-156344: rebuild the selector self-pipe on EOF instead of busy-looping python/cpython#156345 open. - asyncio's self-pipe on Windows is an emulated 127.0.0.1
socketpair(). When its read end hits EOF:- the proactor loop re-arms a
recvthat completes immediately, forever; - the selector loop leaves a permanently-readable socket registered, so
select()never blocks; - nothing is logged either way.
- the proactor loop re-arms a
- Same code from 3.10 through
main, which is why every idle asyncio process on the box goes at once.
What I confirmed on a Windows runner (3.13/3.14,
mcp1.29.0 andmain):- Forcing EOF on the pair inside a running stdio server reproduces your exact py-spy stack at 99% of a core, on both loop types.
- While spinning, the server still answered
tools/callinstantly and still exited on stdin close.- So the cost is the pinned core, not lost service. Worth trying a tool call next time before killing it.
- A forcible teardown (what the TCP stack or a VPN "kill connections" does) looks different:
- one logged
ConnectionResetError, then a loop that goes deaf rather than spinning. - Your silent spin therefore means something on your machine is closing loopback connections gracefully — typically an in-path filter (VPN client, AV web shield, OEM network "optimizer").
- one logged
- If you want to chase the trigger:
- check whether lock → display off → wake alone reproduces it (the upstream reporter's trigger);
netsh winsock show catalogandnetsh wfp show stateshow what's sitting in the loopback path.
SDK position: nothing to change here. The fix belongs in asyncio, and patching private event-loop internals from a library isn't something we want to ship.
Unsupported stopgap until a fixed CPython ships (rebuilds the self-pipe on EOF/error; import before
mcp.run())"""Stopgap for CPython gh-156333 / gh-156344 on Windows: if the asyncio event loop's self-pipe socketpair reaches EOF or errors, rebuild it instead of busy-looping or going deaf to cross-thread wakeups. Import before the event loop is created (top of your entry script, or sitecustomize.py). Pokes asyncio privates; remove once CPython ships a fix.""" import signal import socket import sys import threading if sys.platform == "win32": from asyncio import proactor_events, selector_events def _fresh_pair(loop): ssock, csock = socket.socketpair() ssock.setblocking(False) csock.setblocking(False) if threading.current_thread() is threading.main_thread(): try: old_fd = loop._csock.fileno() if loop._csock is not None else -1 prev = signal.set_wakeup_fd(csock.fileno()) if prev not in (-1, old_fd): signal.set_wakeup_fd(prev) # not ours; put it back except (ValueError, OSError): pass return ssock, csock _orig_proactor = proactor_events.BaseProactorEventLoop._loop_self_reading def _proactor_loop_self_reading(self, f=None): if f is not None and not f.cancelled() and self._self_reading_future is f: exc = f.exception() if isinstance(exc, OSError) or (exc is None and not f.result()): ssock, csock = _fresh_pair(self) old_s, old_c = self._ssock, self._csock self._ssock, self._csock = ssock, csock old_s.close() if old_c is not None: old_c.close() self._self_reading_future = None f = None # arm a read on the new socket via the original code path sys.stderr.write("selfpipe_guard: event loop self-pipe was dead (%s); rebuilt\n" % ("EOF" if exc is None else repr(exc))) return _orig_proactor(self, f) proactor_events.BaseProactorEventLoop._loop_self_reading = _proactor_loop_self_reading def _selector_read_from_self(self): while True: try: data = self._ssock.recv(4096) if not data: raise ConnectionResetError("self-pipe EOF") self._process_self_data(data) except InterruptedError: continue except BlockingIOError: break except OSError as exc: ssock, csock = _fresh_pair(self) self._remove_reader(self._ssock.fileno()) old_s, old_c = self._ssock, self._csock self._ssock, self._csock = ssock, csock old_s.close() if old_c is not None: old_c.close() self._add_reader(ssock.fileno(), self._read_from_self) sys.stderr.write("selfpipe_guard: event loop self-pipe was dead (%r); rebuilt\n" % (exc,)) break selector_events.BaseSelectorEventLoop._read_from_self = _selector_read_from_self
I ran this on the runner against both the EOF and reset cases, both loop types, inside an
mcp1.29.0 stdio server: 0% CPU afterwards, tool calls and stdin-close shutdown keep working. No promises beyond that.Closing as upstream; happy to reopen if something SDK-specific turns up.
Reacted by pangi- Tracked upstream as asyncio: ProactorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF python/cpython#156333 (
Summary
On Windows, a
stdio_server()-based MCP server survives a system suspend/resumecycle as a live process but its event loop starts spinning: one thread burns a
full CPU core indefinitely while the server no longer serves any requests. The
process never exits and never recovers, so it keeps a core pinned until it is
killed manually.
I hit this with two server instances at once, each pinning a core (~2 cores lost
on an 8-core machine) for nearly two hours before I noticed.
Environment
mcpanyiocpython-3.13-windows-x86_64-none)mcp.run()stdio pathWindowsProactorEventLoopPolicy(Python default on Windows)What happens
initialize,tools/list,prompts/list,resources/listall complete.core, forever. No further MCP traffic is logged. The process does not exit.
Evidence that resume is the trigger
The host log recorded the resume precisely:
Accumulated CPU time of the two server processes, sampled at 20:55:11:
So both processes have been burning ~100% of a core continuously since the exact
second of resume, and were healthy before it. A 5-second live sample confirmed
100.3% of one core each at the time of measurement.
Thread-level detail
Only the main thread is hot; the worker threads are idle:
This is what makes me think the spin is in the event loop itself rather than in
the
anyio.wrap_file(...)stdin reader: if the reader thread were hitting EOF ina tight loop, the CPU would show up on a worker thread, not the main one.
Hypothesis (not confirmed with a debugger)
WindowsProactorEventLoopPolicydrives the loop off an IOCP. My guess is thatthe completion port or one of the registered handles is left in a state where
GetQueuedCompletionStatusreturns immediately after resume, soProactorEventLoop._poll()returns with no events and a zero timeout forever.I want to be clear that I have not attached a debugger to confirm this — the
correlation with resume and the main-thread-only CPU profile are what I actually
observed. If a maintainer can suggest what to capture (a
py-spy dumpon thespinning process, for instance) I am happy to reproduce and collect it.
Notes
connected to — so this is not only an orphaned-process problem.
established TCP socket), i.e. it was genuinely idle work-wise.
work is on-demand inside tool calls. Nothing in application code is looping.
Possible mitigations
Two things that might be worth considering, in increasing order of ambition:
suspend/resume would have saved me a couple of hours of process forensics.
the server is useless anyway — exiting cleanly would let the host respawn it,
and would turn a pinned core into a transparent reconnect.
As a local workaround I am considering forcing
WindowsSelectorEventLoopPolicyfor this process, since this particular server only needs sockets and threads
(no subprocess transports). I have not yet verified whether that actually avoids
the spin.