Skip to content

stdio_server busy-loops at 100% CPU on Windows after the host suspends and resumes #3411

Description

@pangi

Summary

On Windows, a stdio_server()-based MCP server survives a system suspend/resume
cycle as a live process but its event loop starts spinning: one thread burns a
full CPU core indefinitely while the server no longer serves any requests. The
process never exits and never recovers, so it keeps a core pinned until it is
killed manually.

I hit this with two server instances at once, each pinning a core (~2 cores lost
on an 8-core machine) for nearly two hours before I noticed.

Environment

mcp 1.29.0
anyio 4.14.2
Python 3.13.14 (uv-managed CPython, cpython-3.13-windows-x86_64-none)
OS Windows 10 Pro N, build 19045
Host Claude Desktop 1.40609.0.0 (stdio transport)
Server a FastMCP server using the default mcp.run() stdio path
Event loop policy WindowsProactorEventLoopPolicy (Python default on Windows)

What happens

  1. The host launches the server over stdio. It initializes normally —
    initialize, tools/list, prompts/list, resources/list all complete.
  2. The machine goes to sleep.
  3. The machine wakes.
  4. From the moment of resume, the server's main thread spins at 100% of one
    core, forever. No further MCP traffic is logged. The process does not exit.

Evidence that resume is the trigger

The host log recorded the resume precisely:

19:02:32 [warn] [event-loop-stall] main process blocked for 8065382ms [likely sleep: power_event]
19:02:32 [info] Timeout detector: tick arrived 8008s late — treating as system wake
19:02:39 [info] [event-loop-stall] OS resume

Accumulated CPU time of the two server processes, sampled at 20:55:11:

PID CPU time wall time since resume (19:02:39) ratio
A 6719 s 6752 s 99.5 %
B 6625 s 6752 s 98.1 %

So both processes have been burning ~100% of a core continuously since the exact
second of resume, and were healthy before it. A 5-second live sample confirmed
100.3% of one core each at the time of measurement.

Thread-level detail

Only the main thread is hot; the worker threads are idle:

PID A
   Id   ThreadState  WaitReason   TotalCPU_s
 5284   Running                      6637.00
13812   Wait         UserRequest        0.00
 6312   Wait         Executive          0.00

This is what makes me think the spin is in the event loop itself rather than in
the anyio.wrap_file(...) stdin reader: if the reader thread were hitting EOF in
a tight loop, the CPU would show up on a worker thread, not the main one.

Hypothesis (not confirmed with a debugger)

WindowsProactorEventLoopPolicy drives the loop off an IOCP. My guess is that
the completion port or one of the registered handles is left in a state where
GetQueuedCompletionStatus returns immediately after resume, so
ProactorEventLoop._poll() returns with no events and a zero timeout forever.

I want to be clear that I have not attached a debugger to confirm this — the
correlation with resume and the main-thread-only CPU profile are what I actually
observed. If a maintainer can suggest what to capture (a py-spy dump on the
spinning process, for instance) I am happy to reproduce and collect it.

Notes

  • Both instances were affected, including the one the host was actively
    connected to — so this is not only an orphaned-process problem.
  • The server had no network connection open to its backend at the time (no
    established TCP socket), i.e. it was genuinely idle work-wise.
  • The server itself has no background tasks or polling loops of its own; all its
    work is on-demand inside tool calls. Nothing in application code is looping.

Possible mitigations

Two things that might be worth considering, in increasing order of ambition:

  1. Document it. Even a note that Windows stdio servers can spin after
    suspend/resume would have saved me a couple of hours of process forensics.
  2. Detect the dead pipe and exit. If stdin is unreadable/broken after resume,
    the server is useless anyway — exiting cleanly would let the host respawn it,
    and would turn a pinned core into a transparent reconnect.

As a local workaround I am considering forcing WindowsSelectorEventLoopPolicy
for this process, since this particular server only needs sockets and threads
(no subprocess transports). I have not yet verified whether that actually avoids
the spin.

Activity

  1. added
    v2Affects the v2 line (2.x on main)
    v1Affects the v1.x maintenance line
    on Aug 28, 2026
  2. anneheartrecord commented on Aug 29, 2026

    @anneheartrecord

    Poked at where a spin could even live in this path, since the SDK-side surface is small: stdio_server() reads stdin through anyio.wrap_file(...), which does plain blocking reads on a worker thread — the stdin handle is never registered with the event loop at all. That fits your thread profile exactly: the reader thread is parked in ReadFile (0 CPU), and in an otherwise idle FastMCP stdio server the only IOCP-registered handles are the proactor's own internals (its self-waking socket pair etc.). So the spin is almost certainly inside ProactorEventLoop/IocpProactor, below the SDK.

    That has an awkward consequence for your mitigation 2: "detect the dead pipe and exit" already works today for the case where the pipe really dies — the reader thread sees EOF, the reader task finishes, and the server shuts down cleanly. In your hang stdin isn't dead (the worker thread is still blocked on it), so an SDK-level dead-pipe check would never fire.

    Which makes your selector-policy experiment the most valuable next data point, because it swaps out exactly the suspect component. If WindowsSelectorEventLoopPolicy avoids the spin, the SDK has a real lever: a stdio server doesn't need the proactor unless user code spawns subprocesses, so the stdio run path could use a selector loop on Windows — or at minimum document that as the workaround. Changing the default is a maintainer call though, since it silently breaks asyncio.create_subprocess_* inside tool handlers.

    For capture: py-spy dump --native --pid <pid>, three or four samples a few seconds apart. If the main thread keeps showing IocpProactor._poll / proactor_events frames with nothing of yours above them, that pins the loop itself, and the native frames should show which handle keeps completing.

    If the maintainers settle on a direction here (selector loop for stdio, docs, or something else), happy to put the PR together — though I can't drive suspend/resume myself, so verification would lean on your repro.

  3. pangi commented on Aug 29, 2026

    @pangi
    Author

    Thanks — that narrows it usefully, and I checked the anyio side against the
    installed source rather than taking it on trust. You're right:

    # anyio 4.14.2, _core/_fileio.py:125
    async def readline(self) -> AnyStr:
        return await to_thread.run_sync(self._fp.readline, limiter=self._limiter)

    AsyncFile routes every read through to_thread.run_sync, so the stdin handle
    is never registered with the loop at all. And I take your point that mitigation 2
    is a dead end for this particular hang: stdin isn't dead here, so a dead-pipe
    check would never fire.

    Baseline capture (healthy, idle server)

    py-spy does work against the uv-managed CPython on Windows, so here is the
    idle baseline with native frames:

    Thread 12292 (idle): "MainThread"
        NtRemoveIoCompletion (ntdll.dll)
        GetQueuedCompletionStatus (KERNELBASE.dll)
        PyInit__overlapped (_overlapped.pyd)
        ...
        _poll (asyncio\windows_events.py:778)
        select (asyncio\windows_events.py:446)
        _run_once (asyncio\base_events.py:2022)
        run_forever (asyncio\base_events.py:683)
        run_until_complete (asyncio\base_events.py:712)
        run (asyncio\runners.py:119)
        run (anyio\_backends\_asyncio.py:2481)
        run (anyio\_core\_eventloop.py:83)
        run (mcp\server\fastmcp\server.py:299)
        main (run.py:7)
    
    Thread 19416 (idle): "AnyIO worker thread"
        NtReadFile (ntdll.dll)
        ReadFile (KERNELBASE.dll)
        read (ucrtbase.dll)
        ...
        run (anyio\_backends\_asyncio.py:1033)
    

    So the worker thread is parked in ReadFile exactly as you described, and the
    main thread sits in GetQueuedCompletionStatus under IocpProactor._poll.

    One caveat worth flagging before the spin dump arrives: a healthy idle
    server already shows _poll (windows_events.py:778) / select on the main
    thread. So the Python frames alone will not distinguish the hang from normal
    idle — the discriminators are py-spy's thread state (idle vs active+gil) and
    the CPU time. I mention it so the eventual spin dump doesn't get read as
    "same frames, therefore nothing to see".

    Capture is armed

    I have a watchdog running that samples CPU of the server processes; anything
    holding >80% of one core for 15s gets four py-spy dump --native samples three
    seconds apart written to a file, and is then killed. So the next suspend/resume
    that reproduces this should produce the dumps without me having to be at the
    machine. I've verified the capture path end to end against a deliberately
    spinning decoy process.

    On ordering

    I'm deliberately not applying the selector-policy workaround yet. Doing so
    would remove my ability to capture the proactor in the spinning state, and that
    dump seems like the more perishable of the two data points. My plan is:

    1. Wait for a reproduction, collect the native dumps (unattended, as above).
    2. Then switch this process to WindowsSelectorEventLoopPolicy via a
      sitecustomize.py in the tool's venv and run for a while across several
      suspend/resume cycles to see whether the spin recurs.

    Say the word if you'd rather have the selector result first and are content to
    treat the spin dump as optional — happy to flip the order.

    Caveats on my side

    This is a single machine, one Windows build, one Python (uv-managed CPython
    3.13.14 on Windows 10 19045). I haven't reproduced it on demand — it has
    happened on natural suspend/resume cycles, so "n" is small and my reproduction
    latency is however long until the next one. If there's a faster way to force the
    condition than actually sleeping the machine, I'm glad to try it.

    And yes — happy to test a patch or a PR branch against the real repro whenever
    you have something you'd like exercised.

  4. pangi commented on Aug 29, 2026

    @pangi
    Author

    Reproduced, captured unattended, and then the selector experiment came back with
    a result I did not expect: it does not fix the spin — but it makes the cause
    visible.
    I think the two runs together point at one root cause for both loops.

    It does not need a long sleep. A ~40 second sleep was enough:

    21:44:10  Kernel-Power 42   entering sleep
    21:44:48  Kernel-Power 131  resume
    21:45:26  watchdog: pid 26360 at 91% of one core
    21:45:36  watchdog: pid 4832  at 160% of one core
    

    1. Proactor loop (stock behaviour)

    Four py-spy dump --native samples per process, three seconds apart. All 8
    samples across both processes:

    Thread 12292 (active): "MainThread"
        NtRemoveIoCompletion (ntdll.dll)
        GetQueuedCompletionStatus (KERNELBASE.dll)
        PyInit__overlapped (_overlapped.pyd)
        ...
        _poll (asyncio\windows_events.py:778)
        select (asyncio\windows_events.py:446)
        _run_once (asyncio\base_events.py:2022)
        run_forever (asyncio\base_events.py:683)
    
    Thread 19416 (idle): "AnyIO worker thread"
        NtReadFile / ReadFile ...
    

    MainThread is always active/active+gil; both worker threads are idle in
    every sample, parked in ReadFile. Nothing of the SDK or of application code
    sits above the loop. Note that these are the same frames a healthy idle server
    shows
    — I have a pre-sleep baseline with an identical stack where MainThread
    reads (idle). The discriminator is thread state and CPU, not the frames.

    2. Selector loop — the spin survives

    I then forced WindowsSelectorEventLoopPolicy for this interpreter. Smoke test
    was clean (initialize OK, 22 tools, clean exit on stdin close). After two more
    suspend/resume cycles the spin came back, at 94% and 162% of a core.

    But now the samples rotate through this:

    sample 1:  _run (asyncio\events.py:89)
               _run_once (asyncio\base_events.py:2060)
    
    sample 2:  _select (selectors.py:305)
               select (selectors.py:314)
               _run_once (asyncio\base_events.py:2022)
    
    sample 3:  _read_from_self (asyncio\selector_events.py:132)
               _run (asyncio\events.py:89)
               _run_once (asyncio\base_events.py:2060)
    
    sample 4:  _add_callback (asyncio\base_events.py:1968)
               _process_events (asyncio\selector_events.py:754)
               _run_once (asyncio\base_events.py:2023)
    

    That is the loop's self-pipe firing continuously: select() returns
    immediately, _process_events queues _read_from_self, _run executes it,
    repeat.

    3. What ties the two together

    Both loops build their self-wakeup out of the same thing:

    # selector_events.py:118 and proactor_events.py:783 - identical
    def _make_self_pipe(self):
        # A self-socket, really. :-)
        self._ssock, self._csock = socket.socketpair()

    On Windows socket.socketpair() is emulated over a loopback TCP connection. My
    hypothesis is that this connection does not survive S3, and comes back with the
    peer closed. Neither loop treats a zero-length read on the self-pipe as "this
    pipe is broken":

    • selector: a closed socket is permanently readable, so select() returns
      immediately forever. _read_from_self does recv() → b'' → break, and
      the next iteration starts over.
    • proactor: _loop_self_reading re-arms with
      f.add_done_callback(self._loop_self_reading) whenever f.result() does not
      raise. A recv on a closed socket completes successfully with b'', so it
      re-arms immediately and completes immediately, forever — which is exactly the
      GetQueuedCompletionStatus returning with no delay that run 1 shows.

    Note the asymmetry that makes this quiet: on the proactor path an error would
    be caught and reported via call_exception_handler without re-arming. It is the
    clean-EOF case that loops silently.

    4. Consequence for the proposed lever

    This is the awkward part: if the mechanism above is right, then "use a selector
    loop for stdio on Windows" would not fix this. I switched loops and kept the
    spin. The problem does not look specific to IOCP; it looks like the self-pipe
    socketpair, which both loops share.

    That also suggests this is more a CPython asyncio issue than an SDK one — a
    self-pipe that reads EOF is broken by definition and could be rebuilt or raise,
    rather than being polled forever. Whether the SDK wants to carry a workaround in
    the meantime (supervise the loop, or detect a spinning self-pipe and exit so the
    host can respawn) is your call. Given how it presents — no error, no log line,
    just a pinned core and a server that has silently stopped answering — even
    detect-and-exit would be a real improvement over the current failure mode.

    5. Caveats

    • One machine, one Windows build (10.0.19045), uv-managed CPython 3.13.14.
    • That the socketpair specifically dies across S3 is inference. I have the
      before/after stacks and the code paths, but I have not directly inspected the
      socket state after resume. If you want that, tell me what to capture and I
      will get it on the next cycle.
    • I would not read anything into the 160% figure: only one thread is ever
      active in the dumps, and one thread cannot exceed one core, so that is most
      likely CPU accounting across the resume boundary.
    • Both server instances spin every time, including the one the host is actively
      connected to.

    6. Impact

    Killing the spinning process frees the core but does not restore service — the
    host reports the server as disconnected and does not respawn it, so it needs a
    restart of the host application.

    I have full dumps for all four runs (proactor and selector, four samples each)
    plus the healthy baseline, and a reproduction that takes about a minute. Happy
    to attach them, run something more targeted, or test a patch.

  5. pangi commented on Sep 1, 2026

    @pangi
    Author

    Two things from a fresh incident today, one of which corrects my previous
    comment
    .

    1. Retraction: the spin does not need a suspend/resume cycle

    Everything I posted so far framed this as a post-S3 failure. That framing is
    wrong. Today's spin hit a machine that never slept:

    07:49:54  boot
    07:58:45  hass server starts, initialize OK, 22 tools
    09:02 .. 11:37:52  watchdog samples 4 processes every 5 min - all 0%
    ~11:40    (see 4 below)
    11:42:53  RUNAWAY pid 26736 - 100% of one core over 15s
    11:43:04  RUNAWAY pid 26672 - 173% of one core over 15s
    
    • no Kernel-Power 42 / 107 / 131 in the System log for the whole session
      (the provider is alive — the boot-time 172 is there)
    • powercfg /lastwake → Wake History Count - 0
    • uptime continuous from 07:49:54; the host's own processes date from 07:58
    • for completeness: this box is S3-capable, and its S3 idle timeout is 2 hours,
      its display timeout 5 minutes. What looked like sleep was the monitors
      switching off.

    So whatever kills the self-pipe, S3 is one way to get there and not the only
    one. If you were scoping a fix or a test around resume specifically, that scope
    is too narrow.

    2. The proactor self-pipe loop, observed rather than inferred

    My previous comment derived the proactor half of the mechanism from reading
    _loop_self_reading; the actual samples only ever showed
    GetQueuedCompletionStatus / _poll, frames a healthy idle server shows too.

    This run was on a stock proactor (I had removed the forced selector policy), and
    8 samples across the two processes rotate through this:

    sample:  _loop_self_reading (asyncio\proactor_events.py:802)   <- re-arm
               recv             (asyncio\windows_events.py:488)
    sample:  _loop_self_reading (asyncio\proactor_events.py:802)
               recv             (asyncio\windows_events.py:494)
                 _register      (asyncio\windows_events.py:724)
                   __init__     (asyncio\windows_events.py:56)
    sample:  _loop_self_reading (asyncio\proactor_events.py:816)   <- re-schedule
    sample:  _poll              (asyncio\windows_events.py:812)
               set_result       (asyncio\windows_events.py:93)
                 call_soon      (asyncio\base_events.py:833)
    sample:  _run_once          (asyncio\base_events.py:2044)
    

    Those two line numbers are the whole cycle, and they are the two ends of it:

    802:  f = self._proactor.recv(self._ssock, 4096)   # re-arm
    ...
    808:  except BaseException as exc:
    809:      self.call_exception_handler({...})       # not taken
    814:  else:
    816:      f.add_done_callback(self._loop_self_reading)   # re-schedule

    Re-arm → register overlapped → immediate completion → set_result →
    call_soon → callback → re-arm, with the else branch at 814 confirming the
    asymmetry: nothing raised, so call_exception_handler is never reached and the
    callback re-schedules itself. Same cycle the selector run showed as
    _select / _process_events / _read_from_self, now on the IOCP path with
    _loop_self_reading named in the stack. Nothing of the SDK or of application
    code appears above the loop in any sample.

    3. What a dead socketpair looks like in the TCP table

    Since you may want to check this from your side, here is the signature, measured
    on a synthetic pair rather than on a live spin:

    • shutdown(SHUT_WR) on one end while both fds stay open → recv() on the
      survivor returns b''. That is the clean-EOF case that re-arms silently
      instead of going through call_exception_handler.
    • both rows then disappear from the TCP table within a couple of seconds,
      even though neither fd was closed. What lingers is a bare
      0.0.0.0:<port> Bound remnant.

    Practical consequence: for a spin that has been running for minutes, do not
    expect to find a half-closed pair. Expect the loopback rows to be absent
    while the loop still polls them. I have extended my watchdog to capture the
    process's TCP sockets at detection time, next to the py-spy samples, so the next
    incident should say directly whether the socketpair rows are still there. If
    there is something else you would rather have captured in that window, say so
    and I will add it.

    4. Candidate trigger — reported, not claimed

    The spin began somewhere in 11:37:52 → ~11:42:40. The only notable system event
    in that window is a Hyper-V virtual switch NIC being torn down and rebuilt
    (ROOT\VMS_MP\0001 deleted → configured → started, plus a network profile
    re-identification) at 11:40:28–34. A TCP/IP-stack-level event would explain a
    dead loopback socketpair as well as S3 does.

    But three byte-for-byte identical rebuilds earlier the same day (09:20:48,
    09:36:58, 09:47:59) did not spin these processes
    , which were alive and at 0%
    throughout. So it is not sufficient on its own, and I am not putting it forward
    as the cause — either it needs a coincidence I have not identified, or the
    trigger is elsewhere and unlogged.

    5. Failure mode, restated because today made it sharper

    The host's MCP log for this server contains nothing at all between
    07:58:49 (startup, tools/list answered) and 11:43:13, which is when my
    watchdog killed the process and the host noticed:

    09:43:13Z [hass] [info]  Server transport closed
    09:43:13Z [hass] [error] Server disconnected.
    

    Three and a half hours of healthy service, then a pinned core and silence. No
    error, no warning, no log line on either side, and the host does not respawn —
    it takes a restart of the host application. Both server instances spun again,
    including the one actively connected.

    Environment as before: Windows 10.0.19045, uv-managed CPython 3.13.14, one
    machine. Full dumps for today's two processes available on request.

  6. maxisbey commented on Sep 1, 2026

    @maxisbey
    Contributor

    Thanks for the dumps — the line numbers were enough to pin this down.

    Root cause: CPython, not the SDK.

    What I confirmed on a Windows runner (3.13/3.14, mcp 1.29.0 and main):

    • Forcing EOF on the pair inside a running stdio server reproduces your exact py-spy stack at 99% of a core, on both loop types.
    • While spinning, the server still answered tools/call instantly and still exited on stdin close.
      • So the cost is the pinned core, not lost service. Worth trying a tool call next time before killing it.
    • A forcible teardown (what the TCP stack or a VPN "kill connections" does) looks different:
      • one logged ConnectionResetError, then a loop that goes deaf rather than spinning.
      • Your silent spin therefore means something on your machine is closing loopback connections gracefully — typically an in-path filter (VPN client, AV web shield, OEM network "optimizer").
    • If you want to chase the trigger:
      • check whether lock → display off → wake alone reproduces it (the upstream reporter's trigger);
      • netsh winsock show catalog and netsh wfp show state show what's sitting in the loopback path.

    SDK position: nothing to change here. The fix belongs in asyncio, and patching private event-loop internals from a library isn't something we want to ship.

    Unsupported stopgap until a fixed CPython ships (rebuilds the self-pipe on EOF/error; import before mcp.run())
    """Stopgap for CPython gh-156333 / gh-156344 on Windows: if the asyncio event loop's
    self-pipe socketpair reaches EOF or errors, rebuild it instead of busy-looping or going
    deaf to cross-thread wakeups. Import before the event loop is created (top of your entry
    script, or sitecustomize.py). Pokes asyncio privates; remove once CPython ships a fix."""
    import signal
    import socket
    import sys
    import threading
    
    if sys.platform == "win32":
        from asyncio import proactor_events, selector_events
    
        def _fresh_pair(loop):
            ssock, csock = socket.socketpair()
            ssock.setblocking(False)
            csock.setblocking(False)
            if threading.current_thread() is threading.main_thread():
                try:
                    old_fd = loop._csock.fileno() if loop._csock is not None else -1
                    prev = signal.set_wakeup_fd(csock.fileno())
                    if prev not in (-1, old_fd):
                        signal.set_wakeup_fd(prev)  # not ours; put it back
                except (ValueError, OSError):
                    pass
            return ssock, csock
    
        _orig_proactor = proactor_events.BaseProactorEventLoop._loop_self_reading
    
        def _proactor_loop_self_reading(self, f=None):
            if f is not None and not f.cancelled() and self._self_reading_future is f:
                exc = f.exception()
                if isinstance(exc, OSError) or (exc is None and not f.result()):
                    ssock, csock = _fresh_pair(self)
                    old_s, old_c = self._ssock, self._csock
                    self._ssock, self._csock = ssock, csock
                    old_s.close()
                    if old_c is not None:
                        old_c.close()
                    self._self_reading_future = None
                    f = None  # arm a read on the new socket via the original code path
                    sys.stderr.write("selfpipe_guard: event loop self-pipe was dead (%s); rebuilt\n" % ("EOF" if exc is None else repr(exc)))
            return _orig_proactor(self, f)
    
        proactor_events.BaseProactorEventLoop._loop_self_reading = _proactor_loop_self_reading
    
        def _selector_read_from_self(self):
            while True:
                try:
                    data = self._ssock.recv(4096)
                    if not data:
                        raise ConnectionResetError("self-pipe EOF")
                    self._process_self_data(data)
                except InterruptedError:
                    continue
                except BlockingIOError:
                    break
                except OSError as exc:
                    ssock, csock = _fresh_pair(self)
                    self._remove_reader(self._ssock.fileno())
                    old_s, old_c = self._ssock, self._csock
                    self._ssock, self._csock = ssock, csock
                    old_s.close()
                    if old_c is not None:
                        old_c.close()
                    self._add_reader(ssock.fileno(), self._read_from_self)
                    sys.stderr.write("selfpipe_guard: event loop self-pipe was dead (%r); rebuilt\n" % (exc,))
                    break
    
        selector_events.BaseSelectorEventLoop._read_from_self = _selector_read_from_self

    I ran this on the runner against both the EOF and reset cases, both loop types, inside an mcp 1.29.0 stdio server: 0% CPU afterwards, tool calls and stdin-close shutdown keep working. No promises beyond that.

    Closing as upstream; happy to reopen if something SDK-specific turns up.

    AI Disclaimer

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    v1Affects the v1.x maintenance linev2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions