Skip to content

Assertion triggered in MiMalloc (free-threaded build) #153176

Description

@weixlu

Bug report

Bug description:

I noticed that tearing down a subinterpreter's thread state aborts on an assertion.

python: Objects/obmalloc.c:215: _PyMem_mi_page_reclaimed: Assertion `tstate == (_PyThreadStateImpl *)_PyThreadState_GET()' failed.
fish: Job 1, './python xxxxx' terminated by signal SIGABRT (Abort)

This can be easily reproduced on a free-threaded debug build by running:

import _testinternalcapi

interpid = _testinternalcapi.create_interpreter()
_testinternalcapi.destroy_interpreter(interpid, basic=True)

Some Investigation

This assertion was added in #145626 . Interestingly, the author @colesbury said in the PR:

In _PyMem_mi_page_maybe_free(), use the page's heap tld to find the correct thread state ...... so the current thread state is not necessarily the owner of the page.

Since _PyMem_mi_page_maybe_free() doesn't assume the page owner to be current thread, maybe that's the same for _PyMem_mi_page_reclaimed ?

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    on Jul 6, 2026
  2. StanFromIreland commented on Jul 6, 2026

    @StanFromIreland
    Member

    Can you please provide a reproducer?

    Wrong tab. Sorry!

  3. zangjiucheng commented on Jul 8, 2026

    @zangjiucheng
    Contributor

    Hey everyone, I can confirm this reproduces in a free-threaded debug build (--with-pydebug --disable-gil):

    Assertion failed: (tstate == (_PyThreadStateImpl *)_PyThreadState_GET()), function _PyMem_mi_page_reclaimed, file obmalloc.c, line 215.
    

    TLDR, @weixlu's theory is correct. _PyMem_mi_page_maybe_free() already documents and handles a page being freed on behalf of a different thread; _PyMem_mi_page_reclaimed() does the same tstate_from_heap() lookup but then wrongly asserts the result equals the current thread. Fix, matching the sibling function:

                 _PyThreadStateImpl *tstate = tstate_from_heap(mi_page_heap(page));
    -            assert(tstate == (_PyThreadStateImpl *)_PyThreadState_GET());
    +            // We may be reclaiming a page belonging to a different thread
    +            // during a stop-the-world event. Find the _PyThreadStateImpl for
    +            // the page.
                 page->retire_expire = 0;
                 llist_insert_tail(&tstate->mimalloc.page_list, &page->qsbr_node);

    The two lines that follow already use tstate (the real page owner), not the current thread, so this only removes an incorrect diagnostic.

    One note for whoever reviews this:
    I propose a tighter fix (assert(tstate == current || world_stopped)), but it's actually wrong. Tracing destroy_interpreter → Py_EndInterpreter, the cross-thread heap cleanup happens in _PyThreadState_DeleteList, which runs after _PyEval_StartTheWorldAll() — it has its own assert(!stoptheworld.world_stopped), because PyThreadState_Clear() needs to run destructors and can't do that while the world is paused. So world_stopped is already false during the actual cross-thread access; a conditional assert would still crash here.

    The real invariant is that the dying thread has been unlinked from interp->threads and marked as shutting down before its heap is touched — that bookkeeping isn't visible at the obmalloc.c layer, so a plain removal (as done in the sibling function) is the correct fix, not a tighter check.

  4. weixlu commented on Jul 8, 2026

    @weixlu
    ContributorAuthor

    @zangjiucheng Hi, thanks for taking the time to look into and reproduce this issue. It looks like you did a very thorough investigation. Yes, I also think we can just remove this assertion.

    Would you like to open a PR for the fix, or would you prefer me to do it? I’d be happy to open one as well :)

  5. zangjiucheng commented on Jul 8, 2026

    @zangjiucheng
    Contributor

    I can make one PR for sure :), and thanks for your suggestion.

  6. kumaraditya303 commented on Jul 22, 2026

    @kumaraditya303
    Contributor

    I think the issue is in destroy_interpreter.

    Removing the following lines fixes the crash:

    diff --git a/Modules/_testinternalcapi.c b/Modules/_testinternalcapi.c
    index 166397a72da..e9950bb2324 100644
    --- a/Modules/_testinternalcapi.c
    +++ b/Modules/_testinternalcapi.c
    @@ -2312,8 +2312,7 @@ destroy_interpreter(PyObject *self, PyObject *args, PyObject *kwargs)
             }
             t2 = PyThreadState_New(interp);
             prev = PyThreadState_Swap(t2);
    -        PyThreadState_Clear(t1);
    -        PyThreadState_Delete(t1);
    +        // t1 is deliberately left alive; Py_EndInterpreter() must clean it up.
             Py_EndInterpreter(t2);
             PyThreadState_Swap(prev);
         }
    
  7. weixlu commented on Jul 22, 2026

    @weixlu
    ContributorAuthor

    @kumaraditya303 Thanks for pointing out! So t1 shouldn't be cleared here. No wonder sam gross thinks PyThreadState_Clear should only be called on the current thread.

    cc @zangjiucheng , maybe this also helps with your PR

  8. zangjiucheng commented on Jul 22, 2026

    @zangjiucheng
    Contributor

    Thanks @kumaraditya303, @weixlu, that's the right call

    I built a free-threaded debug interpreter and traced it. The backtrace shows the crash happens entirely inside PyThreadState_Clear(t1), before Py_EndInterpreter ever runs:

    _PyMem_mi_page_reclaimed              obmalloc.c:215   assert(tstate == current)  <-- fails
    _mi_page_reclaim(heap=t1_heap)        page.c:277
    mi_segment_reclaim                    segment.c:1371
    _mi_abandoned_reclaim_all(t1_heap)    segment.c:1401
    mi_heap_collect_ex(heap=t1_heap)      heap.c:147
    _mi_heap_collect_abandon              heap.c:186
    _PyThreadState_ClearMimallocHeaps(t1) pystate.c:3318
    PyThreadState_Clear(t1)               pystate.c:1902
    destroy_interpreter                   _testinternalcapi.c:2315
    

    The addresses confirm it: the heap being reclaimed into (0x…138c88) sits inside t1's _PyThreadStateImpl, while the current thread is t2.

    What happens: destroy_interpreter(basic=True) calls PyThreadState_Clear(t1) while t2 is current (both were created on the main OS thread). That runs _PyThreadState_ClearMimallocHeaps(t1), and in a debug build mi_heap_collect_ex() treats MI_ABANDON as >= MI_FORCE, so force_main fires and _mi_abandoned_reclaim_all() reclaims segments back into t1's heaps. _PyMem_mi_page_reclaimed then sees tstate_from_heap(page) == t1 while the current thread is t2, and the assertion trips.

    So the assertion guards a real invariant: reclaim should only happen within the current thread's own heap. Unlike the sibling _PyMem_mi_page_maybe_free, which legitimately runs on non-current heaps during _PyThreadState_DeleteList (gh-112532) but never reclaims, reclaim only reaches this point because the helper clears a non-current thread state — matching @colesbury's point that PyThreadState_Clear should only be called on the current thread.

    I'll update #153307 to drop the obmalloc change and fix the helper instead (leave t1 for Py_EndInterpreter to clean up, per your diff), and keep the regression test. Thanks both for digging into this.

    Sorry takes a long time to reply this :D

  9. added a commit that references this issue on Jul 23, 2026
  10. added a commit that references this issue on Jul 24, 2026
  11. added a commit that references this issue on Aug 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions