Repository navigation
Assertion triggered in MiMalloc (free-threaded build) #153176
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Jul 6, 2026 - addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Jul 6, 2026 Can you please provide a reproducer?Wrong tab. Sorry!
Hey everyone, I can confirm this reproduces in a free-threaded debug build (
--with-pydebug --disable-gil):Assertion failed: (tstate == (_PyThreadStateImpl *)_PyThreadState_GET()), function _PyMem_mi_page_reclaimed, file obmalloc.c, line 215.TLDR, @weixlu's theory is correct.
_PyMem_mi_page_maybe_free()already documents and handles a page being freed on behalf of a different thread;_PyMem_mi_page_reclaimed()does the sametstate_from_heap()lookup but then wrongly asserts the result equals the current thread. Fix, matching the sibling function:_PyThreadStateImpl *tstate = tstate_from_heap(mi_page_heap(page)); - assert(tstate == (_PyThreadStateImpl *)_PyThreadState_GET()); + // We may be reclaiming a page belonging to a different thread + // during a stop-the-world event. Find the _PyThreadStateImpl for + // the page. page->retire_expire = 0; llist_insert_tail(&tstate->mimalloc.page_list, &page->qsbr_node);
The two lines that follow already use
tstate(the real page owner), not the current thread, so this only removes an incorrect diagnostic.One note for whoever reviews this:
I propose a tighter fix (assert(tstate == current || world_stopped)), but it's actually wrong. Tracingdestroy_interpreter→Py_EndInterpreter, the cross-thread heap cleanup happens in_PyThreadState_DeleteList, which runs after_PyEval_StartTheWorldAll()— it has its ownassert(!stoptheworld.world_stopped), becausePyThreadState_Clear()needs to run destructors and can't do that while the world is paused. Soworld_stoppedis alreadyfalseduring the actual cross-thread access; a conditional assert would still crash here.The real invariant is that the dying thread has been unlinked from
interp->threadsand marked as shutting down before its heap is touched — that bookkeeping isn't visible at the obmalloc.c layer, so a plain removal (as done in the sibling function) is the correct fix, not a tighter check.@zangjiucheng Hi, thanks for taking the time to look into and reproduce this issue. It looks like you did a very thorough investigation. Yes, I also think we can just remove this assertion.
Would you like to open a PR for the fix, or would you prefer me to do it? I’d be happy to open one as well :)
Reacted by Jiucheng(Oliver)I can make one PR for sure :), and thanks for your suggestion.
Reacted by Xiaowei LuI think the issue is in destroy_interpreter.
Removing the following lines fixes the crash:
diff --git a/Modules/_testinternalcapi.c b/Modules/_testinternalcapi.c index 166397a72da..e9950bb2324 100644 --- a/Modules/_testinternalcapi.c +++ b/Modules/_testinternalcapi.c @@ -2312,8 +2312,7 @@ destroy_interpreter(PyObject *self, PyObject *args, PyObject *kwargs) } t2 = PyThreadState_New(interp); prev = PyThreadState_Swap(t2); - PyThreadState_Clear(t1); - PyThreadState_Delete(t1); + // t1 is deliberately left alive; Py_EndInterpreter() must clean it up. Py_EndInterpreter(t2); PyThreadState_Swap(prev); }
@kumaraditya303 Thanks for pointing out! So t1 shouldn't be cleared here. No wonder sam gross thinks PyThreadState_Clear should only be called on the current thread.
cc @zangjiucheng , maybe this also helps with your PR
Thanks @kumaraditya303, @weixlu, that's the right call
I built a free-threaded debug interpreter and traced it. The backtrace shows the crash happens entirely inside
PyThreadState_Clear(t1), beforePy_EndInterpreterever runs:_PyMem_mi_page_reclaimed obmalloc.c:215 assert(tstate == current) <-- fails _mi_page_reclaim(heap=t1_heap) page.c:277 mi_segment_reclaim segment.c:1371 _mi_abandoned_reclaim_all(t1_heap) segment.c:1401 mi_heap_collect_ex(heap=t1_heap) heap.c:147 _mi_heap_collect_abandon heap.c:186 _PyThreadState_ClearMimallocHeaps(t1) pystate.c:3318 PyThreadState_Clear(t1) pystate.c:1902 destroy_interpreter _testinternalcapi.c:2315The addresses confirm it: the heap being reclaimed into (0x…138c88) sits inside t1's
_PyThreadStateImpl, while the current thread is t2.What happens:
destroy_interpreter(basic=True)callsPyThreadState_Clear(t1)whilet2is current (both were created on the main OS thread). That runs_PyThreadState_ClearMimallocHeaps(t1), and in a debug buildmi_heap_collect_ex()treatsMI_ABANDONas>= MI_FORCE, soforce_mainfires and_mi_abandoned_reclaim_all()reclaims segments back into t1's heaps._PyMem_mi_page_reclaimedthen seeststate_from_heap(page) == t1while the current thread ist2, and the assertion trips.So the assertion guards a real invariant: reclaim should only happen within the current thread's own heap. Unlike the sibling
_PyMem_mi_page_maybe_free, which legitimately runs on non-current heaps during_PyThreadState_DeleteList(gh-112532) but never reclaims, reclaim only reaches this point because the helper clears a non-current thread state — matching @colesbury's point thatPyThreadState_Clearshould only be called on the current thread.I'll update #153307 to drop the obmalloc change and fix the helper instead (leave
t1forPy_EndInterpreterto clean up, per your diff), and keep the regression test. Thanks both for digging into this.Sorry takes a long time to reply this :D
Reacted by Xiaowei Lu- added a commit that references this issue
on Jul 24, 2026
Bug report
Bug description:
I noticed that tearing down a subinterpreter's thread state aborts on an assertion.
This can be easily reproduced on a free-threaded debug build by running:
Some Investigation
This assertion was added in #145626 . Interestingly, the author @colesbury said in the PR:
Since
_PyMem_mi_page_maybe_free()doesn't assume the page owner to be current thread, maybe that's the same for_PyMem_mi_page_reclaimed?CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs