Repository navigation
functools.lru_cache critical section locks self but _LockHeld dict APIs assert self->cache is locked #148180
Description
Activity
- addedextension-modulesC modules in the Modules dirC modules in the Modules dirtype-crashA hard crash of the interpreter, possibly with a core dumpA hard crash of the interpreter, possibly with a core dump
on Apr 6, 2026 I don't think we should fuzz for
gc.get_objectsorgc.get_referrersbugs.Yes, I avoid doing that.
This was found when trying to implement free threading in a C extension, the code looked wrong, but the only way we found to trigger it is with
gc.get_objectsorgc.get_referrers.Should I close this, since the bug isn't triggerable from other paths?
Ah, okay. I jumped to that conclusion.
I'm not sure what we can do about it, but let's leave this open for now. I'm open to suggestions on how we should handle this.
What's your thoughts on the
Py_BEGIN_CRITICAL_SECTION2(self, self->cache);workaround suggested by the LLM? Nothing else should be locking the cache so it should be painless I think?For context, I hit this during zhuyifei1999/guppy3#52 (comment). This is a heap-traversing memory debugging tool I maintain, that answers the question "what objects are there and why are they not being freed". Because I need to track objects in order to look at them more closely, the library obtains refcounts to nearly everything in the heap after a heap traversal, which includes this
functools.lru_cache->cachedict.And because I had a refcounting bug in my library while porting to freethreading, I popped open a
Py_DEBUGbuild to find what the refcounting bug was, only for Python to hit this bug and crash.I hit what looks like the same root cause from an ordinary,
gc-free path, which may be relevant to the "should we close this since it's only triggerable viagc.get_objects/get_referrers?" question above.Concurrent use of a single
functools.lru_cacheon a free-threaded build can make the cache exceed itsmaxsize— nogcintrospection involved, just normal lookups under eviction churn. The over-count is a real, persistent structural state (read quiescently after all threads join viacache_info().currsize, which isPyDict_GET_SIZE(self->cache)), not a transient counter skew.Reproducer (pure stdlib, no
gccalls) -- on Mac / ARM64import functools, threading, time def trial(nthreads=24, seconds=5, maxsize=128, keyspace=8192): @functools.lru_cache(maxsize=maxsize) def f(k): return k * k stop = False def hammer(seed): x = seed | 1 # per-thread PRNG, no shared state while not stop: x = (x * 1103515245 + 12345) & 0x7fffffff f(x % keyspace) # working set >> maxsize => constant eviction ts = [threading.Thread(target=hammer, args=(i,)) for i in range(nthreads)] for t in ts: t.start() time.sleep(seconds) stop = True for t in ts: t.join() return f.cache_info() # quiescent read; currsize == PyDict_GET_SIZE for i in range(40): info = trial() print(i, info) if info.currsize > 128: print("OVER maxsize:", info) break
On a free-threaded arm64 build this prints e.g.
CacheInfo(hits=..., misses=..., maxsize=128, currsize=129)within a few trials.Platform / GIL matrix (Python 3.13.13, same machine)
build result arm64, free-threaded ( PYTHON_GIL=0)~3/40 trials reach currsize=129(maxsize=128)arm64, GIL re-enabled ( PYTHON_GIL=1)0/40 x86-64, free-threaded 0/40 (possibly masked by stronger memory ordering; I haven't confirmed the race is absent vs. merely hidden) So it's a free-threading bug (the GIL-on control is clean), reproducible without any heap-introspection API.
Where it surfaces
In a stock build instrumented only with an
fprintfafter eachlru_cache_append_link(no--with-pydebug, to avoid perturbing timing), the size first passesmaxsizeat the evict-path append inbounded_lru_cache_wrapper(Modules/_functoolsmodule.c, ~L1205 in 3.13.13); a "grow branch taken while already full" probe never fired. That's consistent with two wrapper invocations interleaving whileself's critical section is suspended (e.g. across thePyObject_Callto the user function), making the size-check-then-insert non-atomic — i.e. the same "wrapper guardsself, notself->cache" issue described here, but observable as amaxsizeviolation rather than only the_LockHeldassertion.I'm not certain whether this is identical to the assertion failure above or a sibling of it (and whether the
Py_BEGIN_CRITICAL_SECTION2(self, self->cache)suggestion would close this particular TOCTOU) — you'd know far better than I would. Posting it mainly as evidence that the defect is reachable from normal concurrent use with a user-visible consequence.- added a commit that references this issue
on Jul 16, 2026
Crash report
What happened?
On a debug free-threaded build,
functools.lru_cachecrashes with a critical section assertion failure whengc.get_objects()is called between cache uses.bounded_lru_cache_wrapper(line 1498) acquiresPy_BEGIN_CRITICAL_SECTION(self), locking the lru_cache object's mutex. It then callsbounded_lru_cache_get_lock_held(line 1325) which calls_PyDict_GetItemRef_KnownHash_LockHeldonself->cache. The_LockHeldvariant asserts that the dict's per-object mutex is held (_Py_CRITICAL_SECTION_ASSERT_OBJECT_LOCKED(mp)atdictobject.c:1254), but onlyself's mutex is held —selfandself->cachehave different mutexes.In 3.15 we have this abort message:
This was introduced in gh-131757 (PR #131758), which split the critical section to allow the cached function to execute concurrently. The same pattern was extended in gh-132641 (PR #133787) which added more
_LockHeldcalls following the same assumption.On non-debug builds the assertion is compiled out and the code happens to work in practice (no actual contention in the reproducer), but the lock contract is violated.
Reproducer
Requires a debug free-threaded build (e.g.,
--with-pydebug --disable-gil).Backtrace on 3.14 (trimmed):
Backtrace on 3.15:
Possible fix
Either lock both objects:
Or use the regular (self-locking) dict APIs instead of
_LockHeldvariants for the get path.The initial crash and reproducer were found by @zhuyifei1999. Root cause analysis and issue draft done with assistance from Claude Code.
CPython versions tested on:
3.15, CPython main branch, 3.14
Operating systems tested on:
Linux
Output from running 'python -VV' on the command line:
Python 3.15.0a7+ free-threading build (heads/main:4d0e8ee649c, Mar 29 2026, 23:20:29) [Clang 21.1.2 (2ubuntu6)]