Repository navigation
Calls across stack chunks perform badly #142183
Description
Activity
- addedperformancePerformance or resource usagePerformance or resource usageinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)3.15bugs and security fixesbugs and security fixes
on Dec 2, 2025 If a frames calls a native function taking an array of arguments, then there will be pointers into the frame.
Is this still true? Previously,
PyObject_Vectorcallused thelocalsplusarray from the frame, but I think it now makes a copy to a local variable viaSTACKREFS_TO_PYOBJECTS.Lines 3748 to 3757 in 4172644
STACKREFS_TO_PYOBJECTS(arguments, total_args, args_o); if (CONVERSION_FAILED(args_o)) { DECREF_INPUTS(); ERROR_IF(true); } PyObject *res_o = PyObject_Vectorcall( callable_o, args_o, total_args | PY_VECTORCALL_ARGUMENTS_OFFSET, NULL); STACKREFS_TO_PYOBJECTS_CLEANUP(args_o); Is this still true?
It isn't currently true, but it might become true again with PyNI or similar, and I don't want to prevent that use case.
Plus, we often want to pass pointers to the stack to instrumentation and debug functions.- changed the title
[-]Calls across stack chunks perform badly.[/-][+]Calls across stack chunks perform badly[/+]on Dec 5, 2025 I believe I've run into a pathological case of this in the real world. A tight loop happening at just the right call depth lead to an explosion of mmap/munmap calls when upgarding from Python 3.12 to 3.14 (12x the number of allocs/deallocs), because the size of the frame changed by one pointer. Here's a reproducer Claude Code came up with:
import sys BRANCHES = 1000 # Number of sibling calls at the critical depth def level(d): """Recurse to fill the 16KB chunk, then fan out.""" # 20 locals to make each frame ~35-36 words (FRAME_SPECIALS_SIZE + locals + stack) _a=_b=_c=_d=_e=_f=_g=_h=_i=_j=None _k=_l=_m=_n=_o=_p=_q=_r=_s=_t=None if d > 0: level(d - 1) else: # At the critical depth, each burner() call crosses the chunk boundary, # triggering mmap+munmap. Without chunk caching, this is O(BRANCHES) # kernel calls instead of O(1). for _ in range(BRANCHES): burner() def burner(): """Leaf function that crosses the chunk boundary.""" _a=_b=_c=_d=_e=_f=_g=_h=_i=_j=None _k=_l=_m=_n=_o=_p=_q=_r=_s=_t=None return None def _framesize(code): """Estimate frame size in words (FRAME_SPECIALS_SIZE + locals + stack).""" specials = 10 if sys.version_info >= (3, 13) else 9 return (specials + code.co_nlocals + len(code.co_cellvars) + len(code.co_freevars) + code.co_stacksize) def find_thrash_depth(): """Auto-calibrate: find the depth where burner() just crosses the boundary. The 16KB root chunk holds ~2044 usable words. We compute frame sizes for level(), burner(), and the module code to find the exact depth where the chunk fills to within burner_sz of the top. Returns (depth, level_sz, burner_sz, base). """ try: import _testinternalcapi level_sz = _testinternalcapi.get_co_framesize(level.__code__) burner_sz = _testinternalcapi.get_co_framesize(burner.__code__) base = _testinternalcapi.get_co_framesize(sys._getframe(1).f_code) except (ImportError, AttributeError): level_sz = _framesize(level.__code__) burner_sz = _framesize(burner.__code__) base = _framesize(sys._getframe(1).f_code) usable = 2044 # words in root chunk for d in range(40, 80): remaining = usable - base - (d + 1) * level_sz if 0 < remaining < burner_sz: return d, level_sz, burner_sz, base return None, level_sz, burner_sz, base if len(sys.argv) > 1: depth = int(sys.argv[1]) print(f"Python {sys.version.split()[0]} | depth={depth} | branches={BRANCHES}") else: result, level_sz, burner_sz, base = find_thrash_depth() if result is not None: depth = result print(f"Python {sys.version.split()[0]} | depth={depth} | branches={BRANCHES}" f" | level_sz={level_sz} burner_sz={burner_sz} base={base}") else: # No thrashing depth: level_sz > burner_sz and alignment skips the window. # Fall back to depth 55 for comparison. depth = 55 print(f"Python {sys.version.split()[0]} | depth={depth} (fallback) | branches={BRANCHES}" f" | level_sz={level_sz} burner_sz={burner_sz} base={base}") print(f" No thrashing depth found for this build.") print(f"Run with: strace -e mmap,munmap -c python3 {__file__}") level(depth) print("Done.")
Wouldn't the easy fix be to cache one chunk on the tstate? (Basically a freelist of one chunk per thread.) That seems to fix the pathological case quite handily. It leaves calls that repeatedly alloc/dealloc two or more datastack chunks, but since they're 16Kb that feels like a lot of calls...
(Forgot to mention that this, specifically, made Python 3.14 30%+ slower than 3.12, so quite a big deal.)
What makes what slower, exactly? I assume you're not claiming that 3.14 is 30+% slower than 3.12 in general.
3.12 used the same chunked stack mechanism as 3.14, so you could just as easily construct an example where 3.12 is much slower than 3.14.
What makes what slower, exactly? I assume you're not claiming that 3.14 is 30+% slower than 3.12 in general.
3.12 used the same chunked stack mechanism as 3.14, so you could just as easily construct an example where 3.12 is much slower than 3.14.
Yes, it just showed up because the frame size is different for 3.12. But yes, a real world program, doing a reasonable amount of work (~30 seconds of text processing in my specific test case, often much more) was 37% slower with 3.14 than 3.12 because of this issue. Because it happened to hit the pathological case. Caching one datastack chunk made it go away entirely. Also, padding 'main' with a half dozen or so variables to push the top of the datastack just a little bit also made it go away entirely.
And yes, it reproduces with 3.12 as well if you happen to hit the right depth. Perhaps you could take a look at the reproducer. Hitting the pathological case is 30%+ slower in the reproducer as well:
% hyperfine 'python3.11 repro.py' 'python3.14 repro.py' Benchmark 1: python3.11 repro.py Time (mean ± σ): 10.1 ms ± 0.3 ms [User: 5.8 ms, System: 4.2 ms] Range (min … max): 9.7 ms … 11.2 ms 259 runs Benchmark 2: python3.14 repro.py Time (mean ± σ): 14.1 ms ± 0.3 ms [User: 9.4 ms, System: 4.4 ms] Range (min … max): 13.5 ms … 15.6 ms 187 runs Summary python3.11 repro.py ran 1.39 ± 0.05 times faster than python3.14 repro.pyYou can easily see whether it hits the pathological case by looking at the mmap/munmap count (see the included instructions.)
@Yhg1s
So, we sort of got lucky that this didn't show up in 3.12 or 3.13.
How far back do you want to backport the fix?If Hugo agrees I think it makes sense to backport it to 3.14 and 3.13, but not 3.12 (since it's in security-only mode). Backports will probably require moving the new pointer to the end of the tstate struct to avoid breaking debuggers and what not, but the size increase should be fine.
- added a commit that references this issue
on Mar 11, 2026 - added a commit that references this issue
on Mar 24, 2026 - added a commit that references this issue
on Apr 17, 2026 - added a commit that references this issue
on Apr 27, 2026 - added a commit that references this issue
on Jun 2, 2026 @markshannon Maybe just specialize call sites that hit the boundary? Naively:
_CHECK_STACK_SPACEmiss?CALL_PY_EXACT_ARGS->CALL_PY_EXACT_ARGS_CACHED_STACK,CALL_BOUND_METHOD_EXACT_ARGS->CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK.Maybe just specialize call sites that hit the boundary? Naively:
_CHECK_STACK_SPACEmiss?CALL_PY_EXACT_ARGS->CALL_PY_EXACT_ARGS_CACHED_STACK,CALL_BOUND_METHOD_EXACT_ARGS->CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK.Can you explain. I have no idea what you mean by "just specialize call sites that hit the boundary?"
@markshannon Yes. I don't have a patch ready, and this is a rough idea. I never worked in this code, so likely I'm missing something very obvious here.
That's what I mean by
_CHECK_STACK_SPACEmiss:Line 4600 in 416c346
DEOPT_IF(!_PyThreadState_HasStackSpace(tstate, code->co_framesize)); Instead, do something like (naive, just to sketch the idea):
if (!_PyThreadState_HasStackSpace(tstate, code->co_framesize)) { /* TODO: Check if remaining recursion */ uint8_t op = frame->instr_ptr->op.code; if (op == CALL_PY_EXACT_ARGS) { set_opcode(instr, CALL_PY_EXACT_ARGS_CACHED_STACK); /* an example; more likely: specialize_...() ? */ } else if (op == CALL_BOUND_METHOD_EXACT_ARGS) { set_opcode(instr, CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK); /* an example; more likely: specialize_...() ? */ } DEOPT(); }
Then, both
CALL_PY_EXACT_ARGS_CACHED_STACKandCALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACKcould either continue if there's stack space or try to usetstate->datastack_cached_chunk. (Probably by introducing_CHECK_CACHED_STACK_SPACEor similar?)This doesn't remove the edge but should reduce the deopts, and avoid changing this for everyone.
The Python stack is composed of a series of chunks of memory. These chunks are large, so we cross the boundaries infrequently and assume that it will never be on the fast path.
Unfortunately, it is possible to cross the boundary repeatedly if looping deep in the stack.
We should adjust the stack when crossing the boundary, to avoid doing so repeatedly.
Possible options include:
Both of these assume that we can move the frame. This is only possible if there are no pointers into the frame.
If a frames calls a native function taking an array of arguments, then there will be pointers into the frame.
So if we do any moving, we need to check for these pointers.
While we cannot move the frame, we can copy it, as long as the original remains. This is a bit wasteful, but shouldn't happen too often or waste too much space.
Linked PRs