Skip to content

Calls across stack chunks perform badly #142183

Description

@markshannon

The Python stack is composed of a series of chunks of memory. These chunks are large, so we cross the boundaries infrequently and assume that it will never be on the fast path.

Unfortunately, it is possible to cross the boundary repeatedly if looping deep in the stack.

We should adjust the stack when crossing the boundary, to avoid doing so repeatedly.

Possible options include:

  1. Move the caller to the new chunk, so repeated calls do not cross the boundary.
  2. Instead of using a chunked stack, double the size of the stack and copy the old stack.

Both of these assume that we can move the frame. This is only possible if there are no pointers into the frame.
If a frames calls a native function taking an array of arguments, then there will be pointers into the frame.
So if we do any moving, we need to check for these pointers.

While we cannot move the frame, we can copy it, as long as the original remains. This is a bit wasteful, but shouldn't happen too often or waste too much space.

Linked PRs

Activity

  1. added
    performancePerformance or resource usage
    interpreter-core(Objects, Python, Grammar, and Parser dirs)
    3.15bugs and security fixes
    on Dec 2, 2025
  2. colesbury commented on Dec 3, 2025

    @colesbury
    Contributor

    If a frames calls a native function taking an array of arguments, then there will be pointers into the frame.

    Is this still true? Previously, PyObject_Vectorcall used the localsplus array from the frame, but I think it now makes a copy to a local variable via STACKREFS_TO_PYOBJECTS.

    cpython/Python/bytecodes.c

    Lines 3748 to 3757 in 4172644

    STACKREFS_TO_PYOBJECTS(arguments, total_args, args_o);
    if (CONVERSION_FAILED(args_o)) {
    DECREF_INPUTS();
    ERROR_IF(true);
    }
    PyObject *res_o = PyObject_Vectorcall(
    callable_o, args_o,
    total_args | PY_VECTORCALL_ARGUMENTS_OFFSET,
    NULL);
    STACKREFS_TO_PYOBJECTS_CLEANUP(args_o);

  3. markshannon commented on Dec 3, 2025

    @markshannon
    MemberAuthor

    Is this still true?

    It isn't currently true, but it might become true again with PyNI or similar, and I don't want to prevent that use case.
    Plus, we often want to pass pointers to the stack to instrumentation and debug functions.

  4. changed the title [-]Calls across stack chunks perform badly.[/-] [+]Calls across stack chunks perform badly[/+] on Dec 5, 2025
  5. Yhg1s commented on Mar 9, 2026

    @Yhg1s
    Member

    I believe I've run into a pathological case of this in the real world. A tight loop happening at just the right call depth lead to an explosion of mmap/munmap calls when upgarding from Python 3.12 to 3.14 (12x the number of allocs/deallocs), because the size of the frame changed by one pointer. Here's a reproducer Claude Code came up with:

    import sys
    
    BRANCHES = 1000  # Number of sibling calls at the critical depth
    
    def level(d):
        """Recurse to fill the 16KB chunk, then fan out."""
        # 20 locals to make each frame ~35-36 words (FRAME_SPECIALS_SIZE + locals + stack)
        _a=_b=_c=_d=_e=_f=_g=_h=_i=_j=None
        _k=_l=_m=_n=_o=_p=_q=_r=_s=_t=None
        if d > 0:
            level(d - 1)
        else:
            # At the critical depth, each burner() call crosses the chunk boundary,
            # triggering mmap+munmap. Without chunk caching, this is O(BRANCHES)
            # kernel calls instead of O(1).
            for _ in range(BRANCHES):
                burner()
    
    def burner():
        """Leaf function that crosses the chunk boundary."""
        _a=_b=_c=_d=_e=_f=_g=_h=_i=_j=None
        _k=_l=_m=_n=_o=_p=_q=_r=_s=_t=None
        return None
    
    def _framesize(code):
        """Estimate frame size in words (FRAME_SPECIALS_SIZE + locals + stack)."""
        specials = 10 if sys.version_info >= (3, 13) else 9
        return (specials + code.co_nlocals
                + len(code.co_cellvars) + len(code.co_freevars)
                + code.co_stacksize)
    
    def find_thrash_depth():
        """Auto-calibrate: find the depth where burner() just crosses the boundary.
    
        The 16KB root chunk holds ~2044 usable words. We compute frame sizes for
        level(), burner(), and the module code to find the exact depth where the
        chunk fills to within burner_sz of the top.
    
        Returns (depth, level_sz, burner_sz, base).
        """
        try:
            import _testinternalcapi
            level_sz = _testinternalcapi.get_co_framesize(level.__code__)
            burner_sz = _testinternalcapi.get_co_framesize(burner.__code__)
            base = _testinternalcapi.get_co_framesize(sys._getframe(1).f_code)
        except (ImportError, AttributeError):
            level_sz = _framesize(level.__code__)
            burner_sz = _framesize(burner.__code__)
            base = _framesize(sys._getframe(1).f_code)
    
        usable = 2044  # words in root chunk
    
        for d in range(40, 80):
            remaining = usable - base - (d + 1) * level_sz
            if 0 < remaining < burner_sz:
                return d, level_sz, burner_sz, base
    
        return None, level_sz, burner_sz, base
    
    if len(sys.argv) > 1:
        depth = int(sys.argv[1])
        print(f"Python {sys.version.split()[0]} | depth={depth} | branches={BRANCHES}")
    else:
        result, level_sz, burner_sz, base = find_thrash_depth()
        if result is not None:
            depth = result
            print(f"Python {sys.version.split()[0]} | depth={depth} | branches={BRANCHES}"
                  f" | level_sz={level_sz} burner_sz={burner_sz} base={base}")
        else:
            # No thrashing depth: level_sz > burner_sz and alignment skips the window.
            # Fall back to depth 55 for comparison.
            depth = 55
            print(f"Python {sys.version.split()[0]} | depth={depth} (fallback) | branches={BRANCHES}"
                  f" | level_sz={level_sz} burner_sz={burner_sz} base={base}")
            print(f"  No thrashing depth found for this build.")
    
    print(f"Run with: strace -e mmap,munmap -c python3 {__file__}")
    level(depth)
    print("Done.")

    Wouldn't the easy fix be to cache one chunk on the tstate? (Basically a freelist of one chunk per thread.) That seems to fix the pathological case quite handily. It leaves calls that repeatedly alloc/dealloc two or more datastack chunks, but since they're 16Kb that feels like a lot of calls...

  6. Yhg1s commented on Mar 9, 2026

    @Yhg1s
    Member

    (Forgot to mention that this, specifically, made Python 3.14 30%+ slower than 3.12, so quite a big deal.)

  7. markshannon commented on Mar 10, 2026

    @markshannon
    MemberAuthor

    What makes what slower, exactly? I assume you're not claiming that 3.14 is 30+% slower than 3.12 in general.

    3.12 used the same chunked stack mechanism as 3.14, so you could just as easily construct an example where 3.12 is much slower than 3.14.

  8. Yhg1s commented on Mar 10, 2026

    @Yhg1s
    Member

    What makes what slower, exactly? I assume you're not claiming that 3.14 is 30+% slower than 3.12 in general.

    3.12 used the same chunked stack mechanism as 3.14, so you could just as easily construct an example where 3.12 is much slower than 3.14.

    Yes, it just showed up because the frame size is different for 3.12. But yes, a real world program, doing a reasonable amount of work (~30 seconds of text processing in my specific test case, often much more) was 37% slower with 3.14 than 3.12 because of this issue. Because it happened to hit the pathological case. Caching one datastack chunk made it go away entirely. Also, padding 'main' with a half dozen or so variables to push the top of the datastack just a little bit also made it go away entirely.

    And yes, it reproduces with 3.12 as well if you happen to hit the right depth. Perhaps you could take a look at the reproducer. Hitting the pathological case is 30%+ slower in the reproducer as well:

    % hyperfine 'python3.11 repro.py' 'python3.14 repro.py'
    Benchmark 1: python3.11 repro.py
      Time (mean ± σ):      10.1 ms ±   0.3 ms    [User: 5.8 ms, System: 4.2 ms]
      Range (min … max):     9.7 ms …  11.2 ms    259 runs
    
    Benchmark 2: python3.14 repro.py
      Time (mean ± σ):      14.1 ms ±   0.3 ms    [User: 9.4 ms, System: 4.4 ms]
      Range (min … max):    13.5 ms …  15.6 ms    187 runs
    
    Summary
      python3.11 repro.py ran
        1.39 ± 0.05 times faster than python3.14 repro.py
    

    You can easily see whether it hits the pathological case by looking at the mmap/munmap count (see the included instructions.)

  9. markshannon commented on Mar 11, 2026

    @markshannon
    MemberAuthor

    @Yhg1s
    So, we sort of got lucky that this didn't show up in 3.12 or 3.13.
    How far back do you want to backport the fix?

  10. Yhg1s commented on Mar 11, 2026

    @Yhg1s
    Member

    If Hugo agrees I think it makes sense to backport it to 3.14 and 3.13, but not 3.12 (since it's in security-only mode). Backports will probably require moving the new pointer to the end of the tstate struct to avoid breaking debuggers and what not, but the size increase should be fine.

  11. added a commit that references this issue on Mar 11, 2026
  12. added 2 commits that reference this issue on Mar 11, 2026
  13. added 2 commits that reference this issue on Mar 18, 2026
  14. added a commit that references this issue on Mar 24, 2026
  15. added a commit that references this issue on Apr 17, 2026
  16. added a commit that references this issue on Apr 25, 2026
  17. added a commit that references this issue on Apr 27, 2026
  18. added 3 commits that reference this issue on May 8, 2026
  19. added a commit that references this issue on Jun 2, 2026
  20. maurycy commented on Aug 3, 2026

    @maurycy
    Contributor

    @markshannon Maybe just specialize call sites that hit the boundary? Naively: _CHECK_STACK_SPACE miss? CALL_PY_EXACT_ARGS -> CALL_PY_EXACT_ARGS_CACHED_STACK, CALL_BOUND_METHOD_EXACT_ARGS -> CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK.

    cc @pablogsal @dpdani

  21. markshannon commented on Aug 4, 2026

    @markshannon
    MemberAuthor

    Maybe just specialize call sites that hit the boundary? Naively: _CHECK_STACK_SPACE miss? CALL_PY_EXACT_ARGS -> CALL_PY_EXACT_ARGS_CACHED_STACK, CALL_BOUND_METHOD_EXACT_ARGS -> CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK.

    Can you explain. I have no idea what you mean by "just specialize call sites that hit the boundary?"

  22. maurycy commented on Aug 4, 2026

    @maurycy
    Contributor

    @markshannon Yes. I don't have a patch ready, and this is a rough idea. I never worked in this code, so likely I'm missing something very obvious here.

    That's what I mean by _CHECK_STACK_SPACE miss:

    DEOPT_IF(!_PyThreadState_HasStackSpace(tstate, code->co_framesize));

    Instead, do something like (naive, just to sketch the idea):

    if (!_PyThreadState_HasStackSpace(tstate, code->co_framesize)) {
      /* TODO: Check if remaining recursion */
    
      uint8_t op = frame->instr_ptr->op.code;
    
      if (op == CALL_PY_EXACT_ARGS) {
        set_opcode(instr, CALL_PY_EXACT_ARGS_CACHED_STACK); /* an example; more likely: specialize_...() ? */
      } else if (op == CALL_BOUND_METHOD_EXACT_ARGS) {
        set_opcode(instr, CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK); /* an example; more likely: specialize_...() ? */
      }
    
      DEOPT();
    }

    Then, both CALL_PY_EXACT_ARGS_CACHED_STACK and CALL_BOUND_METHOD_EXACT_ARGS_CACHED_STACK could either continue if there's stack space or try to use tstate->datastack_cached_chunk. (Probably by introducing _CHECK_CACHED_STACK_SPACE or similar?)

    This doesn't remove the edge but should reduce the deopts, and avoid changing this for everyone.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.15bugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usage

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions