Skip to content

Inconsistent .extend() behavior in bytearray #145300

Description

@tim-one

Bug report

Bug description:

Extending a bytearrary can behave very differently when extending via an iterable than when extending a list or array.array in the same way:

>>> from itertools import islice
>>> xs = [0, 1, 2]
>>> xs.extend(islice(xs, 12000))
>>> len(xs)
12003
>>> import array
>>> xs = array.array('Q', [0, 1, 2])
>>> xs.extend(islice(xs, 12000))
>>> len(xs)
12003

Tricky, but it's always worked this way, and is very convenient to extend a sequence with copies of itself. Note that the iterator picks up new elements of the sequence while they're being added.

bytearray doesn't work this way, though. It appears to capture the sequence's length just once at the start, and so can't do more than double the original length.

>>> xs = bytearray([0, 1, 2])
>>> xs.extend(islice(xs, 12000))
>>> len(xs)
6
>>> list(xs)
[0, 1, 2, 0, 1, 2]

Bumped into this when changing old code to switch from lists of small ints to bytearrays instead. Quite a head-scratcher to figure out what went wrong! ;-)

CPython versions tested on:

3.15, 3.14

Operating systems tested on:

No response

Linked PRs

Activity

  1. cmaloney commented on Feb 27, 2026

    @cmaloney
    Contributor

    That bytearray appears to capture the sequence's length just once at the start is right on the nose. Going through the implementation bytearray.extend doesn't extend the underlying bytearray until the iterator is exhausted. bytearray.extend when passed an iterable that doesn't support the buffer protocol starts by allocating a temporary bytearray then exhausts the iterator into that temporary resizing as needed. Only once it has exhausted the iterator does it extend the current object with the newly constructed temporary.

    List on the other hand (list_extend_iter_lock_held) continually resizes itself while doing the extend. In terms of which is the "right" behavior sequence.extend suggests the length 6 is the general sequence behavior pattern:

    >>> xs = [1,2,3]
    >>> xs[len(xs):len(xs)] = xs
    >>> xs
    [1, 2, 3, 1, 2, 3]
    

    The behavior has not changed for a long time in bytearray (since at least 2009 in my look at git blame), and the behavior of list looks to be of a similar vintage so I suspect this is better to document than to normalize. Code which assumes the behavior almost certainly exists and while standardizing would be nice it probably doesn't pass the PEP-387 test of "a large benefit to breakage ratio".

    Secondary bug to me: itertools.islice doesn't support __length_hint__ which means things are less efficient then they could be in the .extend

  2. tim-one commented on Feb 27, 2026

    @tim-one
    MemberAuthor

    Ya, it's possible (likely!) nothing can be changed here. Lists have always defined what happens if the length changes during iteration, and list.extend() follows that model. array.array followed suit. bytearray is the oddball. `coillections.deque' does a third thing:

    RuntimeError: deque mutated during iteration
    

    The docs:

    For the most part, this is the same as writing seq[len(seq):len(seq)] = **iterable.**
    

    isn't of much help, as it doesn't give a clue about what "the most part" means 😦

  3. added a commit that references this issue on Feb 27, 2026
  4. tim-one commented on Feb 28, 2026

    @tim-one
    MemberAuthor

    I wasn't hallucinating 👍: with the help of a bot, we traced down when language guaranteeing what happens when lists mutate during iteration was removed. It was in the Reference manual, but got purged when the iteration protocol was introduced.

    But it's far too late to change that behavior now. The docs definitely need updating. It's just incoherent now:

    xs.extend(xs) # doubles the length for a list
    xs.exitend(iter(xs)) # unbounded memory growth
    

    While it's fine to discourage relying on any specific behavior when containers mutate during iteration, it's a fact that we can't change how lists work without breaking tons of code. So we should own up to that. It was advertised behavior since day #1, and while we stopped advertising it, the implementation still continues to honor it. Nothing about it is accidental.

  5. added 2 commits that reference this issue on Feb 28, 2026
  6. added
    docsDocumentation in the Doc dir
    and removed
    type-bugAn unexpected behavior, bug, or error
    on Mar 2, 2026
  7. tim-one commented on Mar 2, 2026

    @tim-one
    MemberAuthor

    Changed label to "docs", since it seems exceedingly unlikely any behaviors will be changed. The docs should nevertheless reflect reality. More here:

    https://discuss.python.org/t/probably-not-a-pep-mutating-a-sequence-during-iteration-is-fuzzy/106341

  8. added a commit that references this issue on Mar 8, 2026
  9. picnixz commented on Mar 8, 2026

    @picnixz
    Member

    I'm definitely in favor of documenting it but I wonder whether this should be considered a CPython implementation detail or not and whether such details should be documented and how.

    Solution 1

    Let's document it as a CPython implementation detail but only where it matters. Have you checked whether this happens for other implementations like PyPy? The question is: what is the "default" behavior: should list and array.array be the expected way extend should happen or should bytearray be the "canonical" approach? what about xs.extend(iter(xs)) vs xs.extend(xs) for lists as well?

    Solution 2

    Define it as an undefined behavior, or a per-sequence implementation detail, that is, the exact way xs.extend(ys) works depends on the type of xs and the type of ys, whether ys = xs or ys = iter(xs) (or anything else that "accesses" xs while ys is consumed).

    Solution 3

    Something else?


    If we were to introduce new sequences, we would also need to consider this question. So I think it's better that in sequence.extend, we say that it's a per-sequence implementation detail, and that in list, bytearray and array.array, we document the CPython implementation detail. WDYT?

  10. tim-one commented on Mar 8, 2026

    @tim-one
    MemberAuthor

    For a list L,

    L.extend(L) # doubles the list length
    L,.extend(iter(L)) # unbounded time and space use
    

    under CPython. Same for array.array. Likewise for P:yPy. bytearray is the oddball under both, PyPy is generally very good at reproducing CPython's quirks. Don't know about other implementations (I never use others). Anyone else?

    Real code does depend on both aspects of list behavior. Which is hard enough to explain on its own - but is a "fact on the ground", and has been so since the iterator protocol was added.

    All types act the same way as iter does sunder an explicit "for" loop (unbounded expense):

    for x in container:
        container.append(x)

    So, to me, that's the "most obvious" behavior, and I consider bytearray.extend's quirk to be a design error. But one that probably can't be changed now 😦 .

    Since real code does rely on list.extend(iter)'s behavior (don't know about array.array),, I don't think there's any chance of changing it. It should be documented instead. Also for bytearray.extend(iter) different behavior.

    For the future, any new sequence type added "should" follow list. Although the implementation has several cases and I'm not entirely sure what they all do 😉 .

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    docsDocumentation in the Doc dirinterpreter-core(Objects, Python, Grammar, and Parser dirs)

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions