Repository navigation
Quadratic time internal base conversions #90716
Description
Activity
Our internal base conversion algorithms between power-of-2 and non-power-of-2 bases are quadratic time, and that's been annoying forever ;-) This applies to int<->str and int<->decimal.Decimal conversions. Sometimes the conversion is implicit, like when comparing an int to a Decimal.
For example:
>> a = 1 << 1000000000 # yup! a billion and one bits
>> s = str(a)I gave up after waiting for over 8 hours, and the computation apparently can't be interrupted.
In contrast, using the function in the attached todecstr.py gets the result in under a minute:
>>> a = 1 << 1000000000 >>> s = todecstr(a) >>> len(s) 301029996
That builds an equal decimal.Decimal in a "clever" recursive way, and then just applies str to _that_.
That's actually a best case for the function, which gets major benefit from the mountains of trailing 0 bits. A worst case is all 1-bits, but that still finishes in under 5 minutes:
>>> a = 1 << 1000000000 >>> s2 = todecstr(a - 1) >>> len(s2) 301029996 >>> s[-10:], s2[-10:] ('1787109376', '1787109375')
A similar kind of function could certainly be written to convert from Decimal to int much faster, but it would probably be less effective. These things avoid explicit division entirely, but fat multiplies are key, and Decimal implements a fancier * algorithm than Karatsuba.
Not for the faint of heart ;-)
Reacted by tahle, dgw, Oleg Iarygin, Takahiro Ueda, Kavi Gupta and CubedProgrammer- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usagePerformance or resource usage
on Jan 28, 2022 Is this similar to https://bugs.python.org/issue3451 ?
Dennis, partly, although that was more aimed at speeding division, while the approach here doesn't use division at all.
However, thinking about it, the implementation I attached doesn't actually for many cases (it doesn't build as much of the power tree in advance as may be needed). Which I missed because all the test cases I tried had mountains of trailing 0 or 1 bits, not mixtures.
So I'm closing this anyway, at least until I can dream up an approach that always works. Thanks!
Changed the code so that inner() only references one of the O(log log n) powers of 2 we actually precomputed (it could get lost before if
lowas non-zero but withinnhad at least one leading zero bit - now we pass the conceptual width instead of computing it on the fly).The test case here is a = (1 << 100000000) - 1, a solid string of 100 million 1 bits. The goal is to convert to a decimal string.
Methods:
native: str(a)
numeral: the Python numeral() function from bpo-3451's div.py after adapting to use the Python divmod_fast() from the same report's fast_div.py.
todecstr: from the Python file attached to this report.
gmp: str() applied to gmpy2.mpz(a).
Timings:
native: don't know; gave up after waiting over 2 1/2 hours.
numeral: about 5 1/2 minutes.
todecstr: under 30 seconds. (*)
gmp: under 6 seconds.So there's room for improvement ;-)
But here's the thing: I've lost count of how many times someone has whipped up a pure-Python implementation of a bigint algorithm that leaves CPython in the dust. And they're generally pretty easy in Python.
But then they die there, because converting to C is soul-crushing, losing the beauty and elegance and compactness to mountains of low-level details of memory-management, refcounting, and checking for errors after every tiny operation.
So a new question in this endless dilemma: _why_ do we need to convert to C? Why not leave the extreme cases to far-easier to write and maintain Python code? When we're cutting runtime from hours down to minutes, we're focusing on entirely the wrong end to not settle for 2 minutes because it may be theoretically possible to cut that to 1 minute by resorting to C.
(*) I hope this algorithm tickles you by defying expectations ;-) It essentially stands
numeral()on its head by splitting the input by a power of 2 instead of by a power of 10. As a result no divisions are used. But instead of shifting decimal digits into place, it has to multiply the high-end pieces by powers of 2. That seems insane on the face of it, but hard to argue with the clock ;-) The "tricks" here are that the O(log log n) powers of 2 needed can be computed efficiently in advance of any splitting, and that all the heavy arithmetic is done in thedecimalmodule, which implements fancier-than-Karatsuba multiplication and whose values can be converted to decimal strings very quickly.Reacted by Gregory P. Smith, Jan and Vassilis PapadimasAddendum: the "native" time (for built in str(a)) in the msg above turned out to be over 3 hours and 50 minutes.
Somebody pointed me to V8's implementation of str(bigint) today:
https://github.com/v8/v8/blob/main/src/bigint/tostring.cc
They say that they can compute str(factorial(1_000_000)) (which is 5.5 million decimal digits) in 1.5s:
https://twitter.com/JakobKummerow/status/1487872478076620800
As far as I understand the code (I suck at C++) they recursively split the bigint into halves using % 10^n at each recursion step, but pre-compute and cache the divisors' inverses.
The factorial of a million is much smaller than the case I was looking at. Here are rough timings on my box, for computing the decimal string from the bigint (and, yes, they all return the same string):
native: 475 seconds (about 8 minutes)
numeral: 22.3 seconds
todecstr: 4.10 seconds
gmp: 0.74 seconds"They recursively split the bigint into halves using % 10^n at each recursion step". That's the standard trick for "output" conversions. Beyond that, there are different ways to try to use "fat" multiplications instead of division. The recursive splitting all on its own can help, but dramatic speedups need dramatically faster multiplication.
todecstr treats it as an "input" conversion instead, using the
decimalmodule to work mostly in base 10. That, by itself, reduces the role of division (to none at all in the Python code), anddecimalhas a more advanced multiplication algorithm than CPython's bigints have.todecstr treats it as an "input" conversion instead, ...
Worth pointing this out since it doesn't seem widely known: "input" base conversions are _generally_ faster than "output" ones. Working in the destination base (or a power of it) is generally simpler.
In the math.factorial(1000000) example, it takes CPython more than 3x longer for str() to convert it to base 10 than for int() to reconstruct the bigint from that string. Not an O() thing (they're both quadratic time in CPython today).
Reacted by Gregory P. Smith80 remaining items
It seems the num-bigint crate attaches to Python 3 quite well with PyO3. The following image shows the performance of the new multiplication implementation on Ryzen 7 2700X, averaged over 10 runs. For small numbers, the num-bigint crate uses naive, Karatsuba, and Toom-3 before switching to number-theoretic transform. The cross-over point over the CPython native multiplication (version 3.11.3) is around 3,000 bits.
The binding I used is very simple. Further optimizations may be possible.
use pyo3::prelude::*; use num_bigint::BigInt; #[pyfunction] fn mul(x: BigInt, y: BigInt) -> BigInt { &x * &y } #[pymodule] fn num_bigint_pyo3(_py: Python, m: &PyModule) -> PyResult<()> { m.add_function(wrap_pyfunction!(mul, m)?)?; Ok(()) }
import num_bigint_pyo3 a, b = 10**100000, 9**100000 x = num_bigint_pyo3.mul(a, b) assert x == a*b
Here are the str -> int conversion timings with the new multiplication routine.
- 40,000 digits:
int : 0.006869841 ± 0.000509983 sec str_to_int_using_decimal : 0.026808119 ± 0.000423398 sec str_to_int : 0.003381038 ± 0.000521641 sec str_to_int_new_mul : 0.001501775 ± 0.000504214 sec- 400,000 digits:
int : 0.678053117 ± 0.005104713 sec str_to_int_using_decimal : 0.553185940 ± 0.003627799 sec str_to_int : 0.127861691 ± 0.000618099 sec str_to_int_new_mul : 0.023358464 ± 0.000615376 sec- 4,000,000 digits:
int : 68.152991366 ± 0.353029732 sec str_to_int_using_decimal : 9.126699829 ± 0.067300091 sec str_to_int : 4.851845384 ± 0.030668179 sec str_to_int_new_mul : 0.326830316 ± 0.002546696 secCode
import sys sys.set_int_max_str_digits(0) import num_bigint_pyo3 from decimal import * from time import time from random import choices import numpy as np setcontext(Context(prec=MAX_PREC, Emax=MAX_EMAX, Emin=MIN_EMIN)) pow5 = [5] while len(pow5) <= 23: pow5.append(num_bigint_pyo3.mul(pow5[-1], pow5[-1])) def str_to_int_new_mul(s): def _str_to_int(l, r): if r - l <= 3000: return int(s[l:r]) lg_split = (r - l - 1).bit_length() - 1 split = 1 << lg_split return (num_bigint_pyo3.mul(_str_to_int(l, r - split), pow5[lg_split]) << split) + _str_to_int(r - split, r) return _str_to_int(0, len(s)) def str_to_int(s): def _str_to_int(l, r): if r - l <= 3000: return int(s[l:r]) lg_split = (r - l - 1).bit_length() - 1 split = 1 << lg_split return ((_str_to_int(l, r - split) * pow5[lg_split]) << split) + _str_to_int(r - split, r) return _str_to_int(0, len(s)) def str_to_int_using_decimal(s, bits=1024): d = Decimal(s) div = Decimal(2) ** bits divs = [] while div <= d: divs.append(div) div *= div digits = [d] for div in reversed(divs): digits = [x for digit in digits for x in divmod(digit, div)] if not digits[0]: del digits[0] b = b''.join(int(digit).to_bytes(bits//8, 'big') for digit in digits) return int.from_bytes(b, 'big') # Benchmark against int() on random 1-million digits string funcs = [int, str_to_int_using_decimal, str_to_int, str_to_int_new_mul] s = ''.join(choices('123456789', k=1_000_000)) expect = None for f in funcs: t_list = [] for _ in range(10): t = time() result = f(s) t_elapsed = time() - t # f.__name__) if expect is None: expect = result assert result == expect t_list.append(t_elapsed) mean, std = np.mean(t_list), np.std(t_list) print("{0}: {1:.9f} ± {2:.9f} sec".format(f.__name__.ljust(30), mean, std))
Reacted by Bubbler and GalaxySnailHere are the str -> int conversion timings with the new multiplication routine.
It looks very nice to me. The question is what you would propose to do going on from here. If the suggestion is that CPython might depend on this (even in an optional way) then there are many details that would need to be worked out before that could be considered.
It looks very nice to me. The question is what you would propose to do going on from here. If the suggestion is that CPython might depend on this (even in an optional way) then there are many details that would need to be worked out before that could be considered.
Yes, I'd love to see CPython's multiplication improved asymptotically. It would also be great to improve other arithmetic operations too, but as a first step, I think it would be appropriate to limit the scope to multiplication.
Currently the binding through PyO3 receives the raw bits of the integer through _PyLong_ToByteArray and sends it back to Python via _PyLong_FromByteArray. I think that mechanism can still work in the C code, which might help lessen the maintenance burden since the existing APIs don't need to be modified substantially.
On the other hand, more deliberation is needed to port the primitive operations from Rust to C. Rust provides portable primitives such as add-with-carry and 128bit integer. Although it is possible to access the higher 64bit of a 64 x 64 multiplication in most C compilers, this would not be portable. Switching to plain portable C (without compiler intrinsics) is possible but it comes with performance degradations.
For memory allocation, one call of malloc (and a corresponding call of free) would suffice. The temporary buffer has a size linear in the input integer and doesn't need to be a PyObject instance.
Please let me know of other details that need to be worked out. Thanks!
- added3.13only security fixesonly security fixesand removed3.12only security fixesonly security fixes
on Jan 5, 2024 Is there anything left to do in this issue or it can be closed?
Reacted by Tim PetersIt's hard to tell because of the length and complexity of the discussion, but
decimal.Decimal <-> intconversions remain quadratic-time. It's onlystr <-> intconversions that were effectively addressed.- addedextension-modulesC modules in the Modules dirC modules in the Modules dirand removed3.13only security fixesonly security fixes
on Jul 12, 2025 decimal.Decimal <-> int conversion handled by libmpdec's mpd_qexport_/mpd_qimport_ API. We are going to remove libmpdec bundled copy in 3.16.
Is CPython an appopriate place to fix/track libmpdec's issues? I think that fix should rather be submitted to upstream.
If Python wants faster Decimal<->int conversion than libmpdec offers, then it's up to Python to arrange for that in its libmpdec wrapper. The "internal" in this issue's "Quadratic time internal base conversions" is there for a reason.
Whether Stefan Krah wants to build faster conversion inside libmpdec is a different (although possibly related) issue. Making Python's story consistent across the Python types Python offers isn't really libmpdec's problem to solve. Sure, it would make life easier for us if libmpf\dec did do it much faster itself. But unless and until it does, this remains a Python issue.
closing as not planned, if someone is ever actually going to work on this they can reopen it with a plan.
- marked Decimal to int conversion with engineering notation costly despite failing #139957 as a duplicate of this issue
on Oct 11, 2025

Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: