Integers and Overflow
A processor adds integers of a fixed width, 8, 16, 32 or 64 bits, and when a result does not fit, it wraps around like an odometer rolling over. That single fact has forced airliners to be rebooted on a schedule, sent the world's most-watched video past its counter's ceiling, and will stop unpatched 32-bit clocks in January 2038.
Python hides the fact by letting integers grow without limit, and pays for the hiding on every addition. The limit does not disappear. It waits at the edge of the program, where a Python integer is written to a database column, a binary file or another language that never made the same promise.
Fixed Widths and Their Ranges
With n bits there are 2 to the power n patterns, and a fixed-width integer maps each pattern to one number. An unsigned byte holds 0 to 255. An unsigned 32-bit integer holds 0 to 4,294,967,295, which is why IPv4, with 32-bit addresses, has about 4.3 billion of them. A signed 32-bit integer spends half its patterns on negative numbers and holds −2,147,483,648 to 2,147,483,647. A signed 64-bit integer reaches about 9.2 billion billion in each direction.
The width is chosen once, when the variable, the column or the file format is designed, and every value that ever passes through must fit. Nothing grows later. A counter declared as 32 bits in 2010 is still 32 bits when its traffic is a thousand times higher.
Two's Complement by Picture
Negative numbers are stored by counting backwards from the top of the range. Picture the 256 patterns of one byte arranged on a wheel. Counting clockwise from zero, the first half, 0 to 127, are the non-negative numbers, and all of them have a top bit of 0. Keep going and the second half is read as negative: 11111111, the pattern immediately before zero, is −1, and 10000000, directly opposite zero, is −128. This scheme is called two's complement.
Its great advantage is that the adder does not need to know which reading it has. Adding 1 moves one step clockwise on the wheel whether the byte is read as signed or unsigned: 11111111 plus 1 becomes 00000000, which is 255 plus 1 wrapping to 0 in one reading and −1 plus 1 giving 0 in the other. One circuit, which the next chapter builds from gates, serves both, and that is why every modern processor stores negative integers this way.
Overflow Is Silent Arithmetic
The wrap point sits between 127 and −128. So 127 plus 1 in a signed byte is −128, and 0 minus 1 in an unsigned 32-bit integer is 4,294,967,295. The hardware sets a flag when it happens and carries on. Most languages ignore the flag. A few trap. C goes further and declares signed overflow undefined behaviour, which lets its compiler assume overflow never happens and remove checks that were written to catch it.
The wheel also explains a stranger bug. There are 128 negative numbers and only 127 positive ones, so the most negative value has no positive partner. The absolute value of −128 in a signed byte, or of −2,147,483,648 in a signed 32-bit integer, is itself: the result wraps straight back. NumPy, measured on its current version, returns −128 for the absolute value of −128 in an 8-bit array, without a warning.
Overflow in the Wild
Many systems store time as a signed 32-bit count of seconds since the start of 1970. That count runs out at 03:14:07 UTC on 19 January 2038, and the next second wraps to a date in December 1901. A signed 32-bit counter of hundredths of a second runs out after 2 to the power 31 hundredths, about 248 days. That is the span in a 2015 airworthiness directive from the US Federal Aviation Administration, which blamed a software counter that overflows after 248 days of continuous power and required Boeing 787 operators to power-cycle the aircraft's electrical system before any unit had run that long.
In December 2014 YouTube announced it had moved its view counter to 64 bits because one video, "Gangnam Style", was approaching 2,147,483,647 views. A database table whose id column is a signed 32-bit integer hits the same number of rows, and a busy table inserting a hundred rows a second gets there in under a year.
Why Python's Integers Do Not Overflow
A Python int stores as many 30-bit digits as its value needs, so 2 to the power 1,000 is exact: a 302-digit number in a 160-byte object. Adding 1 to the largest 64-bit value produces a bigger integer. No wrap is possible, because no width was ever fixed.
Python pays for that on every value and every operation. Every integer is a heap object, 28 bytes for a small one where a machine integer takes 8. Every addition checks both operands' sizes, allocates a result object unless the answer is a small cached number, and loops over digits. Measured on CPython 3.15, adding two small integers takes about 15 nanoseconds and adding two 10,000-digit integers about a microsecond, where the processor's own addition takes well under one nanosecond. Some operations grow faster than the digits. Converting an integer to a decimal string is quadratic in the number of digits: ten times the digits took about a hundred times as long, from 2 microseconds at 400 digits to 207 at 4,000. That was slow enough to be a denial-of-service attack, so since Python 3.11 converting an integer of more than 4,300 digits to or from a string raises a ValueError by default.
The Boundary Where the Width Comes Back
Python's unbounded int meets a fixed width at every exit. A database column declared as a signed 32-bit integer rejects a larger value, or clips it to the largest value that fits, depending on the database and its settings. A NumPy array has a fixed element type, and arithmetic on 8-bit arrays wraps 127 plus 1 to −128 with no warning. A JSON consumer written in JavaScript parses every number as a 64-bit float, which holds integers exactly only up to 2 to the power 53, so the id 9,007,199,254,740,993 arrives as 9,007,199,254,740,992. A binary file format with a 32-bit length field cannot describe a file larger than 4 gigabytes.
The overflow Python protected you from happens in the next system. The cost arrives as a wrong value stored somewhere else, discovered later by someone who did not write the code, and never as an exception at the line that caused it.
- "Integer overflow throws an error." In the hardware and in most languages it wraps silently. Python raises nothing because it never overflows, and NumPy arrays and database drivers downstream may wrap or truncate instead.
- "Python has no integer limits, so my program has none." The limits reappear the moment a value leaves Python. A signed 32-bit database column stops at 2,147,483,647, and then the insert fails or the value is clipped, depending on the system.
- "The 2038 problem was fixed like Y2K." Systems that moved to a 64-bit time are fixed. Embedded devices, old file formats and database columns that store 32-bit timestamps are not, and some of them are being installed today with twelve years of working life ahead of them.
- "The absolute value is never negative." For the most negative fixed-width integer there is no positive counterpart, so the absolute value is the same negative number. Code that takes the absolute value of a hash and then its remainder to pick a slot has shipped this bug.
- "Unsigned integers are safer because they cannot go negative." They cannot, so 0 minus 1 becomes 4,294,967,295. A loop counting down past zero, or a length computed as one size minus another, turns into a four-billion-step loop or a huge memory allocation.
- Choose 64-bit integers for ids, counters and timestamps by default. The 32-bit ceiling of 2.1 billion is within reach of a busy table or a counter that runs for years.
- Check the width at every boundary where a Python int leaves Python. Database columns, binary formats, arrays and JSON consumers each have one.
- Store time as a 64-bit count or a proper timestamp type, never a 32-bit seconds field. 2038 falls inside the working life of hardware bought now.
- Validate the range of integers from untrusted input before using them as sizes or indexes. An overflowed length is how an integer bug becomes a memory bug, the subject of Chapter 4.
Knowledge Check
A signed 8-bit counter holds 127 and the program adds 1. What does the counter hold afterwards on typical hardware?
- 128, because the hardware widens the value automatically
- 127, because the addition stops at the top of the range
- −128, because the pattern wraps round to the negative half
- 0, because the overflow flag clears the register to zero
Why do processors store negative integers in two's complement rather than, say, a separate sign bit?
- One adder circuit then works for signed and unsigned values alike
- It lets a 32-bit register hold larger numbers than a sign bit would
- It prevents overflow, because negative numbers absorb the wrap
- It gives every number, including zero, both a positive and a negative form
A new building-control device stores its clock as a signed 32-bit count of seconds since 1970 and is expected to run for 15 years. What happens?
- Nothing, because the limit applies only to devices built before 2000
- Its clock wraps to 1901 in January 2038, well within its working life
- Its clock wraps every 248 days, so it must be restarted each year
- The operating system widens the count to 64 bits when it nears the limit
What does Python's arbitrary-precision int cost compared with a machine integer?
- Nothing measurable, because small values use the hardware add
- Precision, because large values are rounded to the nearest float
- A lower ceiling, because values wrap at 2 to the power 62
- Memory and time: a heap object per value and checks per step
A Python service sends record id 9,007,199,254,740,993 in JSON to a browser that parses it with JavaScript. What does the browser see?
- A parse error, because the number is too large for JSON
- A different id, 9,007,199,254,740,992, with no warning
- A negative id, because the value wraps past 2 to the power 63
- The same id as text, because JSON keeps large numbers as strings
You got correct