Everything Is an Interpretation
Four bytes in memory have no type. Read as a 32-bit integer they are one number, read as a floating-point value another, read as text four letters, and handed to the processor as code they are instructions. Computing Foundations established that meaning comes from an agreed code. This topic goes one layer down, to what happens when two parts of a system agree on different codes.
The answer is the most expensive kind of bug there is: nothing fails. The bytes are intact, every reader does its job correctly, and the value that comes out is wrong. No exception is raised, no log line is written, and the wrong value travels on into reports, databases and other systems until a person notices it.
Four Bytes, Four Readings
Take the four bytes 43, 41, 46 and 45, written in hexadecimal. Read as ASCII text they are the word CAFE, one letter per byte. Read as an unsigned 32-bit integer with the first byte most significant, they are 1,128,351,301. Read with the first byte least significant, which is how most processors store integers and the subject of Byte Order and Serialization later in this chapter, they are 1,162,232,131.
Read as a 32-bit floating-point number, the same bits split into a sign, an exponent and a fraction and come out as about 193.27. And handed to a processor as machine code they are instructions. A 32-bit x86 processor decodes them as four one-byte instructions, each adding one to a different register. A 64-bit x86 processor treats them as prefix bytes that modify whatever instruction follows, and a 64-bit ARM processor with the SVE2 vector extension reads all four together as a single vector arithmetic instruction. Every one of these readings was computed on the same four bytes, and every one is correct for its reader.
The Type Lives in the Reader
Nothing in memory marks where an integer ends and a string begins. There is no tag on a byte saying "I am the second half of a float". The program that reads the bytes decides what they mean, from a declaration in its source code, a file format's specification, a header field, or a convention the two sides agreed on long ago and may have forgotten.
So a mismatch produces a plausible wrong value rather than a crash. A reader expecting a 32-bit integer gets 32 bits, and every pattern of 32 bits is a valid integer. A reader expecting a float gets a float. There is nothing for the reader to reject, because the bytes have the shape it asked for. The error lives in the agreement, and the agreement is not in the data.
How Files Announce What They Are
A file is bytes plus a name. The extension at the end of the name is a hint that anyone can change by renaming the file. Many formats therefore announce themselves in their first bytes, a signature often called a magic number. Every PNG image starts with the byte 89 in hexadecimal followed by the letters PNG. Every PDF starts with the characters "%PDF". A ZIP archive starts with the letters PK, and so do the modern Word and Excel formats and Java archives, which are ZIP files inside.
A reader that trusts the extension and ignores the content is how a renamed file gets handed to the wrong parser. An upload form that accepts "photo.png" and passes it to an image library will happily pass along a file that is really a ZIP archive or a script. Checking the signature is better, and it is still a guess: a signature says what the first few bytes claim to be, not that the rest of the file is well formed.
Carrying the Type With the Data
One way to prevent misreadings is to store the type next to every value. Dynamically typed languages do this at run time. Every Python object starts with a reference count and a pointer to its type object, and the integer 1,000 carries a field recording its size and sign before its one 4-byte digit. On CPython 3.15 that integer occupies 28 bytes where a bare 64-bit integer needs 8, which is the overhead Chapter 1 measured. The payoff is that Python can never read an integer as a float by accident: every operation asks the object what it is first.
Statically typed languages make the opposite trade. The compiler settles every type once, before the program runs, and stores bare bytes with no tag at all, which is smaller and faster. Inside the program the compiler keeps the readings consistent, which is the subject of Chapter 13. Outside the program, in a file or on a network, neither approach helps: the type the program had in mind stays behind, and only the bytes travel.
Code Is Data Too
Instructions and data sit in the same memory, in the same bytes, and the processor cannot tell them apart. It runs whatever the program counter points at, a register Chapter 3 describes. So bytes an attacker supplies as data, the contents of a form field or a network message, can be run as code if the program is tricked into jumping to them.
The hardware's answer is to mark regions of memory as non-executable, so that the processor refuses to fetch instructions from pages that are meant to hold data. That closes the simplest version of the attack, and attackers moved to reusing code that is already executable. The bugs that let an attacker redirect the jump in the first place are memory-safety bugs, which Chapter 4 covers.
The Cost of a Wrong Reading
Three failures show the pattern. The word "café" encoded as UTF-8 and read back as Latin-1, an older encoding with one byte for every character, displays as "café", because the two bytes of the accented letter are read as two separate characters. This garbling is called mojibake. A spreadsheet that reads the gene name SEPT2 decides it is the second of September and stores a date. A 2016 survey of genetics papers found that about one in five with gene lists in Excel files had names corrupted this way, and in 2020 the committee that names human genes renamed 27 of them, SEPT1 to SEPTIN1 and MARCH1 to MARCHF1 among them, so that spreadsheets would stop converting them. And a binary field written little-endian and read big-endian turns the number 1 into 16,777,216.
None of these raise an error. All of them corrupt data quietly, and the corruption is usually found far from where it happened, by a person looking at a report. The repair is always the same: make the interpretation explicit at every boundary. Name the encoding when text becomes bytes, name the byte order when a number is written, validate the format when a file is read, and keep the type checks on the side that receives the data, because the sender's intentions do not travel with it.
- "The file knows what it contains." A file is bytes plus a name. The extension is a convention anyone can change, and a reader that guesses from it, or from the first few bytes, can guess wrong without complaint.
- "If the data were misread, the program would crash." A misread integer is still an integer and misread text is still text. Interpretation errors produce wrong values, not exceptions, and so they reach production.
- "Python strings and bytes are interchangeable." A string is a sequence of code points and a bytes object is a sequence of raw bytes. Converting between them always uses an encoding. Code that does not name one relies on a default, and Python changed its own default in version 3.15: files opened without an encoding are now read as UTF-8 everywhere, where older versions used whatever the machine's locale said.
- "Types are a language feature, not a data feature." Every byte that crosses a boundary, a file, a socket or a database column, loses the type the program had in mind. The receiver must be told again, by a schema, a format or a convention.
- Name the encoding, byte order and type explicitly wherever bytes cross a boundary. A default is an agreement with whichever machine happens to run the code.
- Check a file's content signature, not its extension, on any upload path. The extension is user input, and so is the name.
- Keep text as text and bytes as bytes inside the program, and convert only at the edges. Mixed handling in the middle is where mojibake is born.
- Treat any data that will be executed or evaluated as code, wherever it came from. The processor cannot tell the difference, so the program has to.
Knowledge Check
A service writes a 32-bit count to a file, and a second service reads the field with the wrong type. What is the most likely result?
- The reader raises a type error as soon as it opens the field
- The reader returns a plausible value that is quietly wrong
- The operating system refuses to hand the bytes to the reader
- The reader detects the mismatch and converts the value itself
Why does the integer 1,000 take 28 bytes in CPython when a bare 64-bit integer needs 8?
- CPython stores the digits of every integer as decimal text characters
- Every integer reserves spare room so it can grow later without copying
- Each object carries a reference count, a type pointer and a size field
- Python keeps a second copy of each integer for garbage collection
An upload endpoint accepts only files named with a .png extension and passes them to an image library. What is the weakness?
- The name is chosen by the uploader, so any content can arrive as .png
- The filter will reject real PNG files whose extension is in capitals
- PNG files lose their content signature when renamed by the uploader
- PNG images can contain machine code that the image library will run
The bytes 00 00 00 01 are written by one program and read back by another as a 32-bit integer in the opposite byte order. What does the reader get?
- An error, because the byte order does not match the file
- The number 1, because a single set bit reads the same way
- The number 16,777,216, with the 1 in the top byte instead
- A negative number, because the top bit is now set to one
You got correct