Topic 05

Byte Order and Serialization

Representation

A program's data in memory is a web of objects and pointers laid out for one processor. A file or a network message is a flat line of bytes that another program, on another machine, possibly in another language, must read back. Serialization is the translation between the two.

Every choice in that translation has a price in size, speed or safety: which byte of a number comes first, where padding goes, whether numbers are written as digits or as bits, whether the reader needs a schema, and whether the format can describe arbitrary objects. This topic prices each one.

Byte Order

A 32-bit number takes four bytes, and machines disagree on which comes first. Big-endian stores the most significant byte at the lowest address, the way people write digits. Little-endian stores the least significant byte first. The names come from Gulliver's Travels, by way of a 1980 note by the network engineer Danny Cohen, and the argument they describe is about as consequential as which end of an egg to crack, until two machines have to agree.

x86 processors are little-endian, and ARM processors, which can run either way, run little-endian in essentially every phone, laptop and server in use. Network protocols fix big-endian by convention, so it is also called network byte order. A binary format must name its order, or the machine on the other side reads it backwards: the integer 1 written little-endian and read big-endian is 16,777,216.

Two byte orders for one number
The 32-bit number 0x0A0B0C0D in four memory cellsaddress 100address 101address 102address 103Big-endianmost significant first0A0B0C0DLittle-endianleast significant first0D0C0B0ARead little-endian bytes as big-endian and 1 becomes 16,777,216

Alignment and Padding

Processors read memory fastest in naturally aligned chunks: an 8-byte integer at an address that is a multiple of 8. Compilers therefore insert padding between fields. A record of a 1-byte flag, an 8-byte integer and another 1-byte flag occupies 24 bytes on a typical 64-bit system, measured here with Python's ctypes module: the integer must start at offset 8, so seven bytes of padding follow the first flag, and seven more round the record up so the next record in an array is aligned too.

Reorder the same fields largest first, the integer then the two flags, and the record occupies 16 bytes: only six bytes of padding at the end. A million such records differ by 8 megabytes for nothing but field order. Formats that write a struct's raw memory to disk or to the network also write its padding, whatever garbage it holds, and depend on the reader using the same compiler rules.

Field order changes the size of a record
flag, int64, flag: 24 bytesflagpadint64flagpadint64, flag, flag: 16 bytesint64flagflagpad081624

Why Pointers Do Not Travel

An address means something only inside one process: Chapter 10 shows that every process sees its own private address space, so address 4096 in one program is unrelated to address 4096 in another, or in the same program tomorrow. Serializing an object graph therefore means replacing every reference with something the reader can rebuild from: an id, a nested copy, or an offset within the message.

An object graph has to be taken apart to travel
Object graphpointers
→
Replace refsids, copies
→
Flat bytesfile or socket
→
Rebuildnew addresses

Shared references and cycles must be handled explicitly. If two orders point at the same customer object, a format that nests copies writes the customer twice, and the reader gets two customers that no longer change together. If the customer also points back at its orders, a naive nesting serializer never finishes. A format that cannot express shared references either duplicates them or refuses the input.

Text Formats

JSON, CSV and XML are readable by a person, self-describing, and supported by every language. They pay in size and in parsing time. A million 32-bit integers are exactly 4 megabytes in binary. The integers 0 to 999,999 written as a compact JSON array are 6,888,891 bytes, about 6.9 megabytes, and Python's json module with its default spacing writes 7,888,890. On the laptop this book was written on, parsing that JSON took about 70 milliseconds on Python 3.15, while loading the same numbers from the 4-megabyte binary copy took under one millisecond: every JSON number is re-parsed from characters.

JSON also has exactly one number type, and many consumers read it as a 64-bit float, which corrupts integers above 2 to the 53 as the previous topic showed. It has no bytes type, so binary data travels as base64 text, a third larger. It has no dates, so every system invents a string convention. And producers disagree at the edges: Python's json module writes a float NaN as the bare word NaN, which the JSON standard does not allow and strict parsers reject.

Binary Formats and Schemas

Protocol Buffers, MessagePack, Avro and their relatives store numbers in binary and field names as small numeric tags, or not at all. Messages are smaller and faster to read, and the price is that the bytes make sense only with the schema, or to a reader that already knows it. Open a Protocol Buffers message in a text editor and you see tags and raw bytes, not names.

The schema then becomes a contract between every writer and every reader, and it must evolve without breaking either. That is why such formats number their fields and never reuse a number: an old reader skips a field number it does not know, and a new reader supplies a default for one that is missing, so old and new versions of a service can talk during a rolling deploy.

What a Format Costs

The format sets bandwidth, CPU and safety. An internal service that ships JSON pays parsing cost on every request, and at thousands of requests a second that cost shows up in the CPU bill. A binary format pays in tooling instead: no message you can read by eye, no diff that means anything, and a schema registry to run.

The sharpest cost is safety. A format that can rebuild arbitrary objects, such as Python's pickle or Java's native serialization, records which classes to construct and which functions to call as part of the data. Reading it runs code chosen by whoever wrote the bytes, and Python's documentation warns that the pickle module is not secure and that only trusted data should ever be unpickled. Deserializing untrusted input with such a format is remote code execution, reached through what looks like reading a file.

Misconceptions
  • "Byte order is a historical curiosity." Every binary file format and network protocol fixes one. Code that writes an integer's raw memory to disk produces a file that a machine of the other order reads as a different number.
  • "A record's size is the sum of its fields." Padding for alignment can add half again or more: three fields totalling 10 bytes occupy 24. Field order changes the size of every record, and the difference multiplies across a million-row array or a network message.
  • "JSON is a universal lossless format." It has no integers distinct from floats, no bytes, no dates, and no way to represent a shared reference. Large ids, binary blobs and timestamps all need a convention layered on top.
  • "Pickle is a faster JSON for Python objects." Unpickling constructs whatever objects and calls whatever functions the data names, so a crafted pickle executes arbitrary code. It is safe only for bytes the program itself wrote and nobody else could alter.
  • "Binary formats are always faster." For small messages the parsing difference is microseconds and the debugging cost is hours. Binary pays off on high-volume paths and large numeric payloads, not everywhere.
Why It Matters
  • Name the byte order in every binary format and convert explicitly on read and write. Never write a native in-memory integer to a file or socket.
  • Use a self-describing text format at human-facing and low-volume boundaries, and a schema-based binary format on high-volume internal paths. The cost that matters differs at each.
  • Never deserialize untrusted input with a format that can instantiate arbitrary objects. Use a data-only format and validate what comes out.
  • Evolve schemas by adding optional fields, and never reuse a field number or name. Old readers and new writers coexist for longer than any deploy.
RelatedEverything is an interpretation byte order is the classic misreadingCompression a compressed text format approaches binary size, at a CPU costVirtual memory why an address means nothing outside its process (Chapter 10)

Knowledge Check

A little-endian machine writes the 32-bit integer 1 to a file as raw memory. A big-endian reader loads it. What value does the reader see?

  • 1, because the bytes themselves did not change
  • 16,777,216, because the 1 lands in the top byte
  • 128, because the bits inside the byte are reversed
  • An error, because the file's byte order is marked

A record type has a 1-byte flag, an 8-byte integer and a 1-byte flag, in that order, and occupies 24 bytes. Why does reordering it shrink it?

  • Putting the integer first lets the compiler compress the two flags
  • Reordering turns alignment off, so no padding is inserted at all
  • With the integer first, the flags share one padded 8-byte slot
  • Placing the integer first lets it shrink to four bytes when small

A service exchanges a million 32-bit integers per message as a JSON array. What does that cost compared with packed binary?

  • About 70% more bytes, and parsing every number from text
  • Nothing, because JSON stores each integer in four bytes
  • Precision, because every value is rounded to three digits
  • Fewer bytes, but a schema must be shipped with each one

Why must every pointer be replaced when an object graph is serialized?

  • Pointers are too large to write efficiently into a file
  • The operating system forbids writing addresses to a file
  • Pointers are encrypted and cannot be read by the receiver
  • An address has meaning only inside the process that made it

An internal tool loads user-uploaded settings files with pickle because it is convenient. What is the risk?

  • Uploaded files may be slightly larger than their JSON versions
  • A crafted file can make the tool run arbitrary code on load
  • A file written on another machine may load in the wrong byte order
  • Files become unreadable after the tool's interpreter restarts

You got correct