The Unreliable Channel
Each of Millbrook's 30 branches sends Lantern a message whenever a copy is checked out or returned, and the "copies on shelf" line on a search result is computed from those messages. The city network between the branches and the search server does what every network does. It loses some messages, delivers some twice, delivers some in the wrong order and, rarely, delivers some altered.
None of that is a malfunction to be fixed once and forgotten. It is the normal behaviour of every channel between two machines, from a home connection to the fabric inside a data centre. Each of the four failures has its own source, and one argument decides which layer of a system has to deal with them. Its answer is why the rest of this chapter exists.
Loss
A message is lost when something between sender and receiver has nowhere to put it or no way to carry it. A router whose outgoing link is busy holds arriving packets in a queue, and when the queue is full it drops the next one. A radio link fades for a moment. A cable is unplugged for a second during maintenance. A process crashes with a message still sitting in its memory, sent by nobody's definition and received by nobody.
The first cause is by far the most common, and it is not a fault. A router drops packets because more traffic arrives than its outgoing link can carry, and on a busy network that drop is the only signal a sender gets that it should slow down. Loss is how congestion announces itself, which is the subject of the flow control and congestion topic later in this chapter. A network that never lost anything would have to either stay idle or let its queues grow without limit.
Duplication and Reordering
Duplicates are mostly made by the cure for loss. A sender that hears nothing back sends the message again, and when the first copy had arrived and only its reply was lost, the receiver now has two. Reordering comes from the network's freedom to route: two packets can take different paths, or wait in different queues on the same path, and arrive in the opposite order from the one they were sent in.
Both corrupt Lantern's shelf count in ways nobody would notice from one search. A duplicated checkout subtracts twice, so the result shows one copy fewer than the shelf holds. A return that overtakes its own checkout is applied first, so for a moment the branch appears to own one more copy than it has. The figure below puts all three failures on one timeline between branch 7 and Lantern.
Message 1 arrives and is applied. Message 2 is lost halfway, and Lantern never learns that it existed. Message 3 arrives, and so does a resent copy of it, which a naive receiver applies a second time. Messages 4 and 5 leave in order and arrive in the opposite order. Every one of these happens on the city network in a normal week.
Corruption
Now and then a message arrives with different bits from the ones that were sent. Noise on a link flips a bit. A faulty network card garbles a frame. A bad memory chip in a router changes a byte while the packet sits in a buffer. Most damage on the wire is caught by the check every link frame carries, and the frame is thrown away, so corruption on a link usually turns into loss one layer up.
The dangerous remainder is the damage no check was positioned to see, and it arrives looking like valid data. A byte changed inside a router's memory is changed after one link's check and before the next link computes its own, so every check along the path passes. The next topic, on detecting corruption, is about the checks and their limits.
Layers and What Each Can Fix
Networks are built in layers, and each lower layer offers a service to the one above it. The link layer moves frames between two neighbours on one wire or one radio channel. The network layer moves packets between any two machines, on a best-effort basis: it tries, and it promises nothing. A transport such as TCP, running only on the two endpoints, turns that best effort into an ordered stream of bytes with no gaps and no duplicates. How TCP does it in practice belongs to Networking Deep Dive; this chapter teaches the idea underneath it.
Every layer removes some failures, and none removes all of them. The link layer's check sees only its own wire. The network layer adds no guarantees at all. TCP covers loss, duplicates and order within one connection, and nothing that happens to the data after it hands the bytes to the application.
The End-to-End Argument
In 1984, Jerome Saltzer, David Reed and David Clark published the argument that settles where the guarantee has to live, in a paper called "End-to-End Arguments in System Design". A function such as reliable delivery can be completely implemented only by the endpoints, because failures happen outside the layer that handled them: in the sender's disk, in a buggy copy routine, in the receiver's memory, in a crash after the transport has done its job. Lower layers may still add checks, but only as a performance improvement. Only the application can confirm that what it meant to happen, happened.
The paper's central example is a careful file transfer from one computer to another. Even over a perfect network, the file could be read wrongly from the first disk, damaged by a copying bug or lost in a crash halfway through. The only complete defence is to compute a checksum of the file at both ends, compare them and retry on a mismatch. The paper also reports a real case at MIT: a gateway that swapped a pair of bytes about once in every million, while copying data between buffers where no per-hop check could see it. Programmers who trusted the network lost parts of source files and had to restore them from old printed listings.
The Cost in a System You Run
A feed built on "the transport is reliable" works almost perfectly. When a branch's sending process restarts, it loses whatever sat in its buffer, and when Lantern's worker restarts, it loses whatever it had received and not yet applied. Everything else is counted correctly. The error is small, silent and permanent: a shelf count that is off by one for a single record, noticed only when a patron walks to the shelf.
The fix is end to end, and the rest of this chapter builds it. Every message carries its own identity: the branch number and a sequence number the branch assigns. Lantern remembers, per branch, what it has already applied, so it can see a gap, ignore a repeat and put things back in order. The branch keeps a message until Lantern confirms it has applied the change, not merely received the bytes.
- "TCP guarantees delivery." TCP guarantees that the bytes it delivers arrive in order, without gaps or duplicates, or that the connection reports an error. It cannot promise that the receiving application read, processed or stored them, and when a process or connection dies with data in flight, neither side knows exactly how much got through.
- "A send that returned means the other side has it." A send returns when the bytes are in the local kernel's buffer, as the system call topic in Chapter 10 describes. The network, the remote kernel and the remote application are all still ahead of them.
- "Data-centre networks do not drop packets." Every switch has finite buffers, and synchronized bursts, such as many servers answering one request at the same moment, overflow them. Loss is rarer, not absent, and one lost packet can turn a round trip of a fraction of a millisecond into a stall of hundreds of milliseconds while a retransmission timer runs out.
- "If every layer is reliable, the whole system is reliable." A reliable link and a reliable transport still lose the message the receiving process held in memory when it crashed. Reliability composes only when the endpoints check.
- "Packet loss means something is broken." Loss is how a busy network tells senders to slow down. A network with no loss under heavy load is either idle or hiding its congestion in queues that keep growing.
- Design every message path for loss, duplication and reordering at once. All three happen on every real network, and fixing one exposes the others.
- Give every message an identity assigned by its sender. A branch number and a per-branch sequence number let the receiver see gaps, duplicates and order.
- Confirm at the endpoint that the effect happened, not that the bytes arrived. Only the application's own acknowledgement covers the whole path.
- Treat transport reliability as a performance layer, not as the correctness guarantee. It makes end-to-end retries rare; it does not make them unnecessary.
Knowledge Check
Lantern shows one copy fewer on the shelf than a branch really holds. Which failure of the feed most likely caused it?
- A checkout message that was lost somewhere on the way
- A checkout message that was delivered twice
- A return message that was delivered to Lantern twice
- A return that arrived at Lantern before its checkout
What does TCP actually promise the application that receives its data?
- That every byte sent was read and stored by the receiving application
- That no data is lost even if either process crashes mid-transfer
- Delivered bytes arrive in order with no gaps and no copies, or an error is reported
- That each send call returns only after the remote side holds the bytes
A service copies uploaded files from a web server to object storage over TLS. Where does the end-to-end argument say the integrity check belongs?
- In the TLS layer, which already rejects any record altered in transit
- In the network cards, which checksum every frame they send or receive
- Nowhere extra, since TCP already detects and retransmits damaged segments
- In the application: hash the file at the source and after storing it
A network under heavy load shows almost no packet loss. What is the most likely explanation?
- Its switches never drop a packet, so the heavy load has no cost at all
- Its queues are absorbing the excess, so latency is climbing instead
- Its senders are retransmitting fast enough to hide every loss
- Its links are faster than the load, so no queue ever forms
A branch's sending process calls send() and it returns without an error. What does that tell the branch?
- That Lantern has received the message and applied the change
- That the message reached Lantern's machine and waits in its buffer
- That the bytes are in the local kernel's buffer and nothing more
- That the network has accepted the message and will not lose it
You got correct