Chapter Twelve · Reliable Communication
Reliable Communication
How two machines get a message across a link that loses, duplicates, reorders and corrupts it. Five topics name the four failures and the argument that puts the fix at the endpoints, detect damage, build delivery from sequence numbers, acknowledgements and timeouts, pace the sender to what the receiver and the network can take, and end on what no protocol can do.
Chapter 11 kept one machine's threads from losing updates. This chapter keeps two machines from losing messages, and the problem changes character on the way. Threads share memory and at least see the same world. Two machines share only a channel, and the channel drops, repeats, reorders and occasionally alters what it carries. Nothing on the other side can be observed directly. Everything one side knows about the other arrives as a message that could have been lost.
The example throughout is Lantern's availability feed: each of the 30 branches reports checkouts and returns over the city network, and "copies on shelf" on a search result is computed from those reports. A lost message leaves a count too high, a duplicate leaves it too low, and neither raises an error. The chapter builds the fix one idea at a time: detect corruption, number and acknowledge every message, resend on a timeout, keep a window in flight without flooding the network, and make repeats harmless.
TCP is the real implementation of most of this, and it stays with Networking Deep Dive, which teaches the protocol itself. The practice of timeouts, retries and idempotency keys stays with Backend Deep Dive. Agreement among many machines stays with the System Design course. What this chapter owns is the idea underneath all three, and its price, which in a working system is paid as latency and as counts that are quietly wrong.