Topic 02

Race Conditions

Concurrency

Lantern shows "searches today" on the library's staff dashboard, and on a busy Saturday the number is lower than the count of searches in the request log. No search was dropped, nothing was logged as an error, and every line of the counting code is correct on its own. The bug is that adding one to a counter is not one step but three, read, add and write back, and two workers can interleave those steps so that two increments produce one.

This is the lost update, the pattern behind every concurrency bug in this chapter. It never crashes, it rarely shows up in tests, and it gets worse when traffic is highest. This topic takes it apart step by step, runs it on Python 3.15, and names the three families of fix.

Three Steps, Not One

At the machine level an increment is a load from memory into a register, an add and a store back to memory, the instruction cycle of Chapter 3. At the Python level it is several bytecode instructions, the interpreter's own instruction set that Chapter 13 describes. Either way there is a gap between reading the old value and writing the new one. Anything another worker writes into the same variable during that gap is overwritten and lost.

One increment is three steps with a gap in the middle
readcount is 41add41 + 1 = 42, in a registerwritecount becomes 42the gapanything another worker writes here is overwritten
What Python 3.15 compiles one "count += 1" into (from the dis module)
LOAD_GLOBAL              0 (count)
LOAD_SMALL_INT           1
BINARY_OP               13 (+=)
STORE_GLOBAL             0 (count)

The listing above is what Python 3.15 prints when asked to disassemble a function whose body is that one line, with count a global variable. It is four instructions: load the current value of the global, load the constant one, add them and store the result back into the global. The read is the first instruction and the write is the last. Two instructions sit between them, and a thread switch in either place opens the gap.

The Lost Update, Step by Step

Worker A reads 41. Worker B reads 41. A adds one and writes 42. B adds one and writes 42. Two searches happened, the counter moved by one, nothing failed and nothing was reported. Had B read after A's write, it would have seen 42 and written 43. The only difference between the right answer and the wrong one is timing that neither worker controls.

The same two increments, two interleavings
Interleaved: one update lostworker Aworker Bt1read 41t2read 41t3write 42t4write 42counter ends at 42One after the other: both countedworker Aworker Bt1read 41t2write 42t3read 42t4write 43counter ends at 43
Two threads, a million increments each, one shared counter
import threading

count = 0

def work():
    global count
    for _ in range(1_000_000):
        count += 1

threads = [threading.Thread(target=work) for _ in range(2)]
for t in threads: t.start()
for t in threads: t.join()
print(count)   # should be 2000000

The program above starts two threads that each add one to a shared counter a million times, waits for both and prints the total, which should be two million. What it prints depends on the build of Python that runs it, and the next section but one explains why. On the free-threaded build of Python 3.15, where threads really run at the same time, it lost updates on every run on the machine used to check this page, printing totals between about 1.1 and 1.6 million. On the default build of Python 3.15 it printed exactly two million in every one of a hundred runs, even with the interpreter told to switch threads as often as it can.

Lantern's Counter

Lantern's worker processes keep the total in a shared store. For each search a worker reads "searches today", adds one and writes it back. Here the gap is a network round trip to the store plus the store's own work, about a millisecond, instead of two bytecode instructions.

At the Saturday peak of 300 searches a second, another increment arrives inside any given millisecond gap with a probability of roughly one in four: on average 0.3 arrivals land in a millisecond, and the chance of at least one is about 26%. That is a back-of-envelope estimate, not a measurement, but it predicts the shape of the bug. The undercount grows with traffic, so it hides on quiet weekdays and in tests with one user, and shows up on the busiest morning of the week.

Why the Race Is Hard to See

A race needs the scheduler to switch at exactly the wrong instruction. Since Python 3.10, a thread on CPython's default build hands the interpreter to a waiting thread only when it blocks, or at a few fixed points: when a function is called or entered, and at the backward jump at the end of each loop iteration. None of the four instructions of a plain increment is such a point, so on the default build the classic two-thread demo prints the right answer, as it did on 3.15 when this page was checked.

Change the code slightly and the protection vanishes. Put a function call between the read and the write, for example writing the increment as a call to a helper that returns the value plus one, and the same default build of 3.15 lost between about 280 thousand and 760 thousand of the two million updates in the same test. Run the original on the free-threaded build and it loses them too. Put the read and the write on opposite sides of a network call or an await, as Lantern does, and it loses them on any build in any language. A correct result in a test is evidence of luck, not of safety.

Check-Then-Act and Other Shapes

The lost update is a read-modify-write race. The same gap appears in check-then-act: "if the file does not exist, create it", "if the seat is free, book it", "if the username is not taken, register it". Two workers both check, both see the condition hold and both act. It also appears in any rule that spans two variables, such as a balance and its history updated in two separate steps, where a reader between the steps sees one without the other.

The general definition follows: a race condition is any result that depends on timing the program does not control. Everything else in this chapter is a way of taking that timing out of the result.

Three Fixes and Their Prices

There are three families of fix. Make the step indivisible: let the store perform the increment itself, in one operation, or use an atomic instruction of the kind the fifth topic of this chapter describes. Make it exclusive: hold a lock around the read and the write, the subject of the next topic. Or stop sharing: let each worker count its own searches and have the dashboard add the counts up when it reads them. Each has a price, and the last is often the cheapest, because it has no gap and no waiting.

The same race sells one seat twice and spends one balance twice. In a service, that is where Backend Deep Dive takes over, in its Data Access chapter's topic on the oversell, and where PostgreSQL Deep Dive takes over with transactions and constraints.

Misconceptions
  • "The GIL makes Python code thread-safe." The global interpreter lock protects the interpreter's own structures. Nothing in the language promises that it will not switch threads between the read and the write of your increment, a helper call in between already exposes the gap, and the free-threaded build removes even the accidental protection.
  • "I ran it a million times and never lost an update, so it is safe." Races depend on switch timing, core count, load and interpreter version. A test on a quiet laptop samples almost none of the interleavings production will produce.
  • "One line of code is one indivisible operation." Adding one to a counter, incrementing a dictionary entry with its get method, and "if x is not in the set, add it" are each several steps with a gap between the read and the write.
  • "Races only happen between threads." Two processes, two containers or two servers doing read-modify-write against one file, row or key race the same way, and their gap is a network round trip, tens of thousands of times wider than a thread's.
  • "A race shows up as a crash or an exception." A lost update produces a plausible, slightly wrong number and no error. It is found by comparing against an independent count, usually weeks later.
Why It Matters
  • Find every read-modify-write on shared state and make it atomic, locked or unshared. The gap between the read and the write is the bug, wherever it lives.
  • Let the owner of the data do the increment. An increment executed inside the store or the database has no gap in the application to lose it in.
  • Prefer not sharing over sharing carefully. Per-worker counters summed on read cannot lose an update and need no lock.
  • Verify concurrent code under stress and against an independent count, never by one clean run. A clean run on one build proves nothing about the next.
RelatedData race unsynchronized access to the same memory, related but not identical to a race conditionTransactions the same problem solved inside the store (PostgreSQL Deep Dive)Idempotence making repeated or overlapping work harmless (Chapter 12)

Knowledge Check

Two workers each add one to a counter that holds 41. Both read before either writes. What does the counter hold afterwards?

  • 42, because both wrote the value they each computed from 41
  • 43, because the store applies both increments in order
  • 41, because the two writes cancel each other out entirely
  • It is undefined, and reading it again raises an error

Which of these Python statements, run by two threads on a shared object, is free of a gap between reading and writing?

  • Adding one to a shared integer with the plus-equals operator
  • Incrementing a dictionary entry by reading it with get and adding 1
  • Adding an item to a set only after checking it is not already there
  • None of them: each one reads a value and writes it back later

Why does Lantern's "searches today" undercount grow as traffic grows?

  • The shared store drops writes when it gets busy, to protect itself
  • More increments arrive inside each millisecond gap between read and write
  • Each worker process caches the counter locally and flushes it only at night
  • The network reorders the writes at busy times, which reverses them

Which fix is usually the cheapest for a dashboard counter shared by many workers?

  • A lock around every increment, held across the call to the shared store
  • Retrying each increment until the total read back equals the expected value
  • Counting per worker and adding the counts up when the dashboard reads them
  • Running a single worker process so that only one of them ever updates the counter

The classic two-thread "count += 1" demo prints exactly two million on Python 3.15's default build. What is the right conclusion?

  • The global interpreter lock makes plus-equals an atomic operation in Python
  • Python 3.15 fixed lost updates, so shared counters no longer need a lock
  • The interpreter happened not to switch threads inside the increment
  • The two threads ran one after the other, never at the same time as each other

You got correct