Processes: A Private Machine
A process is the operating system's best illusion: every program believes it has a processor and a memory of its own. Neither is true. The kernel builds the illusion out of time slices, the subject of the next topic, and page tables, the subject of the one after, and it tears the illusion down and rebuilds it thousands of times a second on every core.
You start processes every day, from a shell, a container runtime or a web server's worker pool. This topic says what one is made of, and what it costs to create one, to switch between two, and to share memory across them. Each of those prices shows up in the memory graph or the latency graph of a service you run.
What a Process Owns
A process owns an address space: its code, its global data, the heap from Chapter 4 and one stack per thread, as in Chapter 3. It owns a saved copy of the processor's registers, including the program counter that says which instruction runs next, for the moments when it is not running. It owns a table of open files and network connections, each known to the program by a small number. And it carries an identity, a user and a set of permissions, that the kernel checks at every system call.
The kernel keeps all of this in one record per process. "The process" is that record plus the memory it points at. Nothing else is needed to stop a program in the middle of an instruction stream and resume it later, on the same core or a different one, without the program noticing.
The Illusion and Its Upkeep
A context switch is the moment the illusion changes hands. The kernel saves the running thread's registers into its record, picks the next thread to run, restores that thread's registers and, if it belongs to a different process, switches the address space as well. The direct cost is on the order of a microsecond or a few.
The indirect cost is larger and invisible. The processor's caches, and the translation cache the next topic but one describes, still hold the previous process's data. The newly running thread finds almost nothing it needs and refills them at the prices Chapter 4 set out: about a hundred nanoseconds per trip to RAM, for thousands of trips. A switch is cheap to perform and expensive to recover from.
Creating a Process: Copy by Promise
Unix splits creating a process into two ideas. Fork makes a copy of the running process. Exec replaces the copy's program with a new one. A shell running a command forks itself and then execs the command in the child. A web server that preloads its application forks workers and never execs at all.
Copying a whole process sounds expensive, and it would be if anything were copied. Instead, parent and child share every page of memory, and the kernel marks all of those pages read-only in both. As long as neither side writes, they keep sharing. The first write to a page faults, the kernel copies that one page, gives the writer its own copy, and lets the write continue. This is copy-on-write, and the page tables of the next topic but one are what make it possible.
Threads: Processes That Share
A thread is a second set of registers and a second stack inside the same address space. The kernel schedules threads, not processes, so a process with eight threads can run on eight cores at once. Switching between two threads of the same process skips the address-space change, and sharing data between them costs nothing at all: both see the same memory.
That free sharing is what makes the lost updates of Chapter 11 possible. It also means the unit of isolation is the process, not the thread. A thread that writes through a bad pointer corrupts memory every other thread is using, and a thread that triggers a fatal fault ends the whole process, including every request the other threads were serving.
Where the Copy-on-Write Promise Breaks
A pre-fork web server loads its application once and forks workers that share those pages. In a compiled language that sharing can last for days. In CPython it erodes by itself, because merely using an object writes to it. Every object carries a reference count, from Chapter 4, that changes when a variable starts or stops pointing at it, and the cyclic garbage collector writes bookkeeping into the objects it examines. Each such write lands on a shared page, and each shared page that is written gets copied.
So memory per worker climbs toward one full private copy of the application, with no change in the code. Instagram traced this effect in its Python web servers in 2017 and singled out the garbage collector's writes as the cause it could remove. CPython added a call that moves every existing object out of the collector's reach, gc.freeze(), in Python 3.7, and immortal objects, whose reference counts never change, in 3.12, and preventing this copying was among the stated reasons for both. The price came from a layer underneath the code.
The Cost in a System You Run
A server that runs one process per worker pays memory and gets isolation: a crash or a leak stays inside one worker. A server that runs one thread per request pays in races and gets cheap sharing. Chapter 11 opens with that choice, and Backend Deep Dive covers the practice in its topic Processes, Threads and the Event Loop.
The two ideas also collide. Python 3.14 stopped using fork as the default way to start worker processes for the multiprocessing module on Linux, and now starts them from a separate, clean server process. Forking a process that already runs threads copies every lock those threads held, into a child where the threads that would release them do not exist, and the child can wait on such a lock forever. To watch processes, their memory and their states on a real machine, see Linux Deep Dive, chapter Processes and Signals.
A process is an address space plus the resources the kernel tracks for it. Use one where a crash or a leak must stay contained.
A thread is a separately scheduled flow of execution inside a process, sharing all of its memory. Use threads where the work mostly waits and must share data.
A container is ordinary processes that the kernel gives a restricted view of the system and a budget of resources, sharing the host's kernel: no third kind of thing, and no small virtual machine either; Linux Deep Dive covers it in The Kernel and Containers.
- "Forking a 2 GB process copies 2 GB." It copies the page tables and marks every page shared and read-only. The child costs almost nothing until one side writes, though copying the page tables of a process that large still takes on the order of milliseconds.
- "Adding up each worker's resident memory gives the server's usage." Shared pages, such as the code, the preloaded application and mapped files, count in full in every worker that touches them. Eight workers reporting 500 MB each may be using 1.5 GB together, or close to 4 GB once copy-on-write has quietly broken the sharing.
- "Context switches are free, so more busy processes than cores costs nothing." Each switch costs a microsecond or more directly and more again in cold caches. Five hundred busy workers on 8 cores spend a visible share of their time switching, and each runs from caches the previous one filled.
- "A thread that crashes only takes itself down." Threads share one address space. A wild write corrupts everyone's memory, and a fatal fault ends the process and every request in flight inside it.
- "Threads are light and processes are heavy, so threads are always the better choice." Threads are cheaper to create and switch, and they pay for it with isolation. A process per worker contains crashes and leaks and, in CPython's default build, is the only way to run Python code on several cores at once, as the last topic of Chapter 11 explains.
- Load shared read-only data before forking workers, and measure whether it stays shared. Invisible writes, such as reference counts, break copy-on-write one page at a time.
- Size the number of CPU-bound workers to the number of cores, not to the number of requests. Every runnable worker beyond the cores adds switches, not throughput.
- Use processes where a crash, a leak or a CPU-bound job must stay contained, and threads where the work mostly waits and must share memory. The process is the unit of isolation.
- Set the multiprocessing start method explicitly in code that creates worker processes. The default changed in Python 3.14, and forking a process that runs threads can hang the child.
Knowledge Check
What does a context switch between threads of two different processes save, and what does it leave behind cold?
- It saves the whole address space to disk and leaves the registers cold
- It saves the registers and leaves the caches full of the old process's data
- It saves the caches into the process record and restores them all later on
- It saves nothing, because each process runs on its own dedicated core
A Python web server preloads its application and forks 8 workers. Right after start, the workers share most pages. What is the likely picture an hour later?
- Each worker's memory is unchanged, since the code itself never changes
- The workers use less memory, because the kernel merges identical pages
- Memory per worker has climbed toward a private copy of the application
- The parent's memory has doubled, because it holds a copy for each child
One thread in a multi-threaded server dereferences a bad pointer and triggers a fatal fault. What happens to requests being served by the other threads?
- They finish normally, because each thread has its own memory
- They are moved automatically to a newly started replacement thread
- They pause until the faulty thread is restarted by the kernel
- They all fail too, because the fault ends the whole process at once
What makes fork cheap, and what makes it unsafe in a program that already runs several threads?
- Pages are shared until written; locks held by other threads are copied held
- Only the stack is copied; the heap is freed in the child to save memory
- The kernel copies all memory at once, which stops the threads for too long
- Pages are shared until written; threads keep running in both processes
Eight workers each report 500 MB of resident memory. Why can the real total be far from 4 GB?
- Resident memory counts only the stack, so the real total is higher
- The kernel compresses idle workers, so the real total is always lower
- Shared pages are counted in full in every worker that touches them
- Each report includes the page tables of all eight workers together
You got correct