Event Loops, Async and Message Passing
A thread per connection costs a stack and a kernel context switch each, and ten thousand connections cost ten thousand of them. An event loop serves the same ten thousand connections from one thread by never waiting on any of them: it asks the kernel which connections are ready and runs only those.
Async and await are how Python writes such a loop. Channels and actors are how other languages avoid sharing memory at all. And the global interpreter lock, which decided for three decades what Python threads could do, is now optional. Each model has a price, and the chapter closes on how to choose between them.
Blocking vs Non-Blocking Reads
A blocking read parks the thread until data arrives, and the scheduler of Chapter 10 runs something else meanwhile. A non-blocking read returns at once, with the data or with "nothing yet". On its own that only moves the waiting into a busy loop. The missing piece is a system call that waits on thousands of connections together and reports which of them are ready. Linux calls it epoll, and the BSDs and macOS call it kqueue; both are named here and not taught.
A loop around that call is an event loop. It asks the kernel for the ready connections, runs the handler for each one until the handler has to wait again, and goes back to ask. One thread, one kernel call per round and as many connections as memory allows.
Async/Await as Cooperative Scheduling
A coroutine runs until it reaches an await on something that is not ready, and then hands control back to the loop, which runs another coroutine. Resuming a coroutine costs about as much as a function call. With the event loop's own bookkeeping around it, a switch between busy tasks measured under a microsecond on Python 3.15, and it never enters the kernel. A waiting coroutine holds only its own small frame: measured on Python 3.15, an idle task with its coroutine takes about a kilobyte, while every thread reserves a whole stack.
The price is cooperation. The code between two awaits runs without interruption, so a CPU-heavy function, or a call that blocks instead of awaiting, freezes every other connection on the loop until it returns.
import asyncio, time async def polite(name): for _ in range(3): await asyncio.sleep(0.1) # waits, and lets the loop run others print(name, "tick") async def rude(): await asyncio.sleep(0.05) time.sleep(1) # blocks the whole loop for a second async def main(): await asyncio.gather(polite("a"), polite("b"), rude()) asyncio.run(main())
The program above runs three coroutines on one event loop. Two of them are polite: they wait a tenth of a second at a time with an await, and print a tick after each wait. The third sleeps for a second with an ordinary blocking call. On Python 3.15 the polite coroutines' first ticks, due at a tenth of a second, arrived at about 1.1 seconds, after the blocking call returned. For that whole second nothing else on the loop ran. On a server the two polite coroutines are other users' requests.
Races Without Threads
One loop runs one coroutine at a time, so a read-modify-write with no await inside it is safe: no other coroutine can run between the read and the write. That is a real advantage over threads, where a switch can land anywhere.
An await between the read and the write reopens the gap of this chapter's second topic, the same way a thread switch would. On Python 3.15, a thousand coroutines that each read a shared counter, await once and write back the value plus one leave the counter at 1, not 1,000. The difference from threads is that the switch point is written in the code, in the word await, for anyone reviewing it to see.
Channels and Actors
Instead of sharing memory under locks, tasks can send each other messages. A channel is a queue between tasks: Go builds its concurrency around channels, and Python's asyncio has a queue that plays the same role. An actor is a task that owns its state outright and handles one message at a time from its mailbox, as in Erlang and Akka.
The lost update disappears, because only one task ever touches each piece of state. New costs appear in its place: messages are copied, mailboxes grow until memory runs out when a consumer falls behind, and two tasks each waiting for the other's reply deadlock the way two threads with two locks do.
The GIL and Free-Threaded Python
CPython's default build lets one thread execute Python bytecode at a time, enforced by the global interpreter lock. Threads still overlap their waiting, because the lock is released around blocking input and output and inside many C extensions, but CPU-bound Python threads take turns. Measured on Python 3.15's default build, four threads each running the same pure-Python loop took four times as long as one thread running it once: no parallelism at all.
PEP 703 made the lock removable. Python 3.13 shipped an experimental free-threaded build, and in 3.14, under PEP 779, that build became officially supported but not the default; it is still not the default in 3.15. It runs Python threads in parallel: on the same machine, the four threads took about one and a half times as long as one thread instead of four times. It costs single-threaded code a few percent, which the 3.15 documentation puts at about 1% on ARM macOS to about 8% on x86-64 Linux. It needs C extensions built for it, and importing one that is not re-enables the lock for the whole process, with a warning. And it exposes every race the lock had been hiding, as the second topic of this chapter showed.
Choosing a Concurrency Model
An async web service holds thousands of idle connections on one core and stalls, for every user at once, when one handler parses a 50 MB JSON body on the loop. A threaded service spends memory on stacks, gains on CPU-bound work only with free threads or processes, and needs locks around everything it shares. A service built on processes isolates every worker and pays in memory and in copying messages between them.
None of these is right for every workload. The question from the first topic of this chapter decides it: is the work mostly waiting, or mostly computing? Backend Deep Dive covers the practice of choosing per workload in its topic Processes, Threads and the Event Loop.
Async interleaves many waiting tasks on one thread. It is the cheapest per connection, no help for CPU work, and one blocking call stalls every task. Use it for thousands of mostly idle connections.
Threads are interleaved or run in parallel by the kernel's scheduler. Each costs a stack and a switch, and shared memory needs locks. Use them for a few blocking calls, or for CPU work on the free-threaded build.
Processes isolate completely and run in parallel on any Python build, at the cost of memory and of messages between them. Use them for CPU-bound work that must also be contained. For mixed work, run a loop that hands CPU jobs to a pool.
- "Async makes code run in parallel." An event loop runs one coroutine at a time on one thread. Async overlaps waiting, not computing, and a CPU-bound task runs no faster.
- "Async is always faster than threads." For a few dozen connections, or for CPU-bound work, it adds overhead and complexity for nothing. Its win is thousands of mostly idle connections.
- "A blocking call inside one coroutine only slows that request." It stops the loop, so every connection on it waits. The symptom is latency jumping on unrelated endpoints at the same moment.
- "The free-threaded build makes my threaded code N times faster." Only CPU-bound pure-Python work speeds up. Single-threaded code runs a few percent slower, an extension not built for free threading turns the lock back on for the whole process, and races the lock used to hide begin to fire.
- "Message passing removes concurrency bugs." It removes shared-memory races and keeps the rest. Two actors each waiting on the other's reply deadlock, an unbounded mailbox is an out-of-memory crash waiting for a slow consumer, and ordering across different senders is not guaranteed.
- Keep every call on an event loop non-blocking, and hand CPU work and blocking libraries to a thread or process pool. One blocking call stalls every connection on the loop.
- Choose the model by what the work waits on. Thousands of idle connections: a loop. CPU: processes or free threads. A few blocking calls: threads.
- Bound every queue and mailbox. Pushing back at a full queue beats an out-of-memory crash, the same flow control Chapter 12 builds into networks.
- Run the test suite under the free-threaded build before calling Python code thread-safe. It surfaces the races the lock used to hide.
Knowledge Check
One coroutine on an async web server parses a large JSON body with an ordinary blocking function. What happens to the other requests on that loop?
- They continue normally, because each coroutine has its own thread
- They all wait until the parse finishes, because the loop is stuck
- They are moved to another core by the kernel's scheduler
- They are cancelled by the loop after a short default timeout
A coroutine reads a shared counter, awaits a network call, then writes back the value plus one. Is it safe on a single-threaded event loop?
- Yes, because an event loop runs only one coroutine at a time on one thread
- Yes, because await makes the whole function atomic until the coroutine returns
- No, because the await lets other coroutines update the counter in that gap
- No, because asyncio runs each coroutine on a separate thread from a pool
What does the global interpreter lock of CPython's default build serialize?
- All input and output, so two threads cannot wait on the network at once
- The execution of Python bytecode, so CPU-bound threads take turns
- Every access to a shared object, so Python code needs no locks of its own
- Only the start of threads, so running threads are free to go in parallel
When does moving a service to the free-threaded build fail to speed it up?
- When the service runs CPU-bound pure-Python work on several threads
- When the service runs on a machine that has more cores than it has threads
- When the work mostly waits on the network, or a C extension turns the lock back on
- When the service uses threading.Lock, which the free-threaded build removes entirely
What does an actor model give up in exchange for having no shared memory?
- Nothing: with no shared state, deadlock and overload become impossible
- Parallelism: actors must all run on one thread to keep messages ordered
- Correctness: updates to an actor's own state can still be lost in races
- Cheap sharing: messages are copied and mailboxes can grow without bound
You got correct