The Kernel and the System Call
Every program you have written lives in a sandbox. It can compute anything, but it cannot touch a disk, a network card or another program's memory. To do any of those it has to ask the kernel, and asking is not a function call. It is a deliberate trap that switches the processor into a privileged mode, runs kernel code and switches back.
That round trip costs on the order of a hundred times a function call before any work is done. Most of the input and output advice you have heard, from "buffer your writes" to "batch your queries" to "avoid tiny packets", is this one price paid over and over. This topic shows where the price comes from and how to stop paying it per item.
Two Modes on One CPU
A processor runs in one of two modes. In user mode, where your program runs, the instructions that talk to devices, change memory mappings or reconfigure the processor are forbidden: attempting one causes a fault. In kernel mode everything is allowed. The mode is one bit of processor state, not a second computer, and the kernel is ordinary machine code running on the same cores as your program.
Kernel code runs only when something sends the processor across the line, and there are exactly three ways to do that. The program asks, with a system call. A device signals, with an interrupt, such as a network card announcing that a packet has arrived. Or an instruction fails, with a fault, such as a division by zero or a touch of memory that is not mapped. Every entry into the kernel is one of these three.
The Trap
A system call is a numbered request. The program puts the call's number and its arguments into registers, the processor's own small set of named slots that Chapter 3 introduced, and executes one special instruction. The processor switches to kernel mode and jumps to a single fixed entry point that the kernel chose at boot. The kernel looks the number up in a table, checks every argument, does the work, and returns a result or an error code, switching the mode back on the way out.
The checking is not optional. A pointer handed in from user space might point at kernel memory, at another part of the program that the caller is not allowed to overwrite, or at nothing at all, and the kernel treats every such value as hostile until it has verified it. The program cannot jump into the kernel anywhere else, and that single guarded door is all that keeps one program out of another.
What a Crossing Costs
A function call inside your program costs a nanosecond or two. The simplest possible system call, one that asks the kernel for a number it already knows, costs on the order of a hundred nanoseconds on current hardware, and a few hundred on machines running the default protections against the processor attacks disclosed in 2018. The crossing saves and restores registers, switches stacks, and disturbs the caches and branch predictors from Chapter 3 that your program had warmed up. The program then runs cold for a while after it returns.
An empty crossing costs that much. A call that waits on a device adds whatever the device costs, from Chapter 4's ladder: up to a hundred microseconds for a solid-state drive, milliseconds for a spinning disk. A few calls are so common and so simple, reading the current time being the usual example, that mainstream kernels answer them without a trap at all, from a page of memory the kernel keeps updated inside every process. That exception exists because the trap is worth avoiding.
Buffering Is Amortization
Write 100 MiB one byte per call and the program makes about 105 million system calls. At a hundred nanoseconds each, that is roughly ten seconds spent entering and leaving the kernel, with no useful work counted at all. Write the same bytes in 64 KiB chunks and it makes 1,600 calls, less than a fifth of a millisecond of crossing overhead. The bytes are identical. The number of crossings differs by a factor of about 65 thousand.
A buffer is how programs get the second number while writing code that looks like the first. Small writes go into a block of the program's own memory, and only when the block fills does the program cross into the kernel, once, with all of it. This is the same trade as the amortized append in Chapter 1: an occasional expensive operation, spread across many cheap ones. Python's file objects buffer by default, and output to a terminal is buffered one line at a time, so a print statement's output appears a line at a time on a screen and in large bursts when the same program writes to a file.
with open("out.bin", "wb", buffering=0) as f: # no buffer: every write crosses for _ in range(1_048_576): f.write(b"x") with open("out.bin", "wb") as f: # default: a buffer in the process for _ in range(1_048_576): f.write(b"x")
The two loops above write the same mebibyte, one byte at a time. The first opens the file with buffering turned off, so every write becomes a system call. The second uses Python's default, a buffer of at least 128 KiB in Python 3.15, so the same million writes turn into a handful of crossings. On the machine used to check this page, running Python 3.15, the first loop took about two seconds and the second about a twentieth of a second: forty times slower for the same bytes. Each unbuffered write there cost about two microseconds, well above the empty-crossing price, because a write to a file also does file-system work inside the kernel every time.
Crossings Hidden in Ordinary Code
The same arithmetic hides in ordinary code. A logger that flushes after every line makes one crossing per line. A loop that sends each small message as its own network write makes one crossing per message, and usually one tiny packet per message as well. A database driver that makes one round trip per row pays the crossing, the network and the database's own overhead a thousand times for a thousand rows. A web server that reads a large file in 4 KiB pieces makes sixteen times as many calls as one reading 64 KiB pieces.
The fix is the same move in every case: fewer, larger crossings. Batch the rows, buffer the log, send the messages together, read in big chunks. To count the system calls a real process makes, the tracing tools in Linux Deep Dive, in its chapter Performance and Troubleshooting, show every crossing as it happens.
A library call, such as Python's print, is an ordinary function in your program's own memory. It may make no system call, one or several, and a buffered one makes a system call only when its buffer fills.
A system call, such as write, is the crossing into the kernel itself. Reasoning about the cost of input and output starts from counting the second kind, not the first.
- "A system call costs about as much as a function call." The crossing alone is roughly a hundred times dearer before any work is done, which is why a loop of a million tiny reads spends most of its time entering and leaving the kernel.
- "The kernel is a separate program running next to mine, serving requests." There is no kernel process waiting in a queue for your call. After the trap, kernel code runs on your thread's own time slice, so time spent in system calls is charged to the program that made them. A few background kernel threads exist, but they are not what serves your call.
- "Unbuffered output is safer and only slightly slower." On small writes it can be many times slower, and the safety is mostly illusory. Bytes handed to the kernel are still not on the disk, as the last topic in this chapter shows, so unbuffered output shrinks the loss window only for a crash of the program, not of the machine.
- "Python is slow at input and output because the interpreter is slow." A byte-at-a-time loop pays the same crossing per call in any language. The interpreter adds tens of nanoseconds per iteration, the system call adds a hundred or more, and the fix is the number of calls, not the language.
- "Opening or checking a file is one cheap operation." Resolving a path, checking permissions, allocating a handle and reading metadata can take several system calls, and walking a directory tree of a million files makes millions of them. The count is invisible in the source code.
- Count kernel crossings before optimizing an input or output path. The number of system calls, not the number of lines of code, predicts its cost.
- Batch small writes into one buffer and flush on a size or time boundary. Amortization turns a per-item crossing into a per-batch one.
- Flush deliberately where visibility or durability matters, and nowhere else. A flush per log line buys little safety and costs a crossing each time.
- Treat any per-item round trip to the kernel, the network or a database as one smell. A loop with one crossing per item is a batch waiting to be written.
Knowledge Check
A program writes 100 MiB one byte per system call. Another writes the same bytes in 64 KiB chunks. Roughly how much time does each spend on crossings alone, at about 100 ns per crossing?
- About ten seconds for both, since the byte count is the same
- About ten seconds against a fifth of a millisecond or less
- About a tenth of a second against a thousandth of a second
- About ten seconds against roughly five seconds for the chunks
Why does the kernel check every pointer a program passes to a system call?
- Because pointers from user space are usually misaligned in memory
- Because checking is how the kernel measures the call's cost
- Because the pointer could aim at kernel memory or at nothing
- Because the processor refuses to switch modes on bad pointers
A script's print output appears line by line on a terminal but arrives in large bursts when the output is piped to a file. What explains the difference?
- The file system delays writes to files until the file is closed
- Output to a terminal is line-buffered, to a file block-buffered
- The terminal forces a system call per character written to it
- Piped output is compressed first, which takes extra time per line
A service flushes its log file after every line "for safety". What does each flush guarantee?
- That the line is on the disk and will survive a power loss
- That the line is compressed and written as part of the next block
- Nothing at all, since the kernel ignores flushes on log files
- That the line survives a crash of the program, not of the machine
Mainstream kernels answer a request for the current time without a trap. What does that design say about the trap?
- That reading the clock needs no protection, unlike other calls
- That the crossing is costly enough to be worth engineering away
- That the trap is cheap, so skipping it is only for tidiness
- That the clock lives in a device the program may access directly
You got correct