The Instruction Cycle
A processor does one thing forever: fetch the instruction at the address in the program counter, decode what it asks for, execute it, and move on to the next. Every program ever run, an operating system, a Python interpreter, a game, is this loop over a list of simple instructions that each move, add or compare a few values.
The list of instructions a processor understands is its instruction set, and it is the contract between software and silicon. It is why a program built for one family of processors does not run on another, and why two processors with the same clock speed can differ several times over on the same work.
Registers and the Program Counter
A register is a named slot inside the core that holds one 64-bit value and can be read or written in under a nanosecond. A core has only a few dozen of them: 16 general-purpose registers on x86-64, 31 on 64-bit ARM. Everything the core computes passes through registers. Values are loaded into them from memory, combined by the arithmetic and logic unit, and stored back.
One register is special. The program counter holds the address of the next instruction to run. It is the core's entire notion of where it is in the program, and changing it is the only way the core ever goes anywhere else.
Fetch, Decode, Execute
Fetch reads the instruction stored at the address in the program counter. Decode works out which operation it is and which registers or memory address it names. Execute runs it, through the arithmetic unit or by reading or writing memory. Then the program counter advances to the next instruction and the loop starts again.
A jump or a branch is an instruction that writes a new address into the program counter instead of letting it advance. That is all a loop, an if statement or a function call is at this level. A loop is a branch back to an earlier address while a condition holds. An if is a branch forward past the code that should not run. A function call is a jump that first saves where to come back to, which the next topic follows in detail.
What One Instruction Does
Instructions are small. Load a value from memory into a register. Add two registers. Compare two registers. Jump to another address if the last comparison said so. Store a register to memory. A useful program is millions of these, and a line of source code is usually several of them; a Python loop body becomes hundreds of them through the interpreter, as Chapter 13 shows.
Not all of them cost the same. An instruction that works on registers finishes in a fraction of a nanosecond. An instruction that touches memory can wait a hundred times longer if the data is not in a cache, the RAM row of Chapter 4's ladder. That single fact decides more about a program's speed than the number of instructions it runs.
The Instruction Set as a Contract
An instruction set architecture defines the instructions, the registers and the binary encoding that a family of processors promises to understand. x86-64, from Intel and AMD, runs most servers and PCs and uses variable-length instructions from 1 to 15 bytes long. 64-bit ARM runs phones, Apple's Mac processors and a growing share of cloud servers, and uses fixed 4-byte instructions.
The contract is what lets hardware change underneath software. A chip can be redesigned completely, with a new pipeline, new caches and a new manufacturing process, as long as it keeps decoding the same instructions to the same effects. That is how an x86 program compiled decades ago still runs on a processor built this year.
One Contract Is Not Another
Machine code for x86-64 means nothing to an ARM core, and the reverse, because the same bytes decode to different instructions. The four bytes of Chapter 2's first topic were four increments to a 32-bit x86 processor and one vector instruction to a 64-bit ARM one. So every compiled artefact is built per architecture: a native binary, a Python wheel that contains a compiled extension, a container image.
Running the wrong one either fails at once or goes through a translation layer that converts one instruction set to the other, at a cost in speed. Deploying an x86-only image to an ARM fleet, or an ARM laptop's build to x86 servers, is one of the most common ways this contract shows up in daily work.
Clock Speed Is Not Instructions per Second
A modern core decodes several instructions per tick and runs independent ones at the same time, which the topic on pipelines covers. On friendly code it finishes several instructions per cycle. On code where every instruction waits for main memory it finishes one every few hundred cycles, because a hundred-nanosecond trip is a few hundred ticks at 3 gigahertz.
The price of a program is therefore instructions times cycles per instruction, divided by the clock rate, and the middle term varies far more than the clock. It is set mostly by the memory layer in Chapter 4. Two 3-gigahertz chips can therefore differ several times over on the same work, and a faster clock does little for a program that spends its time waiting for memory.
A machine instruction is executed by the processor's circuits directly, usually in a fraction of a nanosecond when its data is in registers or cache. It is what this topic's loop fetches and decodes.
A bytecode instruction is data read by the Python interpreter, a program that itself runs as machine instructions. One bytecode, such as "add these two objects", costs dozens to hundreds of machine instructions. Python's disassembler shows the second kind, never the first, which Chapter 13 explains.
- "The CPU runs my Python code." It runs the CPython interpreter, a compiled program that reads your bytecode as data and acts on it. Your code is two steps removed from the instruction cycle, and that distance is where the tens-to-hundreds-fold gap to compiled code comes from.
- "One line of code is one instruction." A single Python line can execute hundreds of machine instructions. Even in a compiled language a line can become many instructions, or none at all after the compiler optimizes it away.
- "A higher clock speed means proportionally faster code." Instructions per cycle varies more than clock speed between chips and between workloads. A program that waits on memory runs at the speed of memory, not of the clock.
- "A container image runs on any machine." The image contains machine code for one instruction set. An x86-64 image on an ARM host fails, or runs under emulation several times slower, unless it was built for both.
- "x86 is complex and ARM is simple, so ARM is faster." Both families decode their instructions into simple internal operations and use the same speed-up techniques. Speed and power differences come from each chip's design, not from the name of its instruction set.
- Build and publish binaries, wheels and container images for every architecture you deploy to. An ARM fleet cannot run x86-only artefacts at native speed.
- Judge a processor by how fast it runs your workload, not by its clock. Instructions per cycle and memory behaviour decide more than gigahertz.
- Count a hot loop's memory accesses before its instructions. An instruction that waits on memory costs as much as a hundred that do not.
- Treat interpreted bytecode as a layer above the instruction cycle, not as instructions. Reducing the bytecode count helps; imagining it as machine code misleads.
Knowledge Check
At the machine level, what does a loop's branch back to its start actually do?
- It copies the loop's instructions again into a fresh part of memory
- It writes the loop's first address into the program counter again
- It resets every one of the registers to zero so the loop starts cleanly
- It asks the operating system to schedule the loop's body one more time
A team builds its service on x86-64 laptops and deploys the same image to a new ARM server fleet. What happens?
- It runs at full speed, since containers are architecture-neutral
- The kernel recompiles the image for ARM the first time it starts
- It fails, or runs under emulation at a fraction of native speed
- Only the Python files run; compiled extensions are skipped
Two servers both run at 3 GHz, yet one finishes the same job three times faster. Which explanation fits best?
- They complete very different numbers of instructions per cycle
- One of the clocks ticks faster despite the matching label
- The faster server runs fewer instructions for the same code
- The faster server uses a newer version of the same clock
How far is a line of Python from the processor's instruction cycle?
- One step: each line is compiled straight into machine code
- None: the processor reads and runs the Python text directly
- One step: the operating system translates each bytecode
- Two steps: the line becomes bytecode that an interpreter reads
You got correct