Topic 06

Just-in-Time and Why Speed Differs

Languages

The same pure-Python program can run several times faster on PyPy than on CPython, from the same source file. PyPy's own benchmark page puts the average across its suite at about three times the speed of CPython 3.11. Nothing about the language changed; the runtime did.

A just-in-time compiler watches the program run, notices which code is hot and what types actually flow through it, and compiles that code into machine instructions specialized for what it saw. It reaches speeds an ahead-of-time compiler cannot, because it knows the inputs. It pays in warm-up time, in memory and in sudden slowdowns when its guesses stop holding. This topic follows that trade from V8 and the JVM to what CPython itself does in 3.15.

A Language Does Not Have a Speed

A language is a specification: what a program means. CPython and PyPy for Python, or V8 and SpiderMonkey for JavaScript, are implementations with different machinery underneath. "Python is slow" really means "CPython's interpreter, running this code, pays the tax the interpreters topic measured", and the same program can be fast or slow depending on which runtime executes it.

That is not a technicality. It means the first question about a slow program is which runtime, which version and which workload, and only then which language. Rewriting in another language is one way to change the runtime, and the most expensive one.

Hot Code and Tiers

Most programs spend most of their time in a small share of their code. A JIT runtime starts by interpreting everything, counts how often each function or loop runs and compiles only what crosses a threshold. Compiling everything up front would waste effort on code that runs once.

Real runtimes use several tiers. V8 moves JavaScript through an interpreter called Ignition, a quick baseline compiler called Sparkplug, a mid-tier optimizer called Maglev and an optimizing compiler called TurboFan. The JVM's HotSpot, named for the idea, interprets first and then compiles with a fast compiler, C1, and later an optimizing one, C2. PyPy records traces of hot loops, the actual path taken through them, and compiles those.

One function's speed in a tiered runtime (schematic)
time since the function first ranspeed of one callinterpreterbaseline compileroptimizing compilerwarm: count passes thresholdhot: types observeda guard fails:deoptimizerecompiled

The figure is a schematic of one function's life. It starts slow in the interpreter. When its call count passes a threshold, a baseline compiler produces quick, plain machine code. When it is clearly hot, the optimizing compiler uses the types it has observed and the function steps up again. The drop near the end is what happens when those observations stop being true.

Speculation and Guards

The JIT's advantage is knowledge. An addition that has seen two integers ten thousand times can be compiled as a single integer add behind a cheap check that both operands are still integers. If a float arrives one day, the check, called a guard, fails, and execution falls back to the interpreter. The runtime may later recompile the code with a broader assumption. This fallback is called deoptimization.

An ahead-of-time compiler for a dynamic language must handle every type the language allows at that line. The JIT handles only the types that showed up, and pays for the rare exception with a slow path. It is the same guess, check and roll back shape as the branch prediction of Chapter 3: bet on what happened before, verify cheaply, undo when wrong.

CPython's Path

Python 3.11 added the specializing adaptive interpreter, described in PEP 659. It produces no machine code. Instead, instructions rewrite themselves in place into type-specialized versions once they have warmed up. The 3.11 release notes report an average speed-up of 1.25 times over 3.10 on the standard benchmark suite.

One function's addition, as dis prints it on Python 3.15 before and after warming up
def add(x, y):
    return x + y

BINARY_OP                0 (+)       # before any call
BINARY_OP_ADD_INT        0 (+)       # after 1,000 calls with integers
BINARY_OP_ADD_FLOAT      0 (+)       # after 1,000 more calls with floats

The listing shows the addition instruction of a two-line function at three moments. Before any call it is the generic addition. After a thousand calls with integers, CPython 3.15 had rewritten it into an integer-only addition with a guard. After a thousand calls with floats, the guard had failed often enough that the instruction was respecialized for floats. Nothing outside the function changed; the bytecode adapted to the types flowing through it.

Python 3.13 added an experimental JIT compiler, described in PEP 744. It uses a technique called copy-and-patch: machine-code templates are prepared with LLVM when CPython itself is built, and stitched together at run time, so users never need LLVM installed. In 3.15 it is still experimental. The release notes report an average gain of 8 to 9 percent on x86-64 Linux and 12 to 13 percent on ARM64 macOS, and mark those results as not yet final. It is not on by default: a build from source leaves it out unless configured to include it, and the official macOS and Windows installers of 3.14, the first to include it, shipped it switched off.

What a JIT Costs

Warm-up comes first. The first calls run in the interpreter, so a short-lived program, such as a command-line tool, a serverless function starting cold or a test run, can finish before anything is compiled, and pays for the profiling anyway. The JIT also costs memory for its profiles and its compiled code, and compilation work that competes with the program for processor time.

Then there are performance cliffs. A hot function that starts seeing a new type deoptimizes and can run slower than it did before it was compiled, until the runtime recompiles it. And there is compatibility: PyPy runs pure-Python code fast, but code that calls heavily into C extensions written for CPython goes through a compatibility layer and can run slower than on CPython.

Why the Same Program Runs at Different Speeds

Put the ladder together for one loop. CPython's interpreter pays dispatch and type checks on every operation. Its specializing interpreter removes some of the checks for types it has seen. A JIT removes dispatch and most checks for the types it saw. Ahead-of-time code in a statically typed language never had them. So a benchmark means nothing until it names the runtime, the version and whether it measured before or after warm-up.

The choice of runtime is a choice of where the price is paid: at build time, at start-up or on every operation. A long-running service can afford warm-up and gains the most from a JIT. A nightly batch job sits in between. A command-line tool that runs for a second wants ahead-of-time code or a plain interpreter, because it will be gone before any JIT pays off.

Interpreter vs Ahead-of-Time Compiler vs JIT

An interpreter, CPython's default, starts instantly and pays on every operation. Choose it where start-up matters and loops are not hot.

An ahead-of-time compiler, as for Go, Rust and C++, pays at build time and runs fast from the first instruction, but cannot use run-time types. Choose it for short-lived processes and predictable latency. A JIT, as in the JVM, V8 and PyPy, pays at start-up and in memory, and reaches the highest speed on long-running hot code. Choose it for services that run for hours.

Misconceptions
  • "Java and JavaScript are slow because they run on a virtual machine." HotSpot and V8 compile hot code to machine instructions, and a warmed-up service often runs within reach of ahead-of-time code. The virtual machine costs start-up time and memory, not steady-state speed.
  • "A JIT is always faster than an interpreter." A program that exits in a second may never reach compiled code and pays the profiling overhead for nothing. Command-line tools and cold serverless starts are where this bites.
  • "The first run of a benchmark shows how fast the code is." The first iterations measure the interpreter plus the compiler at work. A JIT runtime's steady state appears only after warm-up.
  • "Python has a JIT since 3.13, so Python is fast now." The JIT is experimental and not on by default in 3.15, and its reported average gain is under ten percent on x86-64 Linux. The bigger win so far came from the specializing interpreter in 3.11.
  • "Switching to PyPy speeds up any Python program." Code that spends its time inside C-extension libraries can run slower through PyPy's compatibility layer, and code that waits on input and output gains nothing.
Why It Matters
  • Benchmark after warm-up, and name the runtime and its version beside every number. A figure without them compares nothing.
  • Keep the types flowing through a hot path stable. A loop fed integers and then floats deoptimizes and loses what the JIT built.
  • Match the runtime to the process lifetime. A long-running service can amortize warm-up; a one-second tool cannot.
  • Try an alternative runtime on the unchanged code before rewriting in another language. The implementation, not the language, often sets the speed.
RelatedAhead-of-time compilation all the work before the first run (previous topic)Caching a JIT is a cache of compiled code, keyed by the assumptions it was compiled under (Chapter 1)Branch prediction the hardware makes the same bet with the same rollback (Chapter 3)

Knowledge Check

A team moves a short command-line tool that runs for 200 milliseconds from CPython to a JIT runtime. What is the likely result?

  • It runs several times faster, since the JIT compiles all of its hot loops
  • Little or no gain, and possibly slower, since warm-up never finishes
  • It runs about as fast as an ahead-of-time compiled version of the same tool
  • It fails outright, since JIT runtimes cannot run short-lived processes at all

A JIT compiled an addition as an integer add. A float arrives at that line. What happens?

  • The program raises a TypeError, since the compiled code expects only integers
  • The integer add runs anyway and produces a wrong result
  • The guard fails, execution falls back, and the code may be recompiled
  • The runtime quietly converts the float to an integer and continues running

Why does the first iteration of a benchmark on a JIT runtime mislead?

  • It measures the interpreter and the compiler at work, not the compiled code
  • It is faster than the later iterations, because the processor caches are still empty
  • It measures the garbage collector, which only runs at start-up
  • It includes the time taken to download the runtime's compiler first

What does PEP 659's specializing adaptive interpreter do?

  • Compiles hot Python functions into machine code for the processor
  • Checks type hints before running each function and rejects mismatches
  • Converts Python source into C and compiles that ahead of time
  • Rewrites warmed-up bytecode into versions specialized for the types seen

A data pipeline spends most of its time inside a C-extension numeric library. Why can PyPy make it slower rather than faster?

  • PyPy cannot load C extensions at all, so it falls back to slower pure-Python code
  • PyPy's JIT compiles the extension's C code all over again, which takes longer to run
  • The extension runs through a compatibility layer, and that crossing costs time
  • PyPy lays out objects to use less memory, and that layout slows down numeric code

You got correct