Topic 06

Many Cores, SIMD and Accelerators

Hardware

For thirty years each new processor ran old code faster, mostly by raising the clock. Around 2005 that stopped. Faster clocks made chips too hot to cool, and the growing supply of transistors went into more cores, wider arithmetic units and specialized accelerators instead.

The speed is still there, but it is no longer free. A program gets it only if its work can be split into independent pieces, and the kind of split decides which hardware can use it: many cores for many independent tasks, vector units for the same arithmetic over many values, a GPU for enormous amounts of identical work. This topic matches each kind of work to its hardware and prices the mismatch.

The End of Free Clock Speed

For decades, shrinking transistors let each generation run faster at the same power per area, a relationship known as Dennard scaling. In the mid-2000s it broke down: as transistors kept shrinking, current leaking through them when they were supposed to be off, and the heat of switching them faster, stopped falling in step. Clock speeds levelled out at roughly 3 to 5 gigahertz, where they still are.

Clock speed flattened; transistor counts did not
197519851995200520152025mid-2000stransistors per chip: tens of billionsclock: flat at roughly 3 to 5 GHzlog scaleThe shape of the trend, not data for any one chip

Transistor counts did not level out. They kept growing, into the tens of billions on current large chips, so the question for chip designers changed from how fast to run one instruction stream to what to spend all those transistors on. Their answers are the rest of this topic, and each one only helps a program shaped to use it.

More Cores

The first answer was more complete processors on one chip. A laptop now has roughly 8 to 16 cores, and a single server processor can have more than a hundred. Each core is a full copy of everything the earlier topics in this chapter described, with its own pipeline, predictor and registers.

A single-threaded program uses exactly one of them. The rest help only if the work is split into threads or processes that can run at the same time, which is Chapter 11. That chapter also gives Amdahl's law: the part of a job that cannot be split sets a hard ceiling on the speed-up, whatever the core count.

Hardware Threads and the vCPU

A core spends much of its time stalled, waiting for memory, as the previous topics showed. Simultaneous multithreading, which Intel markets as Hyper-Threading, lets one core run two instruction streams at once and fill one stream's stalls with the other's work. The operating system sees two logical processors, but they share one core's arithmetic units, so together they deliver well under twice one core's throughput.

This matters when reading a cloud bill. On most x86 instance types, a cloud vCPU is one of those hardware threads, so two vCPUs are usually one physical core, not two. Some ARM-based instance types are the exception, where a vCPU is a whole core. Check which one you are buying before sizing CPU-bound work.

SIMD: One Instruction, Many Values

SIMD, single instruction multiple data, applies one operation to a whole register of values at once. A 256-bit vector register holds eight 32-bit floats, and one instruction adds them to eight others. Image codecs, compression, JSON parsers, memory copies and numerical libraries such as NumPy all use these instructions, because their inner loops do the same arithmetic over long runs of values.

One instruction, one lane or eight
Scalar add: one instruction, one resultabc+=Vector add: one instruction, eight resultsa0a1a2a3a4a5a6a7b0b1b2b3b4b5b6b7+c0c1c2c3c4c5c6c7=a 256-bit register holds eight 32-bit floats

Measured on the laptop this book was written on, summing a million 64-bit integers took about 24 milliseconds as a Python loop, about 3 milliseconds with the built-in sum function, and about a tenth of a millisecond with NumPy. That is roughly two hundred times faster than the loop. Most of the gain comes from skipping the interpreter on every element, and SIMD adds its share on top.

GPUs and Other Accelerators

A GPU takes the idea to its limit. It runs thousands of simple threads that perform the same operation on different data. That makes it excellent at graphics, at the matrix arithmetic behind machine learning, and at other large, uniform numerical work, and poor at branchy, sequential logic. GPUs also popularized reduced-precision number formats, 16-bit and smaller floats, that trade the exactness of Chapter 2 for more arithmetic per second. The machine learning courses in this catalogue take it from there.

Data must first travel from main memory to the GPU's own memory and back, so a small job can lose more in transfer than it gains in arithmetic. Chips also carry fixed-function units that do a single job, such as video encoding, encryption or neural-network inference, at a fraction of the power a general-purpose core would need for the same work.

Matching the Work to the Hardware

A sequential, branchy job, such as a parser or a compiler, wants one fast core. Many independent requests, such as a web server's traffic, want many cores. The same arithmetic over large arrays wants SIMD or a GPU. Large, uniform numerical work that repays its transfer wants a GPU.

Which hardware a kind of work can use
Sequential, branchy logic: a parser, a compiler→One fast core
Many independent requests: a web server→Many cores
The same arithmetic over long arrays→SIMD in a vectorized library
Huge, uniform numerical work that repays the transfer→A GPU

The cost in the systems you run is paying for the wrong kind. A 32-vCPU instance running a single-threaded Python process uses one hardware thread and bills for thirty-two. A GPU instance running a job too small to repay its data transfer is slower than the CPU it replaced, and considerably more expensive.

Core vs Hardware Thread vs vCPU

A core is a complete processor with its own pipeline and arithmetic units.

A hardware thread is one of the instruction streams, usually two, that a core can run at once, sharing that core's units. Two threads on one core give well under twice one core's throughput.

A vCPU is what a cloud provider sells, and on most x86 instance types it is one hardware thread. Count physical cores when sizing CPU-bound work, and vCPUs when reading the bill.

Misconceptions
  • "More cores make my program faster." Only the parts written to run in parallel get faster. A single-threaded Python process uses one core on a 64-core machine, and CPU-bound threads in standard CPython take turns on one interpreter lock unless the program runs on the free-threaded build, which Chapter 11 covers.
  • "A vCPU is a core." On most cloud instance types it is one hardware thread, half of a physical core. A CPU-bound job sized as "8 cores" on 8 vCPUs gets about four cores' worth of arithmetic units.
  • "A GPU is a faster CPU." It is a throughput machine for thousands of identical operations. On branchy, sequential or small jobs it is slower than a CPU once the data transfer is counted.
  • "Next year's chips will make this fast enough." Single-core speed now improves by modest percentages per year, not the doubling of the 1990s. A program that needs ten times the speed needs parallelism or a better algorithm, not a hardware refresh.
  • "SIMD is for compiler engineers." Every vectorized library you already use depends on it. Choosing an array operation over a Python loop is choosing SIMD, and it is often the largest single speed-up available without changing the algorithm.
Why It Matters
  • Size CPU-bound work by physical cores and waiting work by concurrency. vCPUs overstate compute, and work that is waiting for the network needs no core at all.
  • Express bulk arithmetic as whole-array operations in a vectorized library instead of a Python loop. It buys back the interpreter's overhead and adds SIMD on top.
  • Move work to a GPU only when it is large, uniform and arithmetic-heavy enough to repay the data transfer. Measure the transfer before the computation.
  • Parallelize the part of a job that takes the time, after measuring which part that is. The serial remainder caps the speed-up whatever the core count.
RelatedConcurrency and parallelism how software splits work, and Amdahl's law (Chapter 11)Free-threaded Python CPython without the interpreter lock (Chapter 11)Pipelines parallelism inside one core, invisible to the program

Knowledge Check

Why did clock speeds stop rising in the mid-2000s while transistor counts kept growing?

  • Transistors could no longer be made any smaller after 2005
  • Faster clocks made chips too hot once power stopped scaling
  • Programs could not use clocks faster than about 3 gigahertz
  • Memory could not deliver data faster than 3 gigahertz

A CPU-bound batch job is sized as needing 8 cores and placed on an x86 instance with 8 vCPUs. What does it get?

  • Eight physical cores, one for each vCPU
  • Sixteen cores, since each vCPU has two threads
  • One core, since vCPUs time-share a single core
  • About four cores' worth of arithmetic units

Which job is the best fit for a GPU?

  • Parsing one large, deeply nested configuration file
  • Multiplying two matrices of ten thousand rows each
  • Adding up the twelve numbers in one web form request
  • Serving thousands of requests that wait on a database

A team moves a single-threaded Python report job from 4 vCPUs to 32 vCPUs. What happens to its run time?

  • It drops about eight times, in step with the vCPUs
  • It halves, since Python spreads loops across cores
  • It barely changes, while the bill grows eight times
  • It gets slower, since more vCPUs add scheduling work

Why is summing a million numbers with NumPy so much faster than a Python loop?

  • It skips the interpreter per element and uses vector instructions
  • It moves the data to the GPU, which adds the numbers in parallel
  • It samples a fraction of the numbers and estimates the total
  • It splits the sum across every core of the machine by default

You got correct