Until the mid-2000s, a new processor meant a faster processor: a higher clock, a deeper pipeline, more out-of-order machinery. Then the clock stopped climbing, and chip makers started spending their transistors on something else: more cores. The Apple M2 Ultra this chapter was measured on has 24 of them.
Several cores on one chip raise three new questions. How do cores that each have a private cache agree on what memory contains? What happens to performance when two cores touch the same cache line? And how much faster does a program really get with 24 cores? This chapter answers them, with measurements.
Why cores multiplied
The physics chapter on limits tells the story from below. For thirty years, Dennard scaling let each generation of smaller transistors run faster at the same power per square millimetre. Around 2005 the supply voltage stopped falling, and with it the free lunch: pushing the clock higher now meant more heat per square millimetre than a chip could shed. That's the power wall.
Moore's law kept delivering transistors, though. A single core couldn't use them well: the extra width and depth of an out-of-order core bring less and less, because ordinary code doesn't have enough independent instructions to feed it. The remaining option was to build several complete cores and let software run several threads at once. IBM's POWER4 put two cores on one die in 2001. Intel and AMD shipped their first dual-core x86 chips in 2005. Today's phones have six to ten cores, and server chips more than a hundred.
The catch: a second core does nothing for a program with one thread. Speed now has to come from the software, split into threads or processes.
Hardware multithreading
Before duplicating whole cores, there is a cheaper way to run several threads: let one core hold the state of several. A core spends much of its time waiting: for a cache miss, a mispredicted branch, a long division. If it has a second thread ready, it can fill those gaps.
Hardware multithreading gives each thread its own copy of the architectural state (registers, program counter, flags), while the threads share the execution units and caches. Three flavours:
| Style | When the core switches thread | Example |
|---|---|---|
| Fine-grained | every cycle, in turn | Sun's UltraSPARC T1 (2005): 8 cores × 4 threads |
| Coarse-grained | only on a long stall, like a cache miss | some older mainframe and server chips |
| Simultaneous (SMT) | never: instructions from several threads issue in the same cycle | Intel Hyper-Threading, AMD Zen, IBM POWER |
Fine-grained switching hides latency well but needs many threads, and a single thread runs slowly. Coarse-grained switching keeps one thread fast but pays a pipeline refill at each switch. Simultaneous multithreading fits an out-of-order core naturally: its scheduler already picks ready instructions from a large window, and it doesn't matter to it which thread they came from. Intel introduced it as Hyper-Threading in 2002, with two threads per core. IBM's POWER8 went up to eight.
SMT is cheap in silicon, but the two threads compete for everything: execution ports, L1 cache, branch predictor, TLB. Two threads on one core typically do more work together than one alone, but much less than two cores. They can also spy on each other through the shared structures, which is why some security-sensitive systems disable SMT. Apple's cores don't have it, and Intel dropped Hyper-Threading from the performance cores of its 2024 laptop and desktop chips. On the M2 Ultra, logical and physical CPU counts are equal:
$ sysctl hw.physicalcpu hw.logicalcpu
hw.physicalcpu: 24
hw.logicalcpu: 24
GPUs take fine-grained multithreading to the extreme, switching among dozens of groups of threads to hide memory latency; the SIMD and GPUs chapter covers them.
Performance and efficiency cores
The 24 cores of the M2 Ultra aren't all the same:
$ sysctl hw.nperflevels hw.perflevel0.name hw.perflevel0.logicalcpu \
hw.perflevel1.name hw.perflevel1.logicalcpu
hw.nperflevels: 2
hw.perflevel0.name: Performance
hw.perflevel0.logicalcpu: 16
hw.perflevel1.name: Efficiency
hw.perflevel1.logicalcpu: 8
The performance cores are wide out-of-order designs, the ones measured in the real cores chapter. The efficiency cores are smaller and slower, but do the same work with much less energy. Both run the same ARM64 instruction set, so the operating system can move any thread to either kind. ARM introduced the idea as big.LITTLE in 2011, Apple used it in every M-series Mac from 2020, and Intel's P-cores and E-cores followed in 2021.
The cost of the difference is easy to measure. A loop of 4.8 billion dependent multiply-adds took 4.6 s on one performance core, and 9.1–9.8 s on one efficiency core (forced there with taskpolicy -b, which also lowers its clock). Split across threads, with each thread doing an equal share:
| Threads | Time | Speedup |
|---|---|---|
| 1 | 4.59–4.63 s | 1 |
| 2 | 2.32–2.34 s | 2.0 |
| 4 | 1.17 s | 3.9 |
| 8 | 0.59 s | 7.8 |
| 16 | 0.31–0.32 s | about 14.5 |
| 24 | 0.24–0.27 s | 17–19 |
Up to 16 threads, each gets a performance core and the speedup follows the thread count. The last 8 threads land on efficiency cores, which add roughly four performance cores' worth: 24 cores, but a speedup of about 19 at best. Having threads take small chunks of work from a shared counter, instead of fixed equal shares, gave the best result, 0.24 s, because fast cores simply take more chunks. The operating system also helps by moving threads between core types as they run.
Private caches, shared memory
Each core has its own L1 caches, and usually its own L2 or an L2 shared by a small cluster. On the M2 Ultra, each group of four performance cores shares a 16 MB L2, and each group of four efficiency cores a 4 MB one:
hw.perflevel0.cpusperl2: 4
hw.perflevel0.l2cachesize: 16777216
hw.perflevel1.cpusperl2: 4
hw.perflevel1.l2cachesize: 4194304
Private caches are fast because they're close and nobody else uses them. But they create a problem the caches chapter could ignore. Suppose core 0 and core 1 have both read variable x into their L1. Core 0 writes x = 1 into its write-back cache. Core 1 reads x again, finds its own copy, and gets the old value. Two copies of the same address now disagree.
The hardware prevents this with a cache-coherence protocol. Its promise is simple: for each individual memory location, all cores see the writes in the same order, and a read returns the latest write. The usual way to guarantee it is the single-writer, multiple-reader rule: at any moment, a line is either writable by exactly one cache, or readable by any number of caches, never both.
MESI
The classic protocol, MESI, tags every line in every cache with one of four states:
| State | Meaning | Other caches may hold it? | Matches memory? |
|---|---|---|---|
| Modified | this cache has written it | no | no, this cache must write it back |
| Exclusive | only this cache has it, unmodified | no | yes |
| Shared | read-only copy | yes | yes |
| Invalid | not here (or stale) | - | - |
The caches watch each other's requests, originally by snooping on a shared bus, where every cache sees every transaction. The main transitions:
- Read miss. The cache asks for the line. If no one else has it, it arrives E; if other caches have it, everyone ends up S. If one cache holds it M, that cache supplies the data (and it's written back), and both end up S.
- Write to an S line. The cache first broadcasts an invalidate: every other copy goes I. Only then does its own copy become M. This is the single-writer rule in action.
- Write to an E line. No one else has it, so the cache silently switches to M, with no bus traffic. That's the point of the E state: a thread that reads and then writes private data doesn't pay for a broadcast.
- Write miss. The cache asks for the line with intent to modify: it gets the data and invalidates all other copies in one transaction.
The demo below runs exactly these rules on two cores. Read x from core 0 and the line arrives E; read it from core 1 and both copies become S; increment it on core 1 and core 0's copy is invalidated; read it again on core 0 and core 1 supplies the modified line and writes it back.
Try it: Click the buttons under each core to read or increment x and y, and watch each line's state (M, E, S, I) and the bus. Then press Play to run two threads that each increment their own counter, with x and y in one line or two.
Nothing has happened yet: every line is Invalid in both caches.
Real chips add states. AMD uses MOESI, whose Owned state lets a modified line be shared without writing it back to memory first. Intel's MESIF adds a Forward state that names the one sharer that answers requests, so several caches don't reply at once. A chip with dozens of cores also can't broadcast every request to everyone: it keeps a snoop filter or a directory that records which cores hold each line, and asks only those. The multiprocessors chapter returns to directories, where they become essential.
Coherence is also what makes atomic instructions work on a multicore chip: a core that executes lock add or ldadd takes the line in state M, does the read and the write in its own cache, and holds off other cores' requests until it's done. The bus lock of older machines is no longer needed, as the bus arbitration chapter explains.
A line moving between cores
Coherence isn't free. When one core writes a line that another core holds, the line has to move. Two threads playing ping-pong through one shared variable (each waits for the other's value, then writes its own) measure it directly. On the M2 Ultra, over several hundred runs of 20,000 round trips, a round trip took 66–70 ns at best, about 110 ns in the median run, and up to about 400 ns in the slowest.
macOS doesn't let a program choose which cores its threads run on, so each run lands on a different pair. The spread is consistent with the chip's layout: two cores in the same cluster share an L2, two cores in different clusters don't, and on the M2 Ultra some pairs are on different dies. Either way, a trip through the coherence protocol costs tens to hundreds of nanoseconds, against about 1 ns for an L1 hit.
False sharing
Coherence works on whole lines, not on variables. Two threads that write different variables in the same line still fight over it: each write invalidates the other core's copy, and the line bounces back and forth exactly as if the data were shared. This is false sharing.
The coherence demo above shows it at the scale of a single line. Press Play with x and y in the same line: two threads each increment their own counter 8 times, and every one of the 16 increments misses: 16 bus transactions and 15 invalidations, the line ping-ponging between the two caches. Switch to separate lines and play again: each thread misses once, then hits 7 times: 2 bus transactions, no invalidations.
The test: each thread increments its own counter 100 million times, with a volatile load, add and store. The counters sit either next to each other (8 bytes apart, in one 128-byte line) or 128 bytes apart (one line each):
static void *work(void *arg) {
volatile long *c = arg; // each thread gets its own counter
for (long i = 0; i < 100000000; i++)
(*c)++;
return NULL;
}
Times on the M2 Ultra, three runs each, with every thread doing the same 100 million increments:
| Threads | Counters packed (8 bytes apart) | Counters padded (128 bytes apart) |
|---|---|---|
| 1 | 0.030–0.031 s | 0.031–0.032 s |
| 2 | 0.147–0.150 s | 0.031 s |
| 4 | 0.162–0.167 s | 0.031–0.033 s |
| 8 | 0.183–0.184 s | 0.032–0.037 s |
| 16 | 0.214–0.226 s | 0.040–0.049 s |
With padding, 8 threads do 8 times the work in the same time as one, and 16 threads barely more: near-perfect scaling, because nothing is shared. Packed, two threads are already 5 times slower than one, though neither ever reads the other's counter. With atomic increments (atomic_fetch_add, relaxed) instead of plain ones, two threads took 12.9–14.1 ns per increment when packed and 2.1 ns when padded, a factor of 6 to 7.
There is a subtlety on this machine. sysctl hw.cachelinesize reports 128 bytes, twice the x86 size. Counters 64 bytes apart, safe on an x86 chip, still share a line here: with atomics, two threads took 3.9–4.2 ns per increment at a 64-byte distance, twice the padded time. Code that pads to 64 bytes because "cache lines are 64 bytes" is still falsely shared on Apple's chips. C++17's std::hardware_destructive_interference_size, and Linux's ____cacheline_aligned in the kernel, exist to get this right per platform.
The fix is layout: give each thread's hot data its own line, with alignas(128) or padding, or keep per-thread data in per-thread structures and combine it at the end. False sharing is invisible in the source code and hard to find without hardware performance counters, which is why it's a classic cause of programs that get slower as threads are added.
Amdahl's law
Suppose a program spends a fraction p of its time in work that can be split among n cores, and the rest, 1 − p, in work that can't: reading input, a sequential phase, a lock everybody queues on. The best possible speedup is:
speedup = 1 / ((1 − p) + p / n)
Gene Amdahl made this argument in 1967, and it's sobering. As n grows, p / n vanishes, and the speedup can never exceed 1 / (1 − p). A program that is 95% parallel can never run more than 20 times faster, whatever the number of cores. Change p below and run it:
Try it: Press Next line to run the highlighted line, or Play to watch; the boxes on the right are the variables, and the ones that just changed light up. Edit the code to try your own changes.
1// Amdahl's law: speedup = 1 / ((1 - p) + p / n)2// p is the parallel fraction, in thousandths (950 = 95%).3long speedup_x100(long p, long n) {4 return 100 * 1000 * n / ((1000 - p) * n + p);5}67int main() {8 long p = 950; // try 500, 900, 990...9 long cores[6] = {1, 2, 8, 24, 100, 1000};10 printf("parallel part: %ld.%ld%%\n", p / 10, p % 10);11 for (int i = 0; i < 6; i++) {12 long s = speedup_x100(p, cores[i]);13 printf("%4ld cores: %ld.%02ld x\n", cores[i], s / 100, s % 100);14 }15 return 0;16}
At 95%, 24 cores give 11.16×, and 1,000 cores only 19.62×. At 99%, 24 cores give 19.51× and 1,000 cores 90.99×. (The demo uses integer arithmetic, so it truncates rather than rounds.) The serial fraction dominates quickly: the last 5% decide what 1,000 cores are worth.
The measured scaling above fits the picture. The spinning loop is almost 100% parallel, so the M2 Ultra's 16 performance cores gave about 14.5×, and what held the 24-core result to 19× was the slower efficiency cores, not a serial part. A real program also has overheads that grow with n (contended locks, lines bouncing between caches, like the shared counters of the threads chapter) and can get slower past some number of cores.
Amdahl's law assumes the problem stays the same size. In practice, bigger machines are used for bigger problems, and the parallel part usually grows with the problem while the serial part doesn't. John Gustafson made that point in 1988: scaled this way, the achievable speedup keeps growing with n. Both views are right; they answer different questions: "how much faster is my job?" versus "how much bigger a job can I do in the same time?".
Takeaways
- When Dennard scaling ended around 2005, clock speeds stalled at the power wall, and extra transistors went into more cores. Performance now has to come from parallel software.
- Hardware multithreading keeps several threads' state in one core. SMT (Hyper-Threading) issues from several threads in the same cycle; Apple's cores have none: 24 physical = 24 logical CPUs on the M2 Ultra.
- The M2 Ultra has 16 performance and 8 efficiency cores. A compute loop sped up about 14.5× on 16 threads and 17–19× on 24, since an efficiency core ran it about half as fast.
- A coherence protocol keeps private caches consistent: one writer or many readers per line. MESI tracks each line as Modified, Exclusive, Shared or Invalid; big chips track sharers with snoop filters or directories.
- Moving a line between two M2 Ultra cores took 66–400 ns per round trip, depending on where the threads ran.
- False sharing: independent counters in one 128-byte line made two threads 5× slower than one (6–7× with atomics); padding to 128 bytes gave near-perfect scaling. On Apple chips, 64-byte padding isn't enough.
- Amdahl's law: with a parallel fraction p, speedup is capped at 1 / (1 − p): 20× for a 95%-parallel program, whatever the core count.