Skip to content

Level 6 · Chapter 6.9

Real cores: x86, ARM and AVR compared

Three classic example processors (Intel's Sandy Bridge Core i7, the Cortex-A9 in TI's OMAP4430 and the 8-bit ATmega168), what changed since 2011 in Intel, AMD and Apple cores, and what an Apple M2 Ultra core actually does, measured: instructions per cycle, load latency at every cache level, and its efficiency cores.

The previous chapters took the ideas one at a time: pipelining, caches, branch prediction, out-of-order execution. A real core uses all of them at once. This chapter looks at real designs: three classic examples from around 2011, what has changed since, and a current core measured directly.

These three examples cover the whole range. The Core i7 is a high-end x86 desktop chip. The OMAP4430 is a phone chip built around two ARM Cortex-A9 cores. The ATmega168 is an 8-bit microcontroller that costs about a dollar. All three were current around 2011. Chips have changed a lot since, so the numbers below say which generation they describe.

The Core i7: an x86 on a RISC-like core

The Core i7 of 2011 is built on Intel's Sandy Bridge microarchitecture. From the outside, it runs the old x86 instruction set: variable-length instructions, few registers, complex addressing. On the inside, it looks like any fast RISC core. The trick is the front end:

  • The decoders translate x86 instructions into simple micro-ops (µops), up to four per cycle.
  • A µop cache of about 1,500 entries keeps decoded µops, so a hot loop skips decoding. Intel said the point was as much power as speed: with the µop cache hitting, the decoders can sleep.
  • The branch predictor steers fetch; its details are a trade secret.
  • The out-of-order engine renames registers and enters each µop in a 168-entry reorder buffer (ROB). Schedulers send ready µops to six ports: three with integer ALUs (also carrying the floating-point and branch units), one for stores and two for loads.
  • A retirement unit commits results in program order, which keeps exceptions precise.

Caches: 32 KB L1 instruction and data caches (8-way, 64-byte lines), a private 256 KB L2 per core, and a shared L3 of up to 20 MB.

A few facts about this chip are worth stating precisely:

  • Eight visible general-purpose registers is the 32-bit view. In 64-bit mode, which is how it's normally run, x86-64 has 16.
  • x86 instructions are 1 to 15 bytes long. Fifteen is the architectural limit; a longer instruction raises an exception.
  • AVX, introduced with Sandy Bridge, has 256-bit registers, twice the width of the 128-bit SSE registers before it; that's the whole point of AVX.
  • The clock rate barely rose after this generation. Sandy Bridge desktop chips turbo-boosted to about 3.9 GHz; the fastest desktop chips today reach about 6 GHz, and most cores run well below that. Performance since has come from wider cores, bigger caches and more cores, the subject of the rest of this level.

The Cortex-A9: RISC without translation

The Cortex-A9 is a 32-bit ARMv7 core, designed by ARM and licensed to chipmakers like Texas Instruments, who built the OMAP4430 around two of them. Since ARM instructions are already simple, register-to-register operations, there is no translation layer: instructions go straight from decode to rename and issue, and execute out of order in integer ALUs, a multiplier, an optional floating-point and NEON SIMD unit, and a load/store unit.

It has a deep pipeline with dynamic branch prediction, 32 KB L1 caches with 32-byte lines, and a 1 MB shared L2. It also has a small fast-loop buffer: a tight loop runs from it while the instruction cache and predictor power down. Saving energy drives phone designs as much as speed.

The ATmega168: the minimal end

The ATmega168 is a different world. It has 32 eight-bit registers, 16 KB of flash for the program, 1 KB of SRAM for data, 512 bytes of EEPROM, timers and I/O ports, all on one chip, and runs at up to 20 MHz. It has no cache, no branch predictor and no out-of-order logic. Its pipeline has two stages: while one instruction executes, the next is fetched. Most instructions take one cycle; loads, taken branches and multiplies take two.

An 8-bit ALU means that 32-bit arithmetic is done in pieces. Here is a + b on 32-bit integers, compiled with LLVM's AVR back end for the ATmega168:

add32:
    add r22, r18
    adc r23, r19
    adc r24, r20
    adc r25, r21
    ret

Four instructions, one per byte, with the carry passed along by adc: the ripple-carry adder of the digital-logic level, done in software. An x86 or ARM64 core does it in one instruction.

This simplicity is a feature. Every instruction's timing is known exactly, so a programmer can count cycles to generate a precise signal on a pin. A desktop core, with caches and speculation, can't promise that.

Two details here are easy to get wrong:

  • The AVR family can reach up to 16 MB of program memory through page registers called RAMPX, RAMPY and RAMPZ. Those exist only on bigger family members: RAMPZ on AVRs with more than 64 KB of flash, the others on the XMEGA line. The ATmega168, with 16 KB of flash, has none of them.
  • The ATmega168 is pipelined, small as it is: the two-stage fetch/execute overlap described above is why most instructions complete in one cycle.

The same shape everywhere

Side by side, the Core i7 and the Cortex-A9 show something that still holds: under very different instruction sets, they have the same execution core. Both execute simple operations with two source registers and one destination, one per cycle per unit, in a deep pipeline with branch prediction and split instruction and data caches. The difference is how they get there. The x86 front end has to break complex instructions into such operations; ARM's instructions already are.

That's exactly the progression from the Mic-1 to the Mic-4: an instruction set that isn't RISC-like, decoded into micro-operations placed in a queue. The Core i7 adds out-of-order execution on top. The ATmega168, in-order and tiny, is closest to the Mic-1.

Fifteen years later: wider and deeper

Since 2011, the organization has stayed the same, and everything in it has grown. Here are the headline numbers vendors have published for their big cores:

Sandy Bridge (2011)Golden Cove (2021)Lion Cove (2024)Zen 4 (2022)Zen 5 (2024)
VendorIntelIntelIntelAMDAMD
x86 instructions decoded per cycle46842 × 4
Reorder buffer entries168512576320448
Integer ALUs35646

Golden Cove is the performance core of Intel's 12th-generation Core chips, and Lion Cove that of the Core Ultra 200 series. Zen 4 and Zen 5 are in AMD's Ryzen 7000 and 9000 and the matching EPYC server chips. Zen 5's decoder is split into two clusters of four.

Apple publishes almost none of these figures for its own cores. Independent analyses of the M1's performance cores in 2020 found an 8-wide decoder and an instruction window of more than 600 entries, bigger than any x86 core of the time. ARM's current Cortex-X cores for phones are similarly wide. The Cortex-A9's descendants haven't shrunk: phone cores now match laptop cores.

Why wider rather than faster? The clock is limited by power: dynamic power grows with frequency and with the square of the voltage, and a higher frequency needs a higher voltage. A wider core gets more done per cycle instead. But the out-of-order chapter showed the limit: width only helps when the program has independent work.

Measuring a real core

The machine this chapter was written on is an Apple M2 Ultra. macOS reports its layout through sysctl:

$ sysctl hw.nperflevels hw.perflevel0 hw.perflevel1 hw.cachelinesize
hw.nperflevels: 2
hw.perflevel0.physicalcpu: 16
hw.perflevel0.l1icachesize: 196608
hw.perflevel0.l1dcachesize: 131072
hw.perflevel0.l2cachesize: 16777216
hw.perflevel0.cpusperl2: 4
hw.perflevel0.name: Performance
hw.perflevel1.physicalcpu: 8
hw.perflevel1.l1icachesize: 131072
hw.perflevel1.l1dcachesize: 65536
hw.perflevel1.l2cachesize: 4194304
hw.perflevel1.cpusperl2: 4
hw.perflevel1.name: Efficiency
hw.cachelinesize: 128

(Some lines trimmed.) There are two kinds of core. The 16 performance cores each have a 192 KB instruction cache and a 128 KB data cache (four times the Sandy Bridge's L1 data cache) and share a 16 MB L2 per cluster of four. The 8 efficiency cores are smaller: 128 KB and 64 KB L1s, and a 4 MB L2 per cluster of four. Cache lines are 128 bytes, twice the x86 size.

How wide is it?

macOS doesn't report the clock frequency, but a dependent chain of adds measures it. Each add x1, x1, #1 needs the previous result, and an add takes one cycle, so the chain runs at exactly one add per cycle. Timing 96 such adds in a loop, 20 million times:

// 96 dependent adds per iteration (the benchmark generates the full list)
__asm__ volatile(
    "1:\n"
    "add x1, x1, #1\n"
    "add x1, x1, #1\n"
    /* ... */
    "subs %0, %0, #1\n"
    "b.ne 1b\n"
    : "+r"(n) : : "x1", "cc");

gave 0.305 to 0.31 ns per add over several runs, so the P-core was running at about 3.3 GHz. Then the same number of adds, spread over k independent registers:

Independent chainsAdds per cycle
11.00
21.97
43.29
84.8
165.9
(nops instead of adds)7.8

With enough independent work, one core finishes almost six adds per cycle, which suggests six integer ALUs, and nearly eight instructions per cycle when they need no ALU at all, which points to an 8-wide front end. The results were the same over repeated runs. This is exactly the out-of-order effect in the simulator:

Out of order · Four independent chains on a 4-wide core

Try it: Switch between in order and out of order, change the width, window, load latency (4 = L1 hit, 12 = L2, 40 = L3) and renaming below, and compare the cycle counts, or press Edit to run your own code.

renaming

Loading emulator…

Twelve adds take 5 cycles at width 4, 8 at width 2, and 14 at width 1. Press Edit and make all twelve add eax, 1: now it takes 14 cycles at every width. Width is worthless without independent instructions.

The efficiency cores, reached by running the same program with taskpolicy -b (background priority), clocked at about 1.6 to 1.9 GHz in that mode and topped out around 4 adds and 5 nops per cycle. The numbers were noisier, since the clock kept changing, but the design is clearly narrower. The chapter on multicore comes back to why a chip has two kinds of core.

How far is memory?

The second measurement is load latency: a chain of pointers, each one stored in a different 128-byte line, visited in a random order so that neither the prefetcher nor any address or value predictor can guess the next one. Each load needs the previous one's result, so the time per step is the latency of wherever the data lives:

Working setTime per load≈ Cycles at 3.3 GHzWhere the data lives
16–128 KiB0.92 ns3L1 data cache
256 KiB – 4 MiB5.5–6.6 ns18–22L2
8–48 MiB7–84 ns, varies by run-L2, system cache, TLB misses
128 MiB – 1 GiB128–135 ns~430DRAM

The steps match the sysctl sizes: latency jumps just after 128 KiB, the L1's size. An L1 hit takes three cycles; DRAM takes more than 400. In the middle, several effects mix. The 16 MB L2 is shared by four cores, the chip has a large system-level cache in front of DRAM whose size Apple doesn't publish, and with 16 KiB pages a random walk over many megabytes also misses in the TLB. Those rows changed from run to run by a factor of two, so the table gives the range.

The caches chapter gave 4–5 cycles for a typical L1 and 60–100 ns for DRAM. The M2's L1 is faster than most, and its DRAM, as seen from one core chasing pointers, slower: the price of a large, shared memory system built for bandwidth, which the chapter on multiprocessors measures.

An 8-bit core next to a 64-bit one

Put the two ends side by side. The ATmega168 at 20 MHz completes at most 20 million simple instructions a second, 8 bits at a time. One M2 performance core, at 3.3 GHz and up to six adds per cycle, does about 20 billion 64-bit adds a second: a thousand times more operations, each on eight times as many bits, and the chip has 24 cores. Yet the microcontroller still sells in the billions, because for a thermostat, a key fob or a motor controller, a fixed, predictable cycle count and a few milliwatts matter more than speed.

Takeaways

  • The Core i7 (Sandy Bridge) decodes x86 into µops, caches them, and runs them out of order on a RISC-like core with a 168-entry ROB. The Cortex-A9 does the same without translation; the ATmega168 is an in-order, two-stage 8-bit core with no cache.
  • x86-64 has 16 general registers and instructions of at most 15 bytes, AVX registers are 256 bits wide, clock rates have barely risen since 2011, and the ATmega168 has a two-stage pipeline but no RAMP registers.
  • Since 2011, cores got wider (4 → 6–8 instructions decoded per cycle) and deeper (168 → 320–576 ROB entries), not much faster in clock.
  • Measured on an M2 Ultra performance core: about 3.3 GHz, up to ~6 adds and ~8 instructions per cycle with independent work, 1 add per cycle without it, and load latencies of 3 cycles (L1), ~20 cycles (L2) and ~430 cycles (DRAM).
  • The same chip mixes big and small cores, and the smallest cores in the world are still 8-bit microcontrollers.

In this level

  1. 6.1The fetch–decode–execute cycle
  2. 6.2Datapath, internal and system buses
  3. 6.3Control units and microcode
  4. 6.4A complete machine: the Mic-1 running IJVM
  5. 6.5Pipelining and hazards
  6. 6.6Caches and the memory hierarchy
  7. 6.7Branch prediction
  8. 6.8Out-of-order execution, register renaming and speculation
  9. 6.9Real cores: x86, ARM and AVR compared
  10. 6.10SIMD, GPUs and coprocessors
  11. 6.11Multicore, multithreading and cache coherence
  12. 6.12Shared-memory multiprocessors and NUMA
  13. 6.13Clusters, message passing and supercomputers