Skip to content

Level 5 · Chapter 5.1

What the ISA promises: memory model, alignment and ordering

The instruction set architecture as a contract between software and hardware: what it specifies and what it leaves open, user and kernel mode, registers across x86, ARM and RISC-V, the address space, what misaligned accesses really cost today, and memory ordering — with store-buffering and message-passing litmus tests run on a real ARM64 chip and under x86 ordering.

The levels above this one — programming languages, assembly, the operating system — all produce or run machine code. The levels below — microarchitecture, logic, devices — build a machine that executes it. The instruction set architecture (ISA) is the line between them: a contract that says exactly what each instruction does, and nothing about how the hardware achieves it. Everything above the line relies only on the contract; everything below is free to change as long as it keeps it. That's why a program compiled for x86-64 in 2005 still runs on a 2026 processor built from a completely different microarchitecture.

This chapter covers what that contract contains: modes, registers, the memory model — and the two places where the hardware underneath leaks through, alignment and ordering.

A contract, written down

Tanenbaum defines the ISA as what a compiler writer needs to know: the memory model, the registers, the data types and the instructions. Whether the machine is pipelined, superscalar or microprogrammed isn't part of it, although it affects performance, which compilers do care about.

For most ISAs the contract is a formal document, written to let several companies build compatible chips. The ARM Architecture Reference Manual defines ARM. The RISC-V specifications, published by RISC-V International, are open for anyone to implement without a license. x86 is described by Intel's Software Developer's Manual, which the book notes ran to 4,161 pages for the first Core i7, and has only grown since. These documents are written like legal texts: an instruction shall raise an exception in some case, while other behavior is implementation defined, left to each chip designer. Code that relies on implementation-defined behavior works on one chip and breaks on another.

The contract also defines privilege levels. Programs run in user mode, where instructions that control the machine — page tables, interrupts, I/O ports, cache control — are forbidden and cause a fault. The operating system runs in kernel mode, where everything is allowed. x86 calls these rings 3 and 0, ARM64 calls them EL0 and EL1 (with EL2 for hypervisors and EL3 for firmware), and RISC-V calls them U-mode and S-mode (with M-mode for firmware). Everything in this level is user mode unless stated otherwise.

Registers

Registers are the ISA's fastest storage: a handful of named locations, each a word wide, that instructions can use directly. They come in two kinds: special-purpose registers — the program counter, the stack pointer, the flags — and general-purpose registers for data.

General-purpose registersNotes
x86-6416 (rax…r15)some have fixed roles in older instructions (rdx for mul/div, rcx for shifts and rep); Intel's APX extension doubles them to 32
ARMv7 (32-bit ARM)16 (r0…r15)the program counter is r15, one of the sixteen, so writing to it is a jump
ARM6431 (x0…x30)plus a register number that reads as zero or means the stack pointer, depending on the instruction
RISC-V32 (x0…x31)x0 is hardwired to zero

The book says RISC machines usually have at least 32 general-purpose registers. That's true of MIPS, SPARC, RISC-V and ARM64, but the most widespread RISC of the book's time, 32-bit ARM, had only 16, program counter included.

The flags register (Tanenbaum's PSW) holds the condition codes — zero, negative, carry, overflow — that compare and branch instructions communicate through, as the flags chapter showed. Not every ISA has one: RISC-V deliberately has no flags at all. Its branches compare two registers directly (blt a0, a1, target), and an overflow check needs an explicit comparison. That removes a register that every instruction would otherwise update, and a dependency that makes out-of-order execution harder.

The memory model: bytes and addresses

At the ISA level, memory is an array of bytes, each with an address, from 0 up to a maximum. On 64-bit ISAs the addresses are 64-bit numbers, but no current chip implements all 2⁶⁴ bytes. x86-64 uses 48-bit virtual addresses (57 with five-level paging), and the unused top bits must be copies of bit 47 — canonical addresses. ARM64 uses 48 or 52 bits.

Most machines use one address space for both code and data — the von Neumann model. Some small microcontrollers, like the ATmega168 AVR that Tanenbaum uses as an example, keep program memory and data memory in separate address spaces — the Harvard model — so address 8 means something different for an instruction fetch and for a load. Modern desktop CPUs are von Neumann at the ISA level, even though their L1 caches are split into instruction and data halves internally.

Alignment: what it costs today

A value is aligned when its address is a multiple of its size: an 8-byte word at address 0, 8, 16… The variables chapter showed compilers inserting padding to guarantee it. The ISA decides what happens when it isn't.

x86 has always allowed misaligned accesses, going back to the 8088 and its 1-byte bus. The simulator, being an x86-64, does too:

Live · Aligned and misaligned loads

Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.

program— ▸ is the next instruction
  1. .data
  2. bytes: .quad 0x0807060504030201, 0x100f0e0d0c0b0a09
  3. .text
  4. lea rsi, [rip+bytes]
  5. mov rax, QWORD PTR [rsi] ; aligned: bytes 0-7
  6. mov rbx, QWORD PTR [rsi+1] ; misaligned: bytes 1-8
  7. mov ecx, DWORD PTR [rsi+6] ; misaligned: bytes 6-9
step 0
Loading emulator…
The instruction that just ran, as the bytes the CPU actually fetched and decoded.

rbx ends up as 0x0908070605040302: the 8 bytes starting at offset 1, assembled in little-endian order.

Tanenbaum argues that misaligned accesses are slow because the memory interface only transfers aligned 8-byte words, so the CPU must make two memory references and splice the result. That was true of the memory bus, but it's no longer what a load sees: loads are served from the L1 cache, which delivers any bytes within a cache line. What costs is crossing a cache line or a page boundary. Measured on an Apple M2 (ARM64, 128-byte cache lines, 16 KiB pages) with a chain of dependent 8-byte loads:

Where the 8-byte load landsTime per load
aligned1.19 ns
misaligned by 1, 60 or 120 bytes, inside one cache line1.19 ns — no difference
crossing a 128-byte cache line1.26 ns
crossing a 16 KiB pageabout 4× slower than aligned loads on separate pages

The page-crossing load needs two address translations and touches two pages, which is where the real cost is. Alignment still matters in a few places: some instructions require it (x86's movaps faults on a vector that isn't 16-byte aligned); RISC-V allows chips to trap on misaligned accesses and emulate them in software, hundreds of times slower; and atomic operations that span two cache lines are a problem everywhere — on x86, such a split lock stalls the whole memory system, and Linux can warn about or kill programs that do it.

Memory ordering: what other cores see

The book raises a harder question. A load after a store to the same address must see the stored value — every ISA guarantees that for a single core. But the core reorders its memory operations internally, as the out-of-order chapter showed, and keeps its stores in a store buffer before they reach the cache. With several cores sharing memory, the question becomes: in what order does another core see my loads and stores? The ISA's answer is its memory consistency model, and the book sketches the range from strict ordering to no guarantees at all. Today's ISAs have precise answers:

  • x86 — TSO (total store order): almost sequential. The one reordering allowed is that a load may complete before an earlier store to a different address has become visible: the store is still waiting in the store buffer.
  • ARM64 and RISC-V — weak ordering (RISC-V calls its model RVWMO): loads and stores to different addresses may become visible in almost any order, unless the program asks for more.

Small litmus tests make this visible. In the store-buffering test, two threads each write one variable and then read the other:

Thread 0          Thread 1
x = 1             y = 1
r1 = y            r2 = x

Interleave them however you like and at least one thread must see the other's write — unless stores can wait in a buffer while later loads go ahead. Run 2 million times on the M2 in native ARM64 mode, both threads read 0 in about 91 % of runs. Run as an x86-64 program under Rosetta 2, which switches the M2's cores into a hardware TSO mode to match x86 rules, it still happened in about 97 % of runs: TSO allows exactly this reordering. With a full fence between the store and the load, it happened 0 times in both modes.

The message-passing test is the one that breaks programs:

Writer            Reader
data = 42         f = flag
flag = 1          d = data      // can f be 1 while d is 0?

TSO forbids that outcome: stores become visible in order, and loads are performed in order. ARM64's and RISC-V's models allow it. In 2 million runs on the M2 it never happened, in either mode — but the architecture permits it, so correct code can't rely on any particular chip not doing it.

Programs restore order with fences and ordered accesses. In C, atomic_store_explicit(&flag, 1, memory_order_release) and the matching memory_order_acquire load make the message-passing pattern correct everywhere, and the compiler picks the instructions each ISA needs. From clang:

x86-64ARM64
release store of flagplain movstlr (store-release)
acquire load of flagplain movldar (load-acquire)
full fencelock or [rsp-64], 0dmb ish

On x86, TSO already provides acquire and release ordering, so the plain movs suffice. Only the full fence costs anything, and clang implements it with a locked instruction, which is cheaper than mfence. On ARM64 the ordering has to be requested at each access. That's the trade-off in the book's terms: the weaker the model, the more freedom the hardware has, and the more the compiler and programmer have to say explicitly. Code that works on x86 by accident — ordinary variables used as flags between threads — can break when ported to ARM. The C and C++ memory models were designed to hide these differences behind atomics.

Takeaways

  • The ISA is a contract: registers, memory model, data types and instructions, specified precisely (shall / implementation-defined), while the microarchitecture is free to change underneath.
  • User mode forbids the instructions that control the machine; kernel mode allows everything.
  • x86-64 has 16 general-purpose registers, ARM64 31, RISC-V 32 with x0 always zero. RISC-V has no flags register.
  • Memory is a byte-addressed array; 64-bit chips implement 48 to 57 address bits. Most CPUs are von Neumann; some microcontrollers are Harvard.
  • Misaligned loads are free inside a cache line on current CPUs; crossing a line costs a little, crossing a page a lot. Some instructions and atomics still require alignment.
  • The memory consistency model says what other cores see. x86's TSO only lets a load pass an earlier store (the store buffer): measured on an M2, 91–97 % of store-buffering runs showed it. ARM64 and RISC-V are weakly ordered. Fences, acquire and release restore order.

In this level

  1. 5.1What the ISA promises: memory model, alignment and ordering
  2. 5.2Instruction formats and encoding
  3. 5.3Addressing modes
  4. 5.4x86, ARM and RISC-V side by side
  5. 5.5Traps, interrupts and exceptions
  6. 5.6Data types the hardware understands
  7. 5.7VLIW, EPIC and the Itanium