Every operand has to say where its value is: inside the instruction, in a register, or somewhere in memory. The different ways of saying it are the ISA's addressing modes. They're a design choice. A rich set makes each instruction more powerful but harder to decode; a minimal set keeps the hardware simple but needs more instructions. The assembly level introduced the three kinds of operands. This chapter looks at the modes themselves, across ISAs.
The modes
| Mode | The operand is… | x86 example | Typical use |
|---|---|---|---|
| immediate | in the instruction | mov eax, 5 | constants |
| register | in a register | mov ebx, eax | local variables the compiler kept in registers |
| direct (absolute) | at a fixed address given in the instruction | mov eax, DWORD PTR [0x404000] | globals, in position-dependent code |
| register indirect | at the address held in a register | mov ecx, DWORD PTR [rsi] | pointers: *p |
| base + displacement | at a register plus a constant | mov edx, DWORD PTR [rsi+8] | locals ([rbp-4]), struct fields (p->y) |
| base + index × scale (+ displacement) | at base + index × 1, 2, 4 or 8 (+ constant) | mov r8d, DWORD PTR [rsi+rdi*4] | arrays: a[i], arrays of structs |
| PC-relative | at an offset from the instruction pointer | mov r10d, DWORD PTR [rip+arr+4] | globals in position-independent code, branch targets |
| stack (implicit) | at the top of the stack | push rax, pop r11 | calls, saved registers |
Tanenbaum's names differ slightly: indexed for a register plus a constant, and based-indexed for two registers plus an optional constant. They're the same ideas.
Watching the address being computed
This program reads the same small array in several ways. Each memory operand's address is computed by the AGU (address generation unit) before the load. The simulator opens at the µarch level, where each step shows that calculation:
- .data
- arr: .long 10, 20, 30, 40
- .text
- mov eax, 5 ; immediate
- mov ebx, eax ; register
- lea rsi, [rip+arr] ; PC-relative: rsi = &arr
- mov ecx, DWORD PTR [rsi] ; register indirect: arr[0] = 10
- mov edx, DWORD PTR [rsi+8] ; base + displacement: arr[2] = 30
- mov edi, 3
- mov r8d, DWORD PTR [rsi+rdi*4] ; base + index×4: arr[3] = 40
- mov r9d, DWORD PTR [rsi+rdi*4-4] ; base + index×4 + disp: arr[2] = 30
- mov r10d, DWORD PTR [rip+arr+4] ; PC-relative: arr[1] = 20
- push rax ; stack: write at rsp - 8
- pop r11 ; stack: read at rsp
For mov r8d, DWORD PTR [rsi+rdi*4], the AGU shows rsi (0x404000) + rdi × 4 = 0x40400c, then the bus read returns 40. The scale factor exists because array elements are 1, 2, 4 or 8 bytes wide: the index stays an element number, and the hardware multiplies it by the element size for free.
From C to addressing modes
Compilers map C expressions onto these modes so directly that you can often read the C back out of the assembly:
| C | Typical x86-64 operand |
|---|---|
local x | a register, or [rbp-4] / [rsp+12] |
*p | [rax] |
p->y (field at offset 8) | [rax+8] |
a[i] (int) | [rax+rcx*4] |
a[i].y (16-byte structs, y at offset 8) | shl rcx, 4 then [rax+rcx+8] — the scale only goes up to 8 |
global g | [rip+g] |
&a[i] | lea rdx, [rax+rcx*4] — the same address, without the load |
That last line is why lea exists, and why compilers also use it for plain arithmetic: it runs the AGU's add-and-scale without touching memory.
What the modes cost in bytes
Richer modes mean longer instructions. The same load, with different addresses:
objdump shows8b 56 08 mov edx, dword ptr [rsi + 0x8]
8B → MOV r32, r/m32 (dst = src): ModRM.reg is the destination, ModRM.rm the source.
mod=01 memory + disp8 · reg=edx · rm=base rsi.
01mod010reg110rm- mod=01
- memory + disp8
- reg=010
- edx
- rm=110
- base rsi
64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.
objdump shows8b 86 00 10 00 00 mov eax, dword ptr [rsi + 0x1000]
8B → MOV r32, r/m32 (dst = src): ModRM.reg is the destination, ModRM.rm the source.
mod=10 memory + disp32 · reg=eax · rm=base rsi.
10mod000reg110rm- mod=10
- memory + disp32
- reg=000
- eax
- rm=110
- base rsi
64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.
[rsi+8] fits its displacement in one byte (8b 56 08). [rsi+0x1000] needs four (8b 86 00 10 00 00). A scaled index adds a SIB byte: mov r8d, DWORD PTR [rsi+rdi*4] is 44 8b 04 be. And in 64-bit mode, an absolute address like [0x404000] can only be encoded through a SIB byte with no base (8b 04 25 00 40 40 00), because the short encoding that meant "absolute address" in 32-bit mode was reassigned to rip-relative addressing. That's one reason 64-bit code uses [rip+…] for its globals.
Addressing branch targets
Branches need addresses too, and they use their own modes:
- PC-relative:
jmp,jccandcallnormally encode a signed offset from the next instruction, 8 bits for nearby targets and 32 bits otherwise. The code keeps working wherever it's loaded, which is what makes position-independent code and ASLR possible. - Register or memory indirect:
jmp rax,call QWORD PTR [rax+16]. The target is computed at run time. That's howswitchjump tables, function pointers and C++ virtual calls work, and why they need an indirect branch predictor. - Stack:
rettakes its target from the top of the stack.
Other ISAs make other choices
- RISC-V is a pure load/store architecture: arithmetic works only on registers, and loads and stores have exactly one memory mode, base register + 12-bit offset. Reading
a[i]takes three instructions: shift i by 2, add the base, thenlwwith offset 0. The decoder stays trivial, and the compiler does the rest. - ARM64 is also load/store, but its loads accept a scaled register offset (
ldr w0, [x1, x2, lsl #2]readsa[i]in one instruction). It also has pre- and post-indexed modes that update the base register as a side effect:ldr x0, [x1], #8loads, then advancesx1by 8, which is ideal for walking an array. - x86 allows a memory operand on most arithmetic instructions (
add eax, [rsi+rdi*4]), with the richest address form of the three, but at most one memory operand per instruction.
How freely modes combine with instructions is called orthogonality. In a fully orthogonal ISA, every instruction accepts every mode for every operand. The old PDP-11 and VAX came close. x86 is far from it — one memory operand, special registers for some instructions. Load/store RISCs sidestep the question by allowing memory only in loads and stores.
A mode that no longer exists: self-modifying code
Before register-indirect addressing existed, programs walked through arrays by rewriting the address inside their own instructions on each iteration. Von Neumann himself suggested it. It's gone from ordinary code: it breaks code sharing between processes, it conflicts with split instruction and data caches, and it's impossible on memory marked non-writable.
But code that writes code still exists. JIT compilers generate machine code at run time, and packed or obfuscated executables decrypt their real code before jumping to it. x86 detects writes to instructions already in the pipeline and flushes them automatically. ARM requires the program to explicitly synchronize its instruction cache with the new code. And both have to work within the OS's W^X rules, making memory writable, then executable, but never both at once.
Takeaways
- An addressing mode says where an operand is: in the instruction (immediate), in a register, at a fixed address (direct), at a register's address (register indirect), at a register plus a constant (base + displacement), at base + index × scale, relative to the PC, or on the stack.
- These modes map directly onto C: locals,
*p,p->field,a[i]and globals each have a characteristic form, andleacomputes the address without loading. - Richer modes cost bytes: 1- vs 4-byte displacements, the SIB byte, and
rip-relative addressing in 64-bit mode. - Branches use PC-relative offsets, indirect targets (jump tables, virtual calls) and the stack (
ret). - RISC-V offers one memory mode, ARM64 adds scaled and auto-indexed forms, and x86 allows memory operands in arithmetic. Orthogonality measures how freely modes combine with instructions.