Almost every instruction a CPU executes reads or writes a register: a small, named storage slot inside the processor itself. Memory holds gigabytes; the registers hold a few dozen values. But they are where the work happens — arithmetic, comparisons and addresses all flow through them.
Why registers exist
Reaching main memory is slow. A register is read in less than a CPU cycle; the closest cache takes a few cycles; main memory can take a hundred or more. Every value an instruction needs from memory has to travel over the bus first.
You can see the difference in the simulator. Open the CPU & buses tab in the demo below and watch the bus cycles counter as you step:
- mov ecx, 5
- mov eax, ecx ; register → register
- mov DWORD PTR [rbp-4], ecx ; register → memory (a store)
- mov edx, DWORD PTR [rbp-4] ; memory → register (a load)
mov eax, ecxcosts 1 read: just fetching the instruction itself. The value moves inside the CPU.- The store costs the fetch plus 1 write on the data bus.
- The load costs the fetch plus 1 read.
That is why compilers work so hard to keep hot variables in registers, and why unoptimized code — which stores every variable in memory — is so much slower.
The general-purpose registers
x86-64 has 16 general-purpose registers, each 64 bits wide. The first eight carry names from the 16-bit 8086 era, when each one had a favourite job:
| Register | Historical role | Typical use today |
|---|---|---|
rax | accumulator | return values, arithmetic |
rbx | base | general purpose (preserved across calls) |
rcx | counter | loop counters, 4th argument |
rdx | data | 3rd argument, high half of mul/div |
rsi | source index | 2nd argument, string source |
rdi | destination index | 1st argument, string destination |
rbp | base pointer | frame pointer (in unoptimized code) |
rsp | stack pointer | always the top of the stack |
r8–r15 | (added by AMD64) | 5th/6th arguments, general purpose |
Only rsp is truly special: it always points to the top of the stack, and push, pop, call and ret use it implicitly. The others are interchangeable as far as the hardware is concerned — their roles come from conventions, which you'll meet in calling conventions.
In 32-bit x86 there are only the first eight, 32 bits wide: eax, ebx, ecx, edx, esi, edi, ebp, esp.
One register, several names
Each 64-bit register can be accessed in smaller pieces. For rax:
| Name | Bits | Part of rax |
|---|---|---|
rax | 64 | all of it |
eax | 32 | bits 0–31 |
ax | 16 | bits 0–15 |
ah | 8 | bits 8–15 |
al | 8 | bits 0–7 |
The same pattern applies to rbx/ebx/bx/bh/bl, rcx, and rdx. rsi, rdi, rbp and rsp have 8-bit forms sil, dil, bpl, spl, and the new registers use suffixes: r8d (32), r8w (16), r8b (8).
The demo below fills rax with a recognizable pattern, then writes to smaller and smaller pieces of it. Step through and watch rax in the register panel:
- mov rax, 0x1122334455667788
- mov al, 0xff ; only bits 0–7 change
- mov ah, 0xee ; only bits 8–15 change
- mov ax, 0x1234 ; only bits 0–15 change
- mov eax, 0xdeadbeef ; …and the upper 32 bits are cleared!
| After | rax |
|---|---|
mov al, 0xff | 0x11223344556677ff |
mov ah, 0xee | 0x112233445566eeff |
mov ax, 0x1234 | 0x1122334455661234 |
mov eax, 0xdeadbeef | 0x00000000deadbeef |
The zero-extension rule
That last line surprises everyone once: writing a 32-bit register clears the upper 32 bits of the 64-bit register, while writing an 8- or 16-bit register leaves the rest untouched.
AMD designed x86-64 this way on purpose. If writing eax had to preserve the top half, the CPU would have to merge the old and new values — an extra dependency that slows everything down. Since most code works with 32-bit ints, making those writes "clean" is a big win.
It also explains two idioms you'll see constantly in compiled code:
mov eax, eaxlooks like it does nothing, but it zero-extendseaxintorax— converting anunsigned intto a 64-bit value.xor eax, eaxis the shortest way to set the wholeraxto 0.
Registers you can't mov into
rip, the instruction pointer, holds the address of the next instruction. You never write it directly:jmp,call,retand conditional jumps change it. It can be read through addressing:lea rax, [rip+message]computes the address ofmessagerelative to the current instruction, which is how position-independent code finds its data.rflagsholds the status flags written by arithmetic and read by conditional jumps — the subject of the next chapter.
There are more registers beyond these: the SIMD registers xmm0–xmm15 (and their wider ymm/zmm forms) for floating point and vector math, segment registers, and control registers used by the operating system. The general-purpose sixteen are the ones you'll read in almost every line of disassembly.
Takeaways
- Registers are the CPU's working storage; memory is reached over the bus and is far slower.
- x86-64 has 16 general-purpose 64-bit registers; only
rsphas a hard-wired role. rax,eax,ax,ahandalare views of the same bits.- Writing a 32-bit register zeroes the top half; writing 8 or 16 bits doesn't.
ripandrflagsare changed by the instructions themselves, not bymov.