At the assembly level, an instruction is a line of text. The CPU fetches bytes. The ISA fixes how every instruction is written as bits: which bits select the operation, which ones name the registers, where constants and addresses go. That layout is the instruction format, and it's one of the hardest things to change once a processor ships. x86 still decodes instructions laid out for the 8086 in 1978.
What an instruction has to say
Every instruction needs an opcode, the field that says what to do. Most also need to say where their operands are. ISAs differ in how many operand addresses each instruction carries:
| Addresses | Example | Style |
|---|---|---|
| 0 | iadd (JVM) | stack machine: operands are implicitly the top of a stack |
| 1 | ADC #5 (6502) | accumulator: one operand is always the accumulator register |
| 2 | add eax, ecx (x86) | the destination is also a source |
| 3 | add w0, w1, w2 (ARM64) | two sources and a separate destination |
Fewer addresses make each instruction shorter but need more instructions: computing a = b + c takes one three-address instruction, but a load, an add and a store on a stack machine. Modern RISC designs are three-address register machines. x86 is two-address, with one operand allowed in memory.
Fixed or variable length
- Fixed length: every instruction is the same size — 32 bits on ARM64 and base RISC-V. Decoding is simple and fast: the next instruction always starts 4 bytes later, so a CPU can decode many at once. But simple instructions waste bits, and constants must fit in the leftover fields.
- Variable length: an instruction takes as many bytes as it needs — 1 to 15 on x86. Code is denser, which helps the instruction cache. But the decoder can't know where the next instruction starts until it has worked out the length of this one.
Many fixed-length ISAs added a compressed 16-bit form for density: Thumb on 32-bit ARM, and the C extension on RISC-V.
Designers weigh several pressures. Shorter instructions mean more instructions per cache line and per fetch cycle, and memory bandwidth is scarce. There has to be room in the opcode space for every operation, and for the ones that will be added over decades. And every address field has to be wide enough for the registers or memory it names.
Expanding opcodes
A fixed format can still give some instructions more opcode bits than others. Take a 16-bit instruction with a 4-bit opcode and three 4-bit register fields. Normally that allows 16 three-register instructions. But if one opcode value, 1111, means "the opcode continues into the next field", the space splits further:
| Opcode bits | Operand fields left | Instructions |
|---|---|---|
4 (0000–1110) | 3 × 4 bits | 15 three-address |
8 (1111 0000–1111 1101) | 2 × 4 bits | 14 two-address |
12 (1111 1110 …, 1111 1111 …) | 1 × 4 bits | 31 one-address |
16 (1111 1111 1111 …) | none | 16 zero-address |
Every one of the 65,536 bit patterns is used: 15 × 4096 + 14 × 256 + 31 × 16 + 16 = 65,536. The same idea — short opcodes for instructions that need many operand bits, longer ones for the rest — shows up in nearly every real encoding, including x86's escape bytes.
x86-64: a field-by-field tour
An x86-64 instruction is a sequence of optional fields around a mandatory opcode:
| Field | Size | Purpose |
|---|---|---|
| prefixes | 0–4 bytes | change the instruction: 66 operand size 16-bit, f0 lock, f3 rep… |
| REX | 0–1 byte | 0x40–0x4f: W = 64-bit operands, R, X, B = 4th bit of register numbers (for r8–r15) |
| opcode | 1–3 bytes | the operation; 0f is the escape byte that opens a second opcode table |
| ModRM | 0–1 byte | mod (2 bits): register or memory, and displacement size; reg (3): a register or opcode extension; rm (3): the other operand |
| SIB | 0–1 byte | scale (2), index (3), base (3), for addresses like [base + index×scale] |
| displacement | 0, 1 or 4 bytes | the constant in an address, like the 8 in [rbx+8] |
| immediate | 0, 1, 2 or 4 bytes (8 for movabs) | a constant operand |
The total is capped at 15 bytes. The opcode itself says little: the decoder has to recognize it fully before it even knows whether a ModRM byte follows. Tanenbaum's text gives 0xFF as the escape byte; it's actually 0x0F. 0xFF is an ordinary opcode, used for inc, dec, call, jmp and push on memory.
Type an instruction into the encoder and it shows every field, with the bits of ModRM and SIB drawn out:
objdump shows03 44 8b 08 add eax, dword ptr [rbx + 4*rcx + 0x8]
03 → ADD r32, r/m32 (dst += src): ModRM.reg is the destination, ModRM.rm the source.
mod=01 memory + disp8 · reg=eax · rm=SIB.
01mod000reg100rm- mod=01
- memory + disp8
- reg=000
- eax
- rm=100
- 100: a SIB byte follows
Scale-Index-Base: address = base + index × scale + disp.
10scale001index011base- scale=10
- ×4
- index=001
- rcx
- base=011
- rbx
64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.
03 44 8b 08: the opcode 03 means "add a register-or-memory operand to a 32-bit register". ModRM 44 is 01 000 100: mod 01 says memory with an 8-bit displacement, reg 000 is eax, and rm 100 means a SIB byte follows. SIB 8b is 10 001 011: scale ×4, index rcx, base rbx. Then comes the displacement, 08.
The encoding has its quirks, and each one shows up in disassembly:
mov eax, DWORD PTR [rsp+8]needs a SIB byte (8b 44 24 08), because rm =100, the code forrsp, was taken to mean "SIB follows".mov eax, DWORD PTR [rbp]needs a zero displacement (8b 45 00), because mod00with rm =101was taken forrip-relative addressing.r8–r15need a REX byte to supply the fourth register bit, andr12also inheritsrsp's SIB rule:
objdump shows45 8b 04 24 mov r8d, dword ptr [r12]
REX = 0100WRXB. R: extends ModRM.reg; B: extends ModRM.rm / SIB.base / opcode reg.
0100fixed0W1R0X1B- fixed=0100
- marks a REX prefix (0x40–0x4F)
- W=0
- default size
- R=1
- ModRM.reg + 8 → r8d
- X=0
- index not extended
- B=1
- SIB.base + 8 → base r12
8B → MOV r32, r/m32 (dst = src): ModRM.reg is the destination, ModRM.rm the source.
mod=00 memory, no displacement · reg=r8d · rm=SIB.
00mod000reg100rm- mod=00
- memory, no displacement
- reg=000
- r8d
- rm=100
- 100: a SIB byte follows
Scale-Index-Base: address = base + disp (needed only because r12 as a base is encoded rm=100, which means "SIB follows").
00scale100index100base- scale=00
- ×1
- index=100
- 100: no index
- base=100
- r12 (with REX.B)
64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.
45 8b 04 24: REX 45 has R set (the destination is register 8 + 0 = r8d) and B set (the base is 8 + 4 = r12). The SIB byte is there only because r12's low three bits, 100, are the same as rsp's.
A few more patterns to recognize:
48 83 c0 01isadd rax, 1.48is REX.W for 64-bit operands, and83is the form with a sign-extended 8-bit immediate, so small constants take one byte instead of four.0f af c1isimul eax, ecx.0fescapes to the two-byte opcode table, where most instructions added after the 8086 live.66 89 d8ismov ax, bx. The66prefix switches a 32-bit instruction to 16 bits.f0 83 07 01islock add DWORD PTR [rdi], 1: an atomic increment.
ARM64 and RISC-V: fixed fields
In a fixed 32-bit format, every field sits at the same place in every instruction of its type. ARM64's add w0, w1, w2 is 0x0b020020:
| Bits | 31 | 30–29 | 28–24 | 23–22 | 21 | 20–16 | 15–10 | 9–5 | 4–0 |
|---|---|---|---|---|---|---|---|---|---|
| Field | sf | op, S | 01011 | shift | 0 | Rm | imm6 | Rn | Rd |
| Value | 0 (32-bit) | 00 | add (shifted register) | 00 | 0 | 2 | 0 | 1 | 0 |
RISC-V's add a0, a1, a2 is 0x00c58533, in the R-type format: funct7 | rs2 | rs1 | funct3 | rd | opcode = 0000000 | 01100 | 01011 | 000 | 01010 | 0110011. The register numbers a2 = 12, a1 = 11 and a0 = 10 can be read straight off the bits. A RISC-V decoder knows where every register field is before it even knows what the instruction does, so it can start reading registers early. That's the point of a regular format.
Why it matters when reading binaries
Because x86 instructions have variable length, a disassembler has to find where each one begins. Start decoding at the wrong byte and you get a different, perfectly valid-looking instruction stream:
decodes tomov eax, 0x909090c3
B8+r → MOV r32, imm32 (dst = src): the register is in the low 3 bits of the opcode: no ModRM byte.
10111opcode000reg- opcode=10111
- B8+r
- reg=000
- eax
64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.
These five bytes are mov eax, 0x909090c3. Skip the first byte — decode c3 90 90 90 — and the same memory reads as ret, followed by nops. Instructions hidden inside other instructions are why disassemblers can be fooled by anti-disassembly tricks, and where return-oriented programming finds many of its "gadgets": useful byte sequences in the middle of a program's own code. On ARM64 or RISC-V, with fixed, aligned instructions, this can't happen: every instruction starts on a 4-byte boundary.
Takeaways
- An instruction format fixes where the opcode and operand fields go. ISAs range from zero-address (stack) to three-address (register) designs.
- Fixed-length formats (ARM64, RISC-V: 32 bits) are easy to decode in parallel. Variable-length ones (x86: 1–15 bytes) are denser but must be decoded one after another.
- Expanding opcodes give short opcodes to instructions that need many operand bits, and long ones to the rest.
- x86-64 encodes an instruction as prefixes, REX, opcode, ModRM, SIB, displacement, immediate. Its quirks — the SIB byte for
rsp, the zero displacement forrbp, REX forr8–r15— are visible in every disassembly. - Variable length lets instructions hide inside others, which matters for disassemblers and exploits.