Skip to content

Level 5 · Chapter 5.2

Instruction formats and encoding

How an instruction becomes bits: opcodes and operand fields, fixed versus variable length, expanding opcodes, and a field-by-field tour of x86-64 encoding — prefixes, REX, opcode, ModRM, SIB, displacement, immediate — with ARM64 and RISC-V for contrast.

At the assembly level, an instruction is a line of text. The CPU fetches bytes. The ISA fixes how every instruction is written as bits: which bits select the operation, which ones name the registers, where constants and addresses go. That layout is the instruction format, and it's one of the hardest things to change once a processor ships. x86 still decodes instructions laid out for the 8086 in 1978.

What an instruction has to say

Every instruction needs an opcode, the field that says what to do. Most also need to say where their operands are. ISAs differ in how many operand addresses each instruction carries:

AddressesExampleStyle
0iadd (JVM)stack machine: operands are implicitly the top of a stack
1ADC #5 (6502)accumulator: one operand is always the accumulator register
2add eax, ecx (x86)the destination is also a source
3add w0, w1, w2 (ARM64)two sources and a separate destination

Fewer addresses make each instruction shorter but need more instructions: computing a = b + c takes one three-address instruction, but a load, an add and a store on a stack machine. Modern RISC designs are three-address register machines. x86 is two-address, with one operand allowed in memory.

Fixed or variable length

  • Fixed length: every instruction is the same size — 32 bits on ARM64 and base RISC-V. Decoding is simple and fast: the next instruction always starts 4 bytes later, so a CPU can decode many at once. But simple instructions waste bits, and constants must fit in the leftover fields.
  • Variable length: an instruction takes as many bytes as it needs — 1 to 15 on x86. Code is denser, which helps the instruction cache. But the decoder can't know where the next instruction starts until it has worked out the length of this one.

Many fixed-length ISAs added a compressed 16-bit form for density: Thumb on 32-bit ARM, and the C extension on RISC-V.

Designers weigh several pressures. Shorter instructions mean more instructions per cache line and per fetch cycle, and memory bandwidth is scarce. There has to be room in the opcode space for every operation, and for the ones that will be added over decades. And every address field has to be wide enough for the registers or memory it names.

Expanding opcodes

A fixed format can still give some instructions more opcode bits than others. Take a 16-bit instruction with a 4-bit opcode and three 4-bit register fields. Normally that allows 16 three-register instructions. But if one opcode value, 1111, means "the opcode continues into the next field", the space splits further:

Opcode bitsOperand fields leftInstructions
4 (0000–1110)3 × 4 bits15 three-address
8 (1111 0000–1111 1101)2 × 4 bits14 two-address
12 (1111 1110 …, 1111 1111 …)1 × 4 bits31 one-address
16 (1111 1111 1111 …)none16 zero-address

Every one of the 65,536 bit patterns is used: 15 × 4096 + 14 × 256 + 31 × 16 + 16 = 65,536. The same idea — short opcodes for instructions that need many operand bits, longer ones for the rest — shows up in nearly every real encoding, including x86's escape bytes.

x86-64: a field-by-field tour

An x86-64 instruction is a sequence of optional fields around a mandatory opcode:

FieldSizePurpose
prefixes0–4 byteschange the instruction: 66 operand size 16-bit, f0 lock, f3 rep…
REX0–1 byte0x40–0x4f: W = 64-bit operands, R, X, B = 4th bit of register numbers (for r8–r15)
opcode1–3 bytesthe operation; 0f is the escape byte that opens a second opcode table
ModRM0–1 bytemod (2 bits): register or memory, and displacement size; reg (3): a register or opcode extension; rm (3): the other operand
SIB0–1 bytescale (2), index (3), base (3), for addresses like [base + index×scale]
displacement0, 1 or 4 bytesthe constant in an address, like the 8 in [rbx+8]
immediate0, 1, 2 or 4 bytes (8 for movabs)a constant operand

The total is capped at 15 bytes. The opcode itself says little: the decoder has to recognize it fully before it even knows whether a ModRM byte follows. Tanenbaum's text gives 0xFF as the escape byte; it's actually 0x0F. 0xFF is an ordinary opcode, used for inc, dec, call, jmp and push on memory.

Type an instruction into the encoder and it shows every field, with the bits of ModRM and SIB drawn out:

Encoding · add with a memory operandInput

objdump shows03 44 8b 08 add eax, dword ptr [rbx + 4*rcx + 0x8]

4 bytes
length
max 15 on x86
none
prefixes
b+i×s+d8
addressing
base + index×scale + disp8
none
immediate
  • 03 → ADD r32, r/m32 (dst += src): ModRM.reg is the destination, ModRM.rm the source.

  • mod=01 memory + disp8 · reg=eax · rm=SIB.

    01
    mod
    000
    reg
    100
    rm
    mod=01
    memory + disp8
    reg=000
    eax
    rm=100
    100: a SIB byte follows
  • Scale-Index-Base: address = base + index × scale + disp.

    10
    scale
    001
    index
    011
    base
    scale=10
    ×4
    index=001
    rcx
    base=011
    rbx

64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.

03 44 8b 08: the opcode 03 means "add a register-or-memory operand to a 32-bit register". ModRM 44 is 01 000 100: mod 01 says memory with an 8-bit displacement, reg 000 is eax, and rm 100 means a SIB byte follows. SIB 8b is 10 001 011: scale ×4, index rcx, base rbx. Then comes the displacement, 08.

The encoding has its quirks, and each one shows up in disassembly:

  • mov eax, DWORD PTR [rsp+8] needs a SIB byte (8b 44 24 08), because rm = 100, the code for rsp, was taken to mean "SIB follows".
  • mov eax, DWORD PTR [rbp] needs a zero displacement (8b 45 00), because mod 00 with rm = 101 was taken for rip-relative addressing.
  • r8–r15 need a REX byte to supply the fourth register bit, and r12 also inherits rsp's SIB rule:
Encoding · A REX prefix and an inherited quirkInput

objdump shows45 8b 04 24 mov r8d, dword ptr [r12]

4 bytes
length
max 15 on x86
45
prefixes
REX
b
addressing
base
none
immediate
  • REX = 0100WRXB. R: extends ModRM.reg; B: extends ModRM.rm / SIB.base / opcode reg.

    0100
    fixed
    0
    W
    1
    R
    0
    X
    1
    B
    fixed=0100
    marks a REX prefix (0x40–0x4F)
    W=0
    default size
    R=1
    ModRM.reg + 8 → r8d
    X=0
    index not extended
    B=1
    SIB.base + 8 → base r12
  • 8B → MOV r32, r/m32 (dst = src): ModRM.reg is the destination, ModRM.rm the source.

  • mod=00 memory, no displacement · reg=r8d · rm=SIB.

    00
    mod
    000
    reg
    100
    rm
    mod=00
    memory, no displacement
    reg=000
    r8d
    rm=100
    100: a SIB byte follows
  • Scale-Index-Base: address = base + disp (needed only because r12 as a base is encoded rm=100, which means "SIB follows").

    00
    scale
    100
    index
    100
    base
    scale=00
    ×1
    index=100
    100: no index
    base=100
    r12 (with REX.B)

64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.

45 8b 04 24: REX 45 has R set (the destination is register 8 + 0 = r8d) and B set (the base is 8 + 4 = r12). The SIB byte is there only because r12's low three bits, 100, are the same as rsp's.

A few more patterns to recognize:

  • 48 83 c0 01 is add rax, 1. 48 is REX.W for 64-bit operands, and 83 is the form with a sign-extended 8-bit immediate, so small constants take one byte instead of four.
  • 0f af c1 is imul eax, ecx. 0f escapes to the two-byte opcode table, where most instructions added after the 8086 live.
  • 66 89 d8 is mov ax, bx. The 66 prefix switches a 32-bit instruction to 16 bits.
  • f0 83 07 01 is lock add DWORD PTR [rdi], 1: an atomic increment.

ARM64 and RISC-V: fixed fields

In a fixed 32-bit format, every field sits at the same place in every instruction of its type. ARM64's add w0, w1, w2 is 0x0b020020:

Bits3130–2928–2423–222120–1615–109–54–0
Fieldsfop, S01011shift0Rmimm6RnRd
Value0 (32-bit)00add (shifted register)0002010

RISC-V's add a0, a1, a2 is 0x00c58533, in the R-type format: funct7 | rs2 | rs1 | funct3 | rd | opcode = 0000000 | 01100 | 01011 | 000 | 01010 | 0110011. The register numbers a2 = 12, a1 = 11 and a0 = 10 can be read straight off the bits. A RISC-V decoder knows where every register field is before it even knows what the instruction does, so it can start reading registers early. That's the point of a regular format.

Why it matters when reading binaries

Because x86 instructions have variable length, a disassembler has to find where each one begins. Start decoding at the wrong byte and you get a different, perfectly valid-looking instruction stream:

Encoding · Five bytes, one instructionInput

decodes tomov eax, 0x909090c3

5 bytes
length
max 15 on x86
none
prefixes
none
addressing
no ModRM byte
imm32
immediate
  • B8+r → MOV r32, imm32 (dst = src): the register is in the low 3 bits of the opcode: no ModRM byte.

    10111
    opcode
    000
    reg
    opcode=10111
    B8+r
    reg=000
    eax

64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.

These five bytes are mov eax, 0x909090c3. Skip the first byte — decode c3 90 90 90 — and the same memory reads as ret, followed by nops. Instructions hidden inside other instructions are why disassemblers can be fooled by anti-disassembly tricks, and where return-oriented programming finds many of its "gadgets": useful byte sequences in the middle of a program's own code. On ARM64 or RISC-V, with fixed, aligned instructions, this can't happen: every instruction starts on a 4-byte boundary.

Takeaways

  • An instruction format fixes where the opcode and operand fields go. ISAs range from zero-address (stack) to three-address (register) designs.
  • Fixed-length formats (ARM64, RISC-V: 32 bits) are easy to decode in parallel. Variable-length ones (x86: 1–15 bytes) are denser but must be decoded one after another.
  • Expanding opcodes give short opcodes to instructions that need many operand bits, and long ones to the rest.
  • x86-64 encodes an instruction as prefixes, REX, opcode, ModRM, SIB, displacement, immediate. Its quirks — the SIB byte for rsp, the zero displacement for rbp, REX for r8–r15 — are visible in every disassembly.
  • Variable length lets instructions hide inside others, which matters for disassemblers and exploits.

In this level

  1. 5.1What the ISA promises: memory model, alignment and orderingPlanned
  2. 5.2Instruction formats and encoding
  3. 5.3Addressing modes
  4. 5.4x86, ARM and RISC-V side by sidePlanned
  5. 5.5Traps, interrupts and exceptionsPlanned
  6. 5.6Data types the hardware understandsPlanned
  7. 5.7VLIW, EPIC and the ItaniumPlanned