Skip to content

Level 3 · Chapter 3.2

A tour of the instruction set

The few dozen x86-64 instructions that make up almost all real code, family by family — moving data, arithmetic, bits and shifts, comparing and choosing, calls and jumps — each one live in the simulator.

The previous chapter showed that 20 mnemonics make up about 93% of the instructions in ordinary programs. This chapter goes through those instructions, grouped by what they do. Each group has a live demo: step through it and watch the registers change.

Moving data

InstructionDoes
mov dst, srccopy
movzx dst, srccopy a smaller value, filling the top with zeros (unsigned)
movsx dst, src / movsxdcopy a smaller value, filling the top with copies of its sign bit (signed)
lea dst, [address]compute an address, without reading memory
xchg a, bswap two values
push src / pop dstput a value on the stack / take one off (rsp moves by 8)
Live · Moving data
program— ▸ is the next instruction
  1. mov ebx, 0xff80
  2. movzx eax, bl ; bl = 0x80 → 128
  3. movsx ecx, bl ; bl = 0x80 → -128
  4. lea edx, [rbx+rbx*4] ; 5 × rbx, no memory access
  5. xchg eax, ecx
  6. push rax
  7. pop rdi
step 0
Loading emulator…
The instructions the compiler generated, and the registers, flags and stack they change.

The same byte, 0x80, becomes 128 with movzx and -128 (0xffffff80) with movsx: it depends on whether the source is treated as unsigned or signed. In C, that's the difference between unsigned char and signed char. lea computes rbx + rbx*4 = 5 × 0xff80 = 0x4fd80 without touching memory or the flags, which is why compilers use it for small multiplications. push then pop moves rax into rdi through the stack, and rsp goes down by 8 and back up.

Arithmetic

InstructionDoes
add / subadd / subtract, setting the flags
inc / decadd / subtract 1 (leaves CF unchanged)
negchange the sign: 0 − x
imul dst, src (and a 3-operand form)signed multiply, keeping the low half
mul src / imul srcfull multiply: rdx:rax = rax × src
div src / idiv srcdivide rdx:rax by src: quotient in rax, remainder in rdx
cdq / cqosign-extend eax into edx (rax into rdx) before a signed divide
Live · Arithmetic
program— ▸ is the next instruction
  1. mov eax, 100
  2. mov ecx, 7
  3. add eax, ecx ; 107
  4. sub eax, 9 ; 98
  5. imul eax, ecx ; 686
  6. add eax, 3 ; 689
  7. cdq ; edx = sign of eax (0)
  8. idiv ecx ; 689 / 7 → eax = 98, edx = 3
  9. neg eax ; -98
step 0
Loading emulator…
The instructions the compiler generated, and the registers, flags and stack they change.

Division is the odd one out. It takes no destination operand: it always divides the 64-bit value edx:eax (or the 128-bit rdx:rax), and it writes two results, the quotient in eax and the remainder in edx. That's why a signed division in compiled code is almost always preceded by cdq or cqo. Dividing by zero doesn't produce a value: it raises a processor exception, which the OS turns into a crash (SIGFPE on Linux).

Division is also slow: often 10 to 40 cycles, depending on the CPU and the operand size, against 1 for an addition. Compilers avoid it when they can: dividing by a constant usually compiles to a multiplication by a "magic number" followed by a shift.

Bits and shifts

InstructionDoes
and / or / xor / notbitwise AND, OR, exclusive OR, NOT
shlshift left: × 2ⁿ
shrshift right, filling with zeros: unsigned ÷ 2ⁿ
sarshift right, copying the sign bit: signed ÷ 2ⁿ, rounding down
rol / rorrotate: bits pushed out one end come back in the other
bt, popcnt, bsf…test a bit, count the 1s, find the first 1
Live · Bits and shifts
program— ▸ is the next instruction
  1. mov eax, 0xb6 ; 1011 0110
  2. and eax, 0x0f ; keep the low 4 bits: 0110 = 6
  3. or eax, 0x80 ; set bit 7: 0x86
  4. xor eax, 0xff ; flip the low 8 bits: 0x79
  5. shl eax, 4 ; × 16: 0x790
  6. shr eax, 8 ; ÷ 256: 7
  7. mov ecx, -16
  8. sar ecx, 2 ; -16 / 4 = -4 (sign kept)
  9. mov edx, -16
  10. shr edx, 2 ; 0x3ffffffc (sign lost)
step 0
Loading emulator…
The instructions the compiler generated, and the registers, flags and stack they change.
  • and with a mask keeps selected bits, or sets them and xor flips them.
  • xor eax, eax is the usual way to set a register to 0, and test eax, eax the usual way to check whether it's 0.
  • sar and shr differ only on negative numbers. sar keeps -16 negative (-4); shr treats it as a huge unsigned number.
  • sar also rounds toward minus infinity: -17 sar 2 is -5, while C's -17 / 4 is -4. That's why a signed division by a power of two compiles to a few extra instructions around the sar.

Comparing and choosing

These instructions compare, and then act on the result through the flags. The flags chapter covers every condition.

InstructionDoes
cmp a, bcompute a − b, set the flags, discard the result
test a, bcompute a AND b, set the flags, discard the result
jcc labeljump if the condition cc holds: je, jne, jl, jg, jb, ja…
setcc r8set a byte to 1 or 0 according to the condition
cmovcc dst, srccopy only if the condition holds, with no jump
Live · Comparing and choosing
program— ▸ is the next instruction
  1. mov eax, 5
  2. cmp eax, 3 ; 5 − 3: greater
  3. setg bl ; bl = 1
  4. mov ecx, 100
  5. cmovg eax, ecx ; greater → eax = 100
  6. test eax, eax ; is eax zero?
  7. jz done ; no: fall through
  8. mov edx, 1
  9. done:
  10. nop
step 0
Loading emulator…
The instructions the compiler generated, and the registers, flags and stack they change.

cmp followed by a conditional jump is how every if, loop and switch is compiled. setcc and cmovcc compute the same kind of decision without jumping, which avoids branch mispredictions.

Calls and jumps

InstructionDoes
jmp targetjump unconditionally; the target can also be in a register (jmp rax)
call targetpush the return address, then jump
retpop the return address into rip
leaveundo a stack frame: mov rsp, rbp + pop rbp
Live · Calling a function
program— ▸ is the next instruction
  1. _start:
  2. mov edi, 6
  3. call triple ; pushes the return address
  4. jmp end
  5. mov eax, 0 ; never runs
  6. triple:
  7. lea eax, [rdi+rdi*2]
  8. ret ; back to the jmp
  9. end:
  10. nop
step 0
Loading emulator…
The instructions the compiler generated, and the registers, flags and stack they change.

call pushes the address of the jmp onto the stack before jumping to triple. ret pops it back, so execution continues right after the call, with eax = 18. How arguments and return values travel is the subject of the calling conventions chapter.

The rest of the instruction set

A few more families show up less often, but you'll meet them:

  • String instructions — movsb, stosb, cmpsb, scasb with a rep prefix — copy, fill, compare or search whole blocks of memory in one instruction. rep stosb is used for memset. The microcode chapter shows how one such instruction becomes a loop.
  • System instructions: syscall asks the operating system for a service, int3 is the one-byte breakpoint debuggers insert (0xcc), hlt stops the CPU until the next interrupt, and nop does nothing — useful as padding.
  • Floating point and SIMD: the SSE and AVX instructions work on the xmm, ymm and zmm registers. They handle float and double (addsd, mulss…) and operate on several values at once (paddd, vaddps…). The old x87 floating-point stack survives mostly in 32-bit code.

Takeaways

  • Moving data: mov, movzx/movsx (unsigned vs signed widening), lea (address arithmetic without memory), push/pop.
  • Arithmetic: add, sub, inc, dec, neg, imul. Division uses rdx:rax implicitly and returns a quotient and a remainder.
  • Bits: and/or/xor with masks. shl/shr/sar multiply and divide by powers of two; sar keeps the sign.
  • Deciding: cmp/test set the flags, then jcc, setcc or cmovcc act on them.
  • Control: jmp, call, ret — plus string, system and SIMD instructions for special jobs.

In this level

  1. 3.1What an instruction is
  2. 3.2A tour of the instruction set
  3. 3.3Registers and the register file
  4. 3.4Flags, cmp and conditional jumps
  5. 3.5Calling conventions (System V, cdecl)
  6. 3.6Assembler, linker and loader
  7. 3.7Directives, sections and macros