The previous chapter showed that 20 mnemonics make up about 93% of the instructions in ordinary programs. This chapter goes through those instructions, grouped by what they do. Each group has a live demo: step through it and watch the registers change.
Moving data
| Instruction | Does |
|---|---|
mov dst, src | copy |
movzx dst, src | copy a smaller value, filling the top with zeros (unsigned) |
movsx dst, src / movsxd | copy a smaller value, filling the top with copies of its sign bit (signed) |
lea dst, [address] | compute an address, without reading memory |
xchg a, b | swap two values |
push src / pop dst | put a value on the stack / take one off (rsp moves by 8) |
- mov ebx, 0xff80
- movzx eax, bl ; bl = 0x80 → 128
- movsx ecx, bl ; bl = 0x80 → -128
- lea edx, [rbx+rbx*4] ; 5 × rbx, no memory access
- xchg eax, ecx
- push rax
- pop rdi
The same byte, 0x80, becomes 128 with movzx and -128 (0xffffff80) with movsx: it depends on whether the source is treated as unsigned or signed. In C, that's the difference between unsigned char and signed char. lea computes rbx + rbx*4 = 5 × 0xff80 = 0x4fd80 without touching memory or the flags, which is why compilers use it for small multiplications. push then pop moves rax into rdi through the stack, and rsp goes down by 8 and back up.
Arithmetic
| Instruction | Does |
|---|---|
add / sub | add / subtract, setting the flags |
inc / dec | add / subtract 1 (leaves CF unchanged) |
neg | change the sign: 0 − x |
imul dst, src (and a 3-operand form) | signed multiply, keeping the low half |
mul src / imul src | full multiply: rdx:rax = rax × src |
div src / idiv src | divide rdx:rax by src: quotient in rax, remainder in rdx |
cdq / cqo | sign-extend eax into edx (rax into rdx) before a signed divide |
- mov eax, 100
- mov ecx, 7
- add eax, ecx ; 107
- sub eax, 9 ; 98
- imul eax, ecx ; 686
- add eax, 3 ; 689
- cdq ; edx = sign of eax (0)
- idiv ecx ; 689 / 7 → eax = 98, edx = 3
- neg eax ; -98
Division is the odd one out. It takes no destination operand: it always divides the 64-bit value edx:eax (or the 128-bit rdx:rax), and it writes two results, the quotient in eax and the remainder in edx. That's why a signed division in compiled code is almost always preceded by cdq or cqo. Dividing by zero doesn't produce a value: it raises a processor exception, which the OS turns into a crash (SIGFPE on Linux).
Division is also slow: often 10 to 40 cycles, depending on the CPU and the operand size, against 1 for an addition. Compilers avoid it when they can: dividing by a constant usually compiles to a multiplication by a "magic number" followed by a shift.
Bits and shifts
| Instruction | Does |
|---|---|
and / or / xor / not | bitwise AND, OR, exclusive OR, NOT |
shl | shift left: × 2ⁿ |
shr | shift right, filling with zeros: unsigned ÷ 2ⁿ |
sar | shift right, copying the sign bit: signed ÷ 2ⁿ, rounding down |
rol / ror | rotate: bits pushed out one end come back in the other |
bt, popcnt, bsf… | test a bit, count the 1s, find the first 1 |
- mov eax, 0xb6 ; 1011 0110
- and eax, 0x0f ; keep the low 4 bits: 0110 = 6
- or eax, 0x80 ; set bit 7: 0x86
- xor eax, 0xff ; flip the low 8 bits: 0x79
- shl eax, 4 ; × 16: 0x790
- shr eax, 8 ; ÷ 256: 7
- mov ecx, -16
- sar ecx, 2 ; -16 / 4 = -4 (sign kept)
- mov edx, -16
- shr edx, 2 ; 0x3ffffffc (sign lost)
andwith a mask keeps selected bits,orsets them andxorflips them.xor eax, eaxis the usual way to set a register to 0, andtest eax, eaxthe usual way to check whether it's 0.sarandshrdiffer only on negative numbers.sarkeeps -16 negative (-4);shrtreats it as a huge unsigned number.saralso rounds toward minus infinity:-17 sar 2is -5, while C's-17 / 4is -4. That's why a signed division by a power of two compiles to a few extra instructions around thesar.
Comparing and choosing
These instructions compare, and then act on the result through the flags. The flags chapter covers every condition.
| Instruction | Does |
|---|---|
cmp a, b | compute a − b, set the flags, discard the result |
test a, b | compute a AND b, set the flags, discard the result |
jcc label | jump if the condition cc holds: je, jne, jl, jg, jb, ja… |
setcc r8 | set a byte to 1 or 0 according to the condition |
cmovcc dst, src | copy only if the condition holds, with no jump |
- mov eax, 5
- cmp eax, 3 ; 5 − 3: greater
- setg bl ; bl = 1
- mov ecx, 100
- cmovg eax, ecx ; greater → eax = 100
- test eax, eax ; is eax zero?
- jz done ; no: fall through
- mov edx, 1
- done:
- nop
cmp followed by a conditional jump is how every if, loop and switch is compiled. setcc and cmovcc compute the same kind of decision without jumping, which avoids branch mispredictions.
Calls and jumps
| Instruction | Does |
|---|---|
jmp target | jump unconditionally; the target can also be in a register (jmp rax) |
call target | push the return address, then jump |
ret | pop the return address into rip |
leave | undo a stack frame: mov rsp, rbp + pop rbp |
- _start:
- mov edi, 6
- call triple ; pushes the return address
- jmp end
- mov eax, 0 ; never runs
- triple:
- lea eax, [rdi+rdi*2]
- ret ; back to the jmp
- end:
- nop
call pushes the address of the jmp onto the stack before jumping to triple. ret pops it back, so execution continues right after the call, with eax = 18. How arguments and return values travel is the subject of the calling conventions chapter.
The rest of the instruction set
A few more families show up less often, but you'll meet them:
- String instructions —
movsb,stosb,cmpsb,scasbwith arepprefix — copy, fill, compare or search whole blocks of memory in one instruction.rep stosbis used formemset. The microcode chapter shows how one such instruction becomes a loop. - System instructions:
syscallasks the operating system for a service,int3is the one-byte breakpoint debuggers insert (0xcc),hltstops the CPU until the next interrupt, andnopdoes nothing — useful as padding. - Floating point and SIMD: the SSE and AVX instructions work on the
xmm,ymmandzmmregisters. They handlefloatanddouble(addsd,mulss…) and operate on several values at once (paddd,vaddps…). The old x87 floating-point stack survives mostly in 32-bit code.
Takeaways
- Moving data:
mov,movzx/movsx(unsigned vs signed widening),lea(address arithmetic without memory),push/pop. - Arithmetic:
add,sub,inc,dec,neg,imul. Division usesrdx:raximplicitly and returns a quotient and a remainder. - Bits:
and/or/xorwith masks.shl/shr/sarmultiply and divide by powers of two;sarkeeps the sign. - Deciding:
cmp/testset the flags, thenjcc,setccorcmovccact on them. - Control:
jmp,call,ret— plus string, system and SIMD instructions for special jobs.