A CPU doesn't run programs the way we think of them. It runs instructions: tiny, fixed operations like "copy this value", "add these two numbers", "jump to that address". A program is a long list of them in memory, and the CPU executes them one after another, billions of times per second.
Assembly language is the human-readable form of those instructions: one line of text per instruction. It's what a disassembler shows you when you open a binary, and what a compiler produces from C just before it becomes machine code.
Anatomy of an instruction
loop: add eax, ecx ; eax = eax + ecx
| Part | Here | Role |
|---|---|---|
| label (optional) | loop: | a name for this instruction's address, so jumps can target it |
| mnemonic | add | the operation |
| operands | eax, ecx | what it operates on |
| comment | ; eax = eax + ecx | ignored by the assembler |
In the Intel syntax used on this site, the first operand is the destination. add eax, ecx means eax ← eax + ecx. Most x86 instructions take two operands, and the destination is also one of the inputs: the old value of eax is replaced by the result.
Three kinds of operand
Every operand is one of three things:
- a register, a named storage cell inside the CPU:
eax,rbx,r8… The registers chapter covers them. - an immediate, a constant written into the instruction itself:
5,0xff,-1. - a memory operand, written in brackets: the value at an address.
[rbp-4]means "the memory at addressrbpminus 4". The different ways to compute that address — base, index, scale, displacement — are the addressing modes of the ISA level.
mov eax, 5 ; immediate → register
mov ecx, eax ; register → register
mov DWORD PTR [rbp-4], ecx ; register → memory
mov edx, DWORD PTR [rbp-4] ; memory → register
DWORD PTR gives the size of a memory access: BYTE (8 bits), WORD (16), DWORD (32) or QWORD (64). With a register operand, the register's name already implies the size — eax is 32 bits — so the keyword is optional. With only a memory operand and an immediate, it's required: mov [rbp-4], 5 doesn't say whether to write 1, 2, 4 or 8 bytes, and the assembler rejects it as ambiguous operand size.
What an instruction changes
Every instruction reads some state and writes some state. The possible effects are few:
- registers — the destination register gets a new value;
- flags — arithmetic and logic instructions record facts about their result, such as zero, negative or overflow (the flags chapter);
- memory — a store writes bytes at an address;
rip, the instruction pointer — always, because it moves on to the next instruction, or to a jump's target.
Step through this program and watch each kind of change:
- mov ecx, 7 ; a register
- mov eax, 5
- add eax, ecx ; a register and the flags
- mov DWORD PTR [rbp-4], eax ; memory
mov only copies: it never touches the flags. add writes eax = 12 and updates the flags. The last mov writes 12 to the stack, at [rbp-4]. And rip advances at every step.
The rules
The CPU only implements certain combinations of operands, and the assembler refuses the others. These are the ones you'll run into:
| Not allowed | Why | Instead |
|---|---|---|
mov DWORD PTR [rbp-4], DWORD PTR [rbp-8] | at most one explicit memory operand per instruction | go through a register |
mov 5, eax | an immediate can't be a destination | mov eax, 5 |
mov eax, bx | both operands must be the same size | movzx eax, bx (zero-extend) |
mov [rbp-4], 5 | size unknown | mov DWORD PTR [rbp-4], 5 |
add rax, 0x100000000 | immediates are at most 32 bits (sign-extended), except in mov | load it into a register first |
All five are errors from a real assembler: invalid operand for instruction or ambiguous operand size. (String instructions like movsb are the exception: they copy memory to memory, with both addresses implicit in rsi and rdi.) The one-memory-operand rule is why compiled code is full of loads into registers: the CPU computes on registers, and memory is mostly read into them and written back from them.
Two ways to write the same thing: Intel and AT&T
The same machine instruction can be written in two assembly syntaxes. Intel syntax is used by Intel's manuals, NASM, MASM, IDA, Ghidra and this site. AT&T syntax is the traditional one of the GNU tools: it's what objdump and gdb show by default on Linux.
| Intel | AT&T | |
|---|---|---|
| operand order | destination, source | source, destination |
| registers | eax | %eax |
| immediates | 5 | $5 |
| size | from the operands, or DWORD PTR | suffix on the mnemonic: b, w, l, q |
| memory | [rbx+rcx*8+16] | 16(%rbx,%rcx,8) |
The same three instructions, disassembled both ways:
Intel AT&T
add eax, ecx addl %ecx, %eax
mov eax, 0x5 movl $0x5, %eax
mov rax, qword ptr [rbx + 8*rcx + 0x10] movq 0x10(%rbx,%rcx,8), %rax
To get Intel syntax from the GNU tools, use objdump -d -M intel or set disassembly-flavor intel in gdb. Once you know the order is reversed, reading one or the other is just a habit.
Behind the text: bytes
The CPU never sees the text. The assembler turns each line into machine code, a few bytes that encode the operation and its operands:
01 c8 add eax, ecx
b8 05 00 00 00 mov eax, 0x5
48 8b 44 cb 10 mov rax, qword ptr [rbx + 8*rcx + 0x10]
c7 45 fc 0a 00 00 00 mov dword ptr [rbp - 0x4], 0xa
c3 ret
48 b8 88 77 66 55 44 33 22 11 movabs rax, 0x1122334455667788
x86 instructions range from 1 to 15 bytes. ret is a single byte, and a mov with a 64-bit constant takes ten. The constants appear little-endian inside the bytes: 05 00 00 00 is 5. How the bits are laid out — prefixes, opcode, ModRM, displacement, immediate — belongs to the instruction set architecture, and the ISA level's chapter on instruction formats takes it apart. At this level, one line of assembly is one instruction, and that's enough.
How many instructions are there?
x86-64 has roughly a thousand mnemonics once you count every SIMD and system instruction, and many more variants of each. But real code uses a small core. Counting the instructions in the x86-64 builds of four ordinary programs — ssh, vim, zip and curl, about 764,000 instructions in total:
| Rank | Mnemonic | Share |
|---|---|---|
| 1 | mov | 31.7% |
| 2 | call | 7.8% |
| 3 | lea | 7.7% |
| 4 | cmp | 6.6% |
| 5 | je | 5.4% |
| 6 | test | 5.3% |
| 7 | pop | 4.4% |
| 8 | push | 4.4% |
| 9 | xor | 4.1% |
| 10 | jmp | 4.0% |
These four programs use 284 different mnemonics, but the top 10 cover 81% of their instructions and the top 20 cover 93%. Nearly a third of everything is mov: most of a program's work is moving data between memory and registers. Learning a few dozen instructions is enough to read most disassembly. The next chapter is a tour of them.
Takeaways
- An instruction is one basic operation. Assembly writes it as a mnemonic followed by operands, destination first in Intel syntax.
- Operands are registers, immediates or memory (
[...]), with a size: byte, word, dword, qword. - An instruction can change registers, flags, memory and
rip— nothing else. - The rules: at most one explicit memory operand, matching sizes, no immediate destination, and immediates limited to 32 bits outside
mov. - AT&T syntax reverses the operands and adds
%,$and size suffixes. - Each line assembles to 1–15 bytes of machine code, and a few dozen mnemonics make up almost all real code.