Skip to content

Level 4 · Chapter 4.1

The fetch–decode–execute cycle

What a CPU does for every single instruction: fetch it, decode it, fetch its operands, execute it, write back — watched micro-operation by micro-operation.

A CPU runs a program by repeating one loop, billions of times per second:

  1. Fetch the next instruction from the address in rip (the program counter), and advance rip past it.
  2. Decode it: work out what operation it is and where its operands are.
  3. Fetch operands: from registers, from the instruction itself, or from memory — which first means computing the address.
  4. Execute it: the ALU computes the result.
  5. Write back the result to a register, or store it to memory.

rip moves on as soon as the instruction has been fetched, which is why, during an instruction, rip already holds the address of the next one — the value rip-relative addressing and call rely on. A jump simply overwrites it with a new target.

Every instruction — a mov, an add, a call — is some combination of these steps. At the assembly level you only see the instruction. One level down, in the microarchitecture, you can see the steps.

The demos in this chapter open on the simulator's CPU & buses tab. Each time you press Step, it replays the micro-operations of the instruction that just ran, one by one, on a diagram of the CPU, the system bus and memory. Use the ◀ ▶ arrows next to µop to go through them at your own pace. To keep the diagram readable, the simulator shows the rip update as each instruction's last µop, once the jump decision is known; the effect is the same.

The simplest instruction: mov eax, 5

Live · mov eax, 5
program— ▸ is the next instruction
  1. mov eax, 5
step 0
Loading emulator…

Five micro-operations:

#StepWhat happens
1Fetchrip is copied into the MAR (memory address register) and put on the address bus; the control unit asserts MEMR (memory read); memory answers on the data bus with the instruction, which lands in the IR (instruction register).
2DecodeThe control unit decodes the IR: operation mov, destination eax, source an immediate.
3OperandThe immediate 5 is part of the instruction itself — it comes straight from the IR, no memory access needed.
4Write-back5 is written into eax.
5Nextrip advances to the next instruction.

One bus cycle in total: the fetch. Everything else happens inside the CPU.

Control-signal names vary between buses: MEMR/MEMW are the ISA-bus names used here, while many textbook buses use a memory-request line (MREQ) plus separate RD and WR lines. Same idea: one line says "this is a memory access", another says which direction.

A full cycle: add eax, DWORD PTR [rbp-4]

Now an instruction that uses every stage. The demo first stores 10 at [rbp-4] and loads 5 into eax, then runs the add:

Live · add with a memory operand
program— ▸ is the next instruction
  1. mov DWORD PTR [rbp-4], 10
  2. mov eax, 5
  3. add eax, DWORD PTR [rbp-4]
step 0
Loading emulator…

The add takes eight micro-operations:

  1. Fetch the instruction into the IR.
  2. Decode: an add, register destination, memory source.
  3. Read eax (5) from the register file onto the internal bus.
  4. Address generation: the AGU computes rbp − 4 = 0x7fffffdc. Before the CPU can read memory, it has to know where.
  5. Memory read: 0x7fffffdc goes into the MAR and onto the address bus, MEMR is asserted, and memory drives 10 onto the data bus into the MDR (memory data register). The simulator highlights the stack region, where that address lives.
  6. Execute: the ALU computes 5 + 10 = 15 and sets the flags.
  7. Write-back: 15 goes into eax.
  8. Next: rip advances.

Two bus cycles: the instruction fetch and the data read. The bus cycles counter under the diagram shows 2 R · 0 W.

Writing memory: push rbx

Live · push rbx
program— ▸ is the next instruction
  1. mov rbx, 42
  2. push rbx
step 0
Loading emulator…

A store runs the memory stage the other way:

  1. Fetch and 2. decode as usual.
  2. Read rbx (42).
  3. Update rsp: rsp ← rsp − 8, making room on the stack. The new rsp is the address to write to.
  4. Memory write: the address goes into the MAR, the value into the MDR, both go out on the buses, and the control unit asserts MEMW (memory write). The data travels from the CPU to memory.
  5. Next.

One read (the fetch) and one write: 1 R · 1 W.

Branches: execute decides the next rip

Live · jl: the last step picks where to go
program— ▸ is the next instruction
  1. mov eax, 3
  2. cmp eax, 5
  3. jl smaller
  4. mov ecx, 0
  5. smaller:
  6. mov ecx, 1
step 0
Loading emulator…

A conditional jump has no data to read or write. Its execute step is a decision: the control unit tests the flags (here SF ≠ OF: 1 ≠ 0, so the jump is taken), and the final step writes the target address into rip instead of the next one. Change 3 to 7 (Edit → Load) and the same last step becomes an ordinary "next instruction".

That's all a jump is at this level: an instruction whose write-back goes to rip. The previous level explains which flags each jump tests.

Code and data share the bus

Look at what the bus spends its time on. In this three-iteration loop, every instruction is fetched from memory, and only the add also touches data:

Live · Where do the bus cycles go?
program— ▸ is the next instruction
  1. mov ecx, 0
  2. loop_top:
  3. add DWORD PTR [rbp-4], 1 ; read-modify-write: 1 data read + 1 data write
  4. inc ecx
  5. cmp ecx, 3
  6. jl loop_top
step 0
Loading emulator…

Press Continue and check the total counter: 19 bus cycles. 13 are instruction fetches (one per instruction executed), and only 6 move data (3 reads and 3 writes of [rbp-4]).

This is the von Neumann architecture: instructions and data live in the same memory and travel over the same bus, so fetching code competes with reading and writing data. That shared path — the von Neumann bottleneck — is why real CPUs add caches, and why they split the first-level cache into separate instruction and data caches (a split cache, or Harvard design at the L1 level), so both can be read in the same cycle.

What this simulator simplifies

The model here runs one instruction at a time, start to finish. Real processors don't:

  • Instruction length: here each instruction occupies one address and the fetch carries its text. Real x86 instructions are 1 to 15 bytes of machine code, and the CPU fetches 16 or more bytes at a time.
  • Pipelining: while one instruction executes, the next is being decoded and the one after that fetched. Several instructions are in flight at once.
  • Micro-ops and out-of-order execution: modern x86 cores translate instructions into internal micro-operations, execute them out of order when their inputs are ready, and retire them in order.
  • Caches: most fetches and loads are served by caches inside the CPU, not by main memory over the system bus.

None of that changes the logical sequence you watched: every instruction is still fetched, decoded, executed, and written back. Pipelines and caches are about doing those steps for many instructions at once, and doing them faster.

Takeaways

  • Every instruction goes through fetch (and rip advance), decode, operand fetch, execute and write-back.
  • The MAR holds the address and the MDR the data for each bus cycle; MEMR and MEMW say which direction.
  • The AGU computes memory addresses; the ALU computes results and flags.
  • A jump is an instruction whose result is written into rip.
  • Code and data share one bus in a von Neumann machine — and instruction fetches are most of the traffic.

In this level

  1. 4.1The fetch–decode–execute cycle
  2. 4.2Datapath, internal and system busesPlanned
  3. 4.3Control units and microcodePlanned
  4. 4.4Pipelining and hazardsPlanned
  5. 4.5Caches and the memory hierarchyPlanned
  6. 4.6Branch prediction and out-of-order executionPlanned
  7. 4.7Multicore and parallel machinesPlanned