Skip to content

Level 4 · Chapter 4.2

Datapath, internal and system buses

The part of the CPU that actually computes — registers, ALU, AGU and the buses between them — and the system bus that connects it to memory: widths, bus cycles, wait states, masters and arbitration.

The previous chapter followed one instruction through its steps. This one looks at the hardware those steps run on.

Inside a CPU, the part that holds and transforms data is the datapath: the registers, the ALU that computes on them, and the buses that carry values between the two. Around it sit the control unit, which decides what the datapath does on each clock cycle, and the bus interface, which connects the CPU to memory through the system bus.

In the simulator's CPU & buses view, the datapath is the left half of the CPU box: register file, internal bus, ALU and AGU. The bus interface with its MAR and MDR is the border between the CPU and the system bus.

The datapath cycle

Most instructions boil down to the same short trip: take two values out of registers, run them through the ALU, and store the result in a register. That trip is the datapath cycle, and it's the heart of the CPU. How fast it runs largely sets how fast the machine runs.

Live · add eax, ebx: one trip through the datapath
program— ▸ is the next instruction
  1. mov eax, 5
  2. mov ebx, 7
  3. add eax, ebx
step 0
Loading emulator…

The add takes seven micro-operations, and five of them are the datapath cycle:

  1. Fetch and 2. decode, as in every instruction.
  2. Read eax (5) from the register file onto the internal bus.
  3. Read ebx (7) the same way.
  4. ALU: 5 + 7 = 12, and the flags are updated.
  5. Write-back: 12 into eax.
  6. Next: rip advances.

No memory is involved: 1 R · 0 W — just the fetch.

One bus or three

The simulator moves the two operands one after the other over a single internal bus. That is a real design: with one bus, the first operand is parked in a holding register at the ALU's input while the second one travels, and the result comes back over the same bus. It takes few wires, but several transfers per operation.

Faster designs give the datapath more buses: two to carry the operands to the ALU at the same time, and a third to carry the result back. The register file then needs two read ports and a write port — it can deliver two registers and accept one in the same cycle. Modern x86 cores go much further, with register files that have many ports feeding several ALUs at once.

What happens inside one clock cycle

add eax, ebx reads eax and writes eax. How can a register be read and written in the same cycle without mixing up old and new values? Through timing. In a typical design, one clock cycle goes like this:

  1. The control unit sets up the control signals: which registers drive the buses, which ALU operation, which register gets the result.
  2. The selected registers put their values on the buses.
  3. The ALU — a combinational circuit that is always computing — produces a stable result once its inputs are stable.
  4. The result propagates back to the registers.
  5. On the clock edge, the destination register latches the result.

The old value of eax is on the bus for the whole cycle, and the new value is only stored at the very end. So reading and writing the same register in one cycle is safe.

This is also what limits the clock speed. The clock period must be longer than steps 1 to 4 added together — the slowest path through the datapath. Make the path shorter, or cut it into pieces, and the clock can tick faster. Cutting it into pieces is exactly what pipelining does, as a later chapter in this level shows.

Computing addresses: the AGU

Not every addition happens in the ALU. Memory operands like [rbx+rcx*4+8] need an address computed first, and x86 CPUs have a dedicated address generation unit (AGU) for that. It can work at the same time as the ALU.

lea ("load effective address") shows the AGU on its own. It computes an address and stores it in a register, without touching memory:

Live · lea: the AGU without a memory access
program— ▸ is the next instruction
  1. mov rbx, 0x1000
  2. mov rcx, 3
  3. lea rax, [rbx+rcx*4+8]
step 0
Loading emulator…

The AGU computes 0x1000 + 3×4 + 8 = 0x1014, and it goes straight into rax. There is no data bus cycle: 1 R · 0 W. That's why compilers often use lea for ordinary arithmetic, such as lea eax, [rdi+rdi*2] to multiply by 3.

Leaving the CPU: MAR and MDR

To reach memory, the datapath goes through two registers of the bus interface:

  • the MAR (memory address register) holds the address and drives the address bus;
  • the MDR (memory data register) holds the data going to memory or coming back, on the data bus.
Live · A store, then a load
program— ▸ is the next instruction
  1. mov eax, 42
  2. mov DWORD PTR [rbp-8], eax
  3. mov edx, DWORD PTR [rbp-8]
step 0
Loading emulator…

The demo stops after the store: eax (42) is read onto the internal bus, the AGU computes rbp − 8 = 0x7fffffd8, the address goes into the MAR and the value into the MDR, and the control unit asserts MEMW. That's 1 R · 1 W.

Press Step for the load: the AGU computes the same address, the MAR drives it, MEMR is asserted, and memory puts 42 on the data bus into the MDR, which then goes to edx. That's 2 R · 0 W: the fetch plus the data read.

The system bus

A bus is a set of shared wires that several devices are connected to. The system bus groups its lines into three sets:

LinesCarryDirection
Addresswhich memory location or devicefrom the CPU (or another master)
Datathe value being read or writtenboth ways
Controlwhat kind of operation, and when: read or write, memory or I/O, wait, interrupt, clock…various

The exact control lines depend on the bus. The simulator uses the ISA bus names MEMR and MEMW. Many textbook buses use a MREQ line ("this is a memory access") together with RD and WR.

Width

With n address lines, a bus can select 2ⁿ different locations. This is a fixed cost: every line is a wire, a pin and a trace on the board. Intel's history shows the problem. The 8088 in the original IBM PC had 20 address lines, so it could address 1 MB. The 80286 needed 24 lines for 16 MB, and the 80386 needed 32 lines for 4 GB. Each extension had to stay compatible with the old one.

The data bus width sets how many bits move per transfer. Bus bandwidth is width × transfers per second, so there are two ways to raise it: make the bus wider, or make it faster. Faster is hard, because signals on different lines don't arrive at exactly the same time (bus skew), and the gap gets worse as the clock rises.

Some buses save wires by multiplexing: the same lines carry the address first, then the data. The 8086's 16 lines AD0–AD15 worked this way. It's cheaper but slower, because address and data can no longer travel together.

Bus cycles and wait states

On a synchronous bus, everything is timed by a bus clock, and each transfer takes a whole number of bus cycles. A read goes like this:

  1. The CPU puts the address on the address lines.
  2. It asserts the control lines for a memory read.
  3. If the memory can't answer in time, it asserts WAIT, and the CPU inserts wait states — extra bus cycles — until the memory is ready.
  4. The memory drives the data lines, and the CPU latches the data at a fixed point in the cycle, then releases the control lines.

For example, a 100 MHz bus has 10 ns cycles. A read then takes a fixed minimum number of cycles, plus one more for every wait state the memory needs.

An asynchronous bus has no master clock. Instead, the two sides use a handshake: the master signals "address and command are ready", the slave answers "data is ready", the master acknowledges, and the slave releases. Each transfer takes exactly as long as that pair of devices needs. It's more flexible, but harder to design.

Masters, slaves and arbitration

A device that starts a transfer is a master, and the one that answers is a slave. The CPU is usually the master and memory the slave. Memory is always a slave. But other devices can be masters too. A disk or network controller doing DMA (direct memory access) writes into memory on its own, without the CPU copying each word.

When several masters want the bus at the same moment, someone has to choose. That's bus arbitration:

  • In centralized arbitration, a single arbiter grants the bus. With daisy chaining, the grant signal passes from device to device, and the first one that wants the bus keeps it, so a device's position sets its priority.
  • In decentralized arbitration, the devices settle it among themselves over shared request lines.

Buses in a modern PC

The single shared system bus is mostly history. Since the mid-2000s — AMD's Athlon 64 in 2003, Intel's Nehalem in 2008 — the memory controller has been inside the CPU chip. Memory is reached through dedicated DDR channels: 64 data bits per DDR4 channel, and two independent 32-bit subchannels per DDR5 module. Devices are attached through PCI Express, which despite its name isn't a bus at all. It's a set of point-to-point serial lanes (x1, x4, x16…) that carry packets. And inside the chip, cores and caches talk over a ring or a mesh interconnect.

The model from this chapter still holds inside all of that. Every transfer still has an address, data and a command saying read or write — only the wires carrying them have changed.

Takeaways

  • The datapath is the registers, the ALU and the buses between them. The datapath cycle — read registers, compute, write back — is the heart of the CPU.
  • Registers are read early in the cycle and written on the clock edge, so an instruction can read and write the same register. The slowest path through the datapath sets the clock period.
  • The AGU computes addresses; lea uses it without touching memory.
  • The MAR and MDR connect the datapath to the address and data buses, and control lines say what kind of transfer it is.
  • Bus width limits how much memory can be addressed and how much data moves per transfer. Slow devices add wait states, and arbitration decides which master uses the bus.

In this level

  1. 4.1The fetch–decode–execute cycle
  2. 4.2Datapath, internal and system buses
  3. 4.3Control units and microcode
  4. 4.4Pipelining and hazardsPlanned
  5. 4.5Caches and the memory hierarchyPlanned
  6. 4.6Branch prediction and out-of-order executionPlanned
  7. 4.7Multicore and parallel machinesPlanned