The datapath chapter introduced the system bus: address, data and control lines, bus cycles, wait states, masters and arbitration. This chapter goes one level down, to what happens on those wires nanosecond by nanosecond. Two questions run through it: when is a value on a wire valid, and who is allowed to drive the wire at all?
Reading a timing diagram
Hardware datasheets describe buses with timing diagrams: time runs left to right, and each signal gets a row.
- A single wire, like the clock or a control line, is drawn high or low. Its edges are drawn slanted, because no signal changes in zero time.
- A group of wires, like the address or data bus, is drawn as two parallel lines when its value is valid, crossing where the value changes. Shading means "don't care, or not driven".
- Control signals are asserted (doing their job) or negated, whichever voltage that means. A name with a bar or a trailing
#, like MREQ# or RD#, is asserted low.
Around each edge, the diagram marks timing parameters with minimum or maximum values. They are a contract between chips: the CPU guarantees, for example, that the address will be valid at most 1 ns after a clock edge, and the memory promises that its data will be valid a certain time after it sees the address. The same setup and hold rules as for a flip-flop apply to the receiver: data must be stable for a while before the sampling edge and must stay stable for a while after it.
A synchronous read, worked out
On a synchronous bus, everything happens relative to a bus clock. Take a bus clocked at 200 MHz, so each bus cycle lasts 5 ns, and a CPU whose datasheet says:
| Parameter | Value | Meaning |
|---|---|---|
| T_AD | ≤ 1 ns | address valid after the rising edge that starts T1 |
| T_ML | ≥ 0.5 ns | address stable before MREQ# is asserted |
| T_DS | ≥ 0.5 ns | data stable before the edge that samples it (setup) |
| minimum read | 2 cycles | data sampled at the rising edge that ends T2, unless WAIT# is asserted |
With no wait states, the memory has at most 2 × 5 − 1 − 0.5 = 8.5 ns between the address becoming valid and the data having to be on the bus. Each wait state adds a 5 ns cycle. A memory needing t ns therefore needs the smallest number of wait states k such that 8.5 + 5k ≥ t:
| Memory access time | Wait states | Read takes |
|---|---|---|
| 8 ns | 0 | 2 cycles, 10 ns |
| 12 ns | 1 | 3 cycles, 15 ns |
| 15 ns | 2 | 4 cycles, 20 ns |
Here is the 12 ns case, cycle by cycle:
| Signal | T1 (0–5 ns) | T2 (5–10 ns) | T3 (10–15 ns) | at 15 ns |
|---|---|---|---|---|
| CLK | rises at 0 | rises at 5 | rises at 10 | rises: sampling edge |
| ADDRESS | valid by 1 ns | valid | valid | released |
| MREQ#, RD# | asserted at 2.5 ns | asserted | asserted | negated |
| WAIT# | - | asserted by the memory | negated | - |
| DATA | not driven | not driven | valid by 13 ns | sampled |
The CPU looks at WAIT# at the end of T2, sees it asserted, and inserts one cycle. The memory's data are valid at 1 + 12 = 13 ns, 2 ns before the sampling edge at 15 ns, which covers the 0.5 ns setup time with room to spare.
Two lessons follow. First, a synchronous bus rounds everything up to whole cycles: a memory that needs 8.6 ns costs the same as one that needs 13.4 ns. Second, the spacing between MREQ# and the address (T_ML) matters: MREQ# typically drives the chip selects, and the address decoder that produces them must see a stable address first, or two chips can be selected at once for an instant.
Wait states today: CAS latency
Classic shared buses, with many chips hanging off long wires, ran at 5 to 133 MHz. Short point-to-point connections run much faster: a DDR5-4800 memory interface is clocked at 2.4 GHz. The principle of counting whole cycles is the same, and so are the wait states.
Modern DRAM interfaces are synchronous, and their wait states have a name: the CAS latency (CL), the number of clock cycles between a READ command and the first data. JEDEC's standard speed grades give:
| Memory | Clock | Cycle | CAS latency | In nanoseconds |
|---|---|---|---|---|
| DDR3-1600 (grade K) | 800 MHz | 1.25 ns | 11 cycles | 13.75 ns |
| DDR4-3200 (grade AA) | 1600 MHz | 0.625 ns | 22 cycles | 13.75 ns |
| DDR5-4800 (grade B) | 2400 MHz | 0.417 ns | 40 cycles | 16.7 ns |
The clock doubled and tripled, and the number of wait cycles grew with it: the time the DRAM cells themselves need has barely moved. What the faster interfaces bought is bandwidth (more data per nanosecond once it starts flowing), not a shorter wait.
Seen from a program, the whole trip is longer still. Chasing pointers through a randomly shuffled array on this chapter's Apple M2 Ultra, each load took about 0.92 ns when the array was 16 KiB, well inside the L1 cache (about 3 cycles at the 3.24 GHz the core ran at), and 120–133 ns when it spanned 64 MiB to 1 GiB and nearly every load went to DRAM (two runs each). That's around 400 core cycles: the CAS latency is only a part, the rest being the caches that missed on the way, address translation, the memory controller's queues and the on-chip network. To a CPU core, main memory is a device that asks for hundreds of wait states.
Asynchronous buses and the full handshake
A synchronous bus is simple, but it has to be timed for its slowest device, and it can't take advantage of a faster one without a new clock. An asynchronous bus has no clock. Each step is triggered by the previous one, in a full handshake between the master and the slave. For a read:
| Step | Master | Slave |
|---|---|---|
| 1 | puts the address on the bus, asserts MREQ and RD, then asserts MSYN ("master ready") | |
| 2 | sees MSYN, reads its memory, puts the data on the bus, asserts SSYN ("slave ready") | |
| 3 | sees SSYN, latches the data, negates MSYN and releases the address | |
| 4 | sees MSYN negated, negates SSYN and releases the data |
Each event is caused by the one before it, never by a clock tick, so the transfer takes exactly as long as this particular pair of devices needs, and a fast device never waits for a slow one's timing. Because each signal goes up and comes back down, this is also called a four-phase handshake.
Most buses remained synchronous because they are easier to design and verify. But the handshake idea is everywhere:
- Motorola's 68000, the processor of the first Macintosh, ran an asynchronous bus: it asserted an address strobe and waited for the memory to answer with DTACK ("data transfer acknowledge");
- on the I²C bus, still used on every motherboard to talk to small chips, a slow device can hold the clock line low (clock stretching), which is a handshake grafted onto a clocked bus;
- inside chips, the ARM AXI interconnect moves every address and data word with a VALID/READY pair: the sender asserts VALID, the receiver asserts READY, and the transfer happens on the clock edge where both are high. It is a handshake sampled by a clock, the best of both worlds;
- PCI Express replaces the handshake with credits: a receiver announces how much buffer space it has, and the sender stops when the credits run out (PCI Express chapter).
Arbitration: who drives the bus
A shared bus can have only one driver at a time. When several masters (the CPU, a disk controller doing DMA, a network card) want it at once, an arbiter decides. The simplest centralized scheme is the daisy chain. Every device can pull a shared request line. The arbiter answers on a single grant line, which is wired through the devices in series: a device that isn't requesting passes the grant on, and one that is requesting keeps it.
Try it: Click an input switch in the circuit (or its button above) to toggle it. The gates and the truth table follow.
| REQ1 | REQ2 | REQ3 | GRANT | GNT1 | GNT2 | GNT3 |
|---|---|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0 | 0 | 1 | 1 | 0 | 0 | 1 |
| 0 | 1 | 0 | 1 | 0 | 1 | 0 |
| 0 | 1 | 1 | 1 | 0 | 1 | 0 |
| 1 | 0 | 0 | 1 | 1 | 0 | 0 |
| 1 | 0 | 1 | 1 | 1 | 0 | 0 |
| 1 | 1 | 0 | 1 | 1 | 0 | 0 |
| 1 | 1 | 1 | 1 | 1 | 0 | 0 |
One grant line, wired through the devices in series. A device that is requesting keeps the grant (GNT); one that isn't passes it on. So the device closest to the arbiter always wins: turn on REQ2 and REQ3, then REQ1. Each device adds a gate delay to the grant's trip, which is why long daisy chains are slow.
Unit-delay model: every gate takes one step to react. With auto off, toggle switches and press step to watch the change travel gate by gate.
Devices 2 and 3 are both requesting, and device 2, closer to the arbiter, has the grant (GNT2). Then:
- Turn off REQ2: the grant ripples one device further, and GNT3 lights up. From a standing start, REQ3 alone needs 4 gate delays to be granted, against 2 for REQ1: each device on the chain adds delay.
- Turn REQ2 back on, and then REQ1. Device 1 now wins. But watch closely with auto off: for two gate delays, GNT1 and GNT2 are both on, because GNT1 rises before the "pass it on" signal to device 2 has fallen.
That overlap is why real arbiters never switch the grant in the middle of a transfer. A device that wins asserts a third line, BUSY (or an acknowledge), for as long as it uses the bus, and the grant only moves once BUSY is released. It is also why the arbitration is usually done during one transfer to decide the next, so that no bus cycles are lost.
The daisy chain's fixed priorities have a darker side: a device near the arbiter that requests constantly can starve the ones behind it forever. Real systems add several request and grant lines with different priority levels, or make the arbiter round-robin, giving the grant to each requester in turn. The PCI bus gave every slot its own request and grant pins to a central arbiter whose algorithm the specification left open, precisely so that it could be fair.
Decentralized arbitration
Arbitration can also be done without any arbiter. In one scheme, each device reads the arbitration line from its neighbor and passes it on only if it doesn't want the bus itself, like a daisy chain whose head is tied permanently to "grant".
A more elegant trick uses open-drain (or open-collector) lines, where any device can pull the line to 0 and nobody drives it to 1; a resistor does. The line is then the AND of what all devices send, and two devices can transmit at the same time without damage. On I²C and on the CAN bus used in cars, competing masters send their message headers bit by bit while listening: a device that sends a 1 but reads back a 0 knows that someone with a higher priority is transmitting, and quietly drops out. The winner never even notices there was a contest, and its message goes through undamaged.
Other kinds of bus cycles
Block transfers. Fetching a whole cache line one word at a time would repeat the address phase each time. Instead, the master sends one address and the slave returns consecutive words on consecutive cycles, a burst. DDR memory works only this way: a DDR4 read returns a burst of 8 transfers of 64 bits, and a DDR5 read a burst of 16 transfers on a 32-bit subchannel. Both are 64 bytes, one cache line on most processors, delivered by a single READ command.
The memory controller also overlaps commands to different banks: while one bank is still opening a row (ACTIVATE), another can be sending its burst, and a third closing its row (PRECHARGE). On DDR3, for example, four accesses can be overlapped this way. Since everything on the interface is synchronous, the controller knows in which cycle each answer will come back and schedules the commands like a pipeline.
Locked cycles. Two CPUs incrementing the same counter must not both read the old value before either writes. Old buses had a read-modify-write cycle that kept the bus locked between the read and the write; the 8086's lock prefix asserted a LOCK# pin for exactly that. The instruction still exists:
; x86-64: __atomic_fetch_add(&counter, 1, __ATOMIC_SEQ_CST)
lock inc QWORD PTR [rip+counter]
; ARM64 (v8.1 and later): the same, as one atomic instruction
mov w8, #1
adrp x9, counter
add x9, x9, :lo12:counter
ldaddal x8, x8, [x9]
Today the lock almost never reaches a bus. The core takes the cache line in exclusive ownership through the cache-coherence protocol, performs the read and the write in its own cache, and simply delays answering other cores' requests for that line until it's done. Only an atomic operation that straddles two cache lines still falls back to locking the whole memory system, which is so disruptive that recent Linux kernels can detect such split locks and warn about or slow down the program that causes them.
Interrupt cycles. On classic PCs, the CPU acknowledged an interrupt with a special bus cycle during which the 8259A interrupt controller put a vector number on the data bus. PCI Express devices instead send an ordinary memory write to a special address, as the interrupts chapter explains: the interrupt has become just another transaction on the bus.
Takeaways
- Timing diagrams give each signal a row and each edge a guaranteed minimum or maximum time. Setup and hold rules decide when data may be sampled.
- On a synchronous bus, a slow device asks for wait states, whole cycles at a time. On a 200 MHz bus with a 2-cycle minimum read, a 12 ns memory needs one wait state.
- DRAM's wait states are its CAS latency: 11 cycles for DDR3-1600, 22 for DDR4-3200, 40 for DDR5-4800, about 14 to 17 ns every time. A load that misses all the caches on the M2 Ultra took 120–133 ns.
- An asynchronous bus replaces the clock with a full handshake (MSYN/SSYN, or DTACK); VALID/READY, clock stretching and credits carry the idea on.
- A daisy chain grants the bus to the requesting device closest to the arbiter, one gate delay per device. BUSY lines, priority levels, round-robin and bitwise arbitration on open-drain lines (I²C, CAN) fix its problems.
- Beyond single reads and writes, buses do bursts (64 bytes per DDR4 or DDR5 read), locked read-modify-write cycles (now done inside the cache) and interrupt cycles (now memory writes).