The datapath can move values and compute, but it doesn't decide anything on its own. On every clock cycle, something has to tell it which register goes on the bus, which operation the ALU performs, which register gets the result, and whether memory is read or written. That something is the control unit.
On each cycle, the control unit produces two things:
- the value of every control signal for this cycle;
- which step comes next.
There are two classic ways to build one: as fixed logic (hardwired control), or as a small program stored inside the CPU (microprogrammed control, or microcode).
Control signals
A control signal is one wire that turns one thing on or off. For the datapath of the previous chapter, the signals include:
- bus sources: which register drives the internal bus —
eax?rbp? the MDR? - ALU function: add, subtract, AND, pass through…
- destinations: which register loads the result at the end of the cycle;
- memory: load the MAR, load the MDR, assert MEMR or MEMW;
- sequencing: update
rip, load the IR.
Some combinations are forbidden. Two registers must never drive the same bus at once — in real hardware that can even damage the circuit. So designers encode the bus source as a small number and let a decoder turn it into one enable line. With 16 possible sources, 4 bits are enough instead of 16 wires.
Tanenbaum's example machine, the Mic-1, needs 29 control signals, and encoding brings them down to 24 bits.
Microinstructions and the control store
A microinstruction is one cycle's worth of control signals, written down as a row of bits, plus a field saying which microinstruction comes next. A microprogram is a list of these rows, and it's stored in a small, fast memory inside the CPU: the control store, usually a ROM.
The control store has its own address and data registers:
- the MPC (microprogram counter) holds the address of the current microinstruction;
- the MIR (microinstruction register) holds the microinstruction itself. Its bits drive the control signals: most go straight to the datapath, and encoded fields such as the bus source go through a decoder first.
Mic-1's microinstructions are 36 bits wide: 24 control bits, 9 bits of next address, and 3 bits that choose how the next address is formed. Its control store holds 512 of them. That's the whole machine: a clock, a datapath, a control store, and a tiny loop that reads one microinstruction per cycle.
A microprogram for one instruction
Here is add eax, DWORD PTR [rbp-4] — the instruction from the fetch–decode–execute chapter — written as a microprogram for the simulator's datapath. Each row is one step, and the columns are groups of control signals:
| Step | Bus source | ALU / AGU | Loads | Memory | Next |
|---|---|---|---|---|---|
| fetch | rip | — | MAR, then IR | MEMR | decode |
| decode | IR | — | — | — | jump to the add r, [m] routine |
| 1 | eax | — | ALU input | — | 2 |
| 2 | rbp | AGU: rbp − 4 | MAR | — | 3 |
| 3 | — | — | MDR | MEMR | 4 |
| 4 | MDR | ALU: add | flags | — | 5 |
| 5 | ALU output | — | eax | — | 6 |
| 6 | — | — | rip ← next | — | fetch |
Step through it in the simulator and compare. Each micro-operation shown on the diagram matches one row:
- mov DWORD PTR [rbp-4], 10
- mov eax, 5
- add eax, DWORD PTR [rbp-4]
The fetch and decode rows are shared by every instruction. Only the rows after decode are specific to add. A different instruction, such as mov or push, has its own short routine elsewhere in the control store, and all of them end by going back to fetch.
Choosing the next step
A normal program runs its instructions in address order unless there's a jump. A microprogram doesn't: every microinstruction says explicitly which one comes next. That makes three kinds of "next" possible:
- Unconditional: go to address n. Most rows do this.
- Conditional: go to one of two addresses, depending on a status bit from the ALU — for instance, whether the last result was zero. Mic-1 does this by OR-ing the saved Z or N bit into the top bit of the next address, so the two possible successors are n and n + 256.
- Multiway (dispatch): go to an address computed from the opcode. This is what decode is. Mic-1 ORs the opcode byte into the next address, which jumps straight to the routine for that instruction in a single step, out of up to 256 possible destinations.
Conditional micro-branches are how one instruction can do different things. jl in the simulator shows the idea one level up — the control unit tests the flags and chooses between two possible next values of rip:
- mov eax, 3
- cmp eax, 5
- jl smaller
- mov ecx, 0
- smaller:
- mov ecx, 1
The third micro-operation reads SF=1 OF=0 → TAKEN. Change 3 to 7 and it becomes not taken.
One instruction, many steps
With a conditional micro-branch, a microprogram can loop, so one machine instruction can do an amount of work that depends on the data. x86 has instructions like that. rep stosb fills rcx bytes at [rdi] with al:
- .bss
- buf: .zero 8
- .text
- _start:
- lea rdi, [rip+buf]
- mov ecx, 4
- mov al, 0x41
- rep stosb
That single instruction takes 15 micro-operations: fetch and decode, then four rounds of write a byte, advance rdi, decrement rcx, then the rip update. The bus counter shows 1 R · 4 W: one instruction fetch and four data writes. Set ecx to 100 and the same one instruction takes 100 rounds.
On real Intel and AMD processors, rep string instructions are among the instructions that run from microcode: the decoder hands them to a microcode sequencer, which feeds out the loop's internal operations.
Hardwired or microprogrammed?
A hardwired control unit computes the control signals directly with logic gates, from the opcode and a small state counter. A microprogrammed one looks them up in the control store.
| Hardwired | Microprogrammed | |
|---|---|---|
| Speed | fast: no control-store read on the critical path | a control-store read every cycle |
| Complex instructions | hard: every case is more logic | easy: just more microcode |
| Design and bug fixes | changes mean redesigning logic | changes mean rewriting the microprogram |
| Same instruction set on different hardware | difficult | natural |
History has swung between the two:
- 1951: Maurice Wilkes proposes microprogramming. A simple machine interprets a richer instruction set, and needs far fewer of the unreliable vacuum tubes of the time.
- 1960s–70s: microcode dominates. IBM's System/360 (1964) used it to run one instruction set on a whole family of machines, most of them microprogrammed, with very different hardware from the cheapest model to the fastest. Instruction sets grow richer, and microprograms grow bigger and slower.
- 1980s: the RISC designs drop microcode. They use simple instructions, run directly by hardwired control, one per cycle.
- Today: x86 is a hybrid.
Microcode in a modern x86 CPU
A current Intel or AMD core doesn't interpret most instructions with microcode:
- The decoders translate each x86 instruction into internal micro-ops (µops), mostly with hardwired logic. A simple instruction like
add eax, ebxbecomes one µop.add eax, [rbp-4]becomes a load µop plus an add µop. - The decoded µops are kept in a µop cache, so a hot loop doesn't need decoding again.
- Only complex instructions —
repstring operations,cpuid, many system instructions — go to the microcode sequencer, which streams longer µop sequences from an on-chip ROM.
Note that the simulator's micro-operations are smaller than Intel's µops. They're closer to the rows of a microprogram: one bus transfer or one ALU operation each.
The microcode can also be patched. CPU vendors publish microcode updates that the firmware or the operating system loads at every boot — the update is not permanent, so it has to be reapplied each time. On Linux, the current revision appears on the microcode line of /proc/cpuinfo. These updates fix processor bugs (errata) and add security mitigations. Several of the Spectre defenses of 2018 were delivered this way, as new control registers. In 2019, a microcode update against the MDS vulnerabilities made an old instruction, verw, also clear internal CPU buffers. A processor already in your machine gained a new behavior — which only makes sense if part of the CPU is software.
Takeaways
- On every cycle, the control unit sets the datapath's control signals and chooses the next step.
- A microinstruction is one cycle's control signals plus a next address. The microprogram lives in the control store, addressed by the MPC and read into the MIR.
- Next addresses can be unconditional, conditional on ALU status, or a multiway dispatch on the opcode, which is how decode works.
- Micro-branches let a single instruction loop, as
rep stosbdoes. - Hardwired control is faster; microcode makes complex instructions and fixes easier. Modern x86 decodes most instructions in hardware, uses microcode for the complex ones, and accepts microcode updates at boot.