Skip to content

Level 4 · Chapter 4.3

Control units and microcode

What drives the datapath on every clock cycle: control signals, microinstructions and the control store, choosing the next step, hardwired versus microprogrammed control, and microcode in today's x86 CPUs.

The datapath can move values and compute, but it doesn't decide anything on its own. On every clock cycle, something has to tell it which register goes on the bus, which operation the ALU performs, which register gets the result, and whether memory is read or written. That something is the control unit.

On each cycle, the control unit produces two things:

  1. the value of every control signal for this cycle;
  2. which step comes next.

There are two classic ways to build one: as fixed logic (hardwired control), or as a small program stored inside the CPU (microprogrammed control, or microcode).

Control signals

A control signal is one wire that turns one thing on or off. For the datapath of the previous chapter, the signals include:

  • bus sources: which register drives the internal bus — eax? rbp? the MDR?
  • ALU function: add, subtract, AND, pass through…
  • destinations: which register loads the result at the end of the cycle;
  • memory: load the MAR, load the MDR, assert MEMR or MEMW;
  • sequencing: update rip, load the IR.

Some combinations are forbidden. Two registers must never drive the same bus at once — in real hardware that can even damage the circuit. So designers encode the bus source as a small number and let a decoder turn it into one enable line. With 16 possible sources, 4 bits are enough instead of 16 wires.

Tanenbaum's example machine, the Mic-1, needs 29 control signals, and encoding brings them down to 24 bits.

Microinstructions and the control store

A microinstruction is one cycle's worth of control signals, written down as a row of bits, plus a field saying which microinstruction comes next. A microprogram is a list of these rows, and it's stored in a small, fast memory inside the CPU: the control store, usually a ROM.

The control store has its own address and data registers:

  • the MPC (microprogram counter) holds the address of the current microinstruction;
  • the MIR (microinstruction register) holds the microinstruction itself. Its bits drive the control signals: most go straight to the datapath, and encoded fields such as the bus source go through a decoder first.

Mic-1's microinstructions are 36 bits wide: 24 control bits, 9 bits of next address, and 3 bits that choose how the next address is formed. Its control store holds 512 of them. That's the whole machine: a clock, a datapath, a control store, and a tiny loop that reads one microinstruction per cycle.

A microprogram for one instruction

Here is add eax, DWORD PTR [rbp-4] — the instruction from the fetch–decode–execute chapter — written as a microprogram for the simulator's datapath. Each row is one step, and the columns are groups of control signals:

StepBus sourceALU / AGULoadsMemoryNext
fetchrip—MAR, then IRMEMRdecode
decodeIR———jump to the add r, [m] routine
1eax—ALU input—2
2rbpAGU: rbp − 4MAR—3
3——MDRMEMR4
4MDRALU: addflags—5
5ALU output—eax—6
6——rip ← next—fetch

Step through it in the simulator and compare. Each micro-operation shown on the diagram matches one row:

Live · add eax, [rbp-4]: one row per micro-operation
program— ▸ is the next instruction
  1. mov DWORD PTR [rbp-4], 10
  2. mov eax, 5
  3. add eax, DWORD PTR [rbp-4]
step 0
Loading emulator…

The fetch and decode rows are shared by every instruction. Only the rows after decode are specific to add. A different instruction, such as mov or push, has its own short routine elsewhere in the control store, and all of them end by going back to fetch.

Choosing the next step

A normal program runs its instructions in address order unless there's a jump. A microprogram doesn't: every microinstruction says explicitly which one comes next. That makes three kinds of "next" possible:

  1. Unconditional: go to address n. Most rows do this.
  2. Conditional: go to one of two addresses, depending on a status bit from the ALU — for instance, whether the last result was zero. Mic-1 does this by OR-ing the saved Z or N bit into the top bit of the next address, so the two possible successors are n and n + 256.
  3. Multiway (dispatch): go to an address computed from the opcode. This is what decode is. Mic-1 ORs the opcode byte into the next address, which jumps straight to the routine for that instruction in a single step, out of up to 256 possible destinations.

Conditional micro-branches are how one instruction can do different things. jl in the simulator shows the idea one level up — the control unit tests the flags and chooses between two possible next values of rip:

Live · jl: the control unit picks one of two successors
program— ▸ is the next instruction
  1. mov eax, 3
  2. cmp eax, 5
  3. jl smaller
  4. mov ecx, 0
  5. smaller:
  6. mov ecx, 1
step 0
Loading emulator…

The third micro-operation reads SF=1 OF=0 → TAKEN. Change 3 to 7 and it becomes not taken.

One instruction, many steps

With a conditional micro-branch, a microprogram can loop, so one machine instruction can do an amount of work that depends on the data. x86 has instructions like that. rep stosb fills rcx bytes at [rdi] with al:

Live · rep stosb: a loop inside one instruction
program— ▸ is the next instruction
  1. .bss
  2. buf: .zero 8
  3. .text
  4. _start:
  5. lea rdi, [rip+buf]
  6. mov ecx, 4
  7. mov al, 0x41
  8. rep stosb
step 0
Loading emulator…

That single instruction takes 15 micro-operations: fetch and decode, then four rounds of write a byte, advance rdi, decrement rcx, then the rip update. The bus counter shows 1 R · 4 W: one instruction fetch and four data writes. Set ecx to 100 and the same one instruction takes 100 rounds.

On real Intel and AMD processors, rep string instructions are among the instructions that run from microcode: the decoder hands them to a microcode sequencer, which feeds out the loop's internal operations.

Hardwired or microprogrammed?

A hardwired control unit computes the control signals directly with logic gates, from the opcode and a small state counter. A microprogrammed one looks them up in the control store.

HardwiredMicroprogrammed
Speedfast: no control-store read on the critical patha control-store read every cycle
Complex instructionshard: every case is more logiceasy: just more microcode
Design and bug fixeschanges mean redesigning logicchanges mean rewriting the microprogram
Same instruction set on different hardwaredifficultnatural

History has swung between the two:

  • 1951: Maurice Wilkes proposes microprogramming. A simple machine interprets a richer instruction set, and needs far fewer of the unreliable vacuum tubes of the time.
  • 1960s–70s: microcode dominates. IBM's System/360 (1964) used it to run one instruction set on a whole family of machines, most of them microprogrammed, with very different hardware from the cheapest model to the fastest. Instruction sets grow richer, and microprograms grow bigger and slower.
  • 1980s: the RISC designs drop microcode. They use simple instructions, run directly by hardwired control, one per cycle.
  • Today: x86 is a hybrid.

Microcode in a modern x86 CPU

A current Intel or AMD core doesn't interpret most instructions with microcode:

  • The decoders translate each x86 instruction into internal micro-ops (µops), mostly with hardwired logic. A simple instruction like add eax, ebx becomes one µop. add eax, [rbp-4] becomes a load µop plus an add µop.
  • The decoded µops are kept in a µop cache, so a hot loop doesn't need decoding again.
  • Only complex instructions — rep string operations, cpuid, many system instructions — go to the microcode sequencer, which streams longer µop sequences from an on-chip ROM.

Note that the simulator's micro-operations are smaller than Intel's µops. They're closer to the rows of a microprogram: one bus transfer or one ALU operation each.

The microcode can also be patched. CPU vendors publish microcode updates that the firmware or the operating system loads at every boot — the update is not permanent, so it has to be reapplied each time. On Linux, the current revision appears on the microcode line of /proc/cpuinfo. These updates fix processor bugs (errata) and add security mitigations. Several of the Spectre defenses of 2018 were delivered this way, as new control registers. In 2019, a microcode update against the MDS vulnerabilities made an old instruction, verw, also clear internal CPU buffers. A processor already in your machine gained a new behavior — which only makes sense if part of the CPU is software.

Takeaways

  • On every cycle, the control unit sets the datapath's control signals and chooses the next step.
  • A microinstruction is one cycle's control signals plus a next address. The microprogram lives in the control store, addressed by the MPC and read into the MIR.
  • Next addresses can be unconditional, conditional on ALU status, or a multiway dispatch on the opcode, which is how decode works.
  • Micro-branches let a single instruction loop, as rep stosb does.
  • Hardwired control is faster; microcode makes complex instructions and fixes easier. Modern x86 decodes most instructions in hardware, uses microcode for the complex ones, and accepts microcode updates at boot.

In this level

  1. 4.1The fetch–decode–execute cycle
  2. 4.2Datapath, internal and system buses
  3. 4.3Control units and microcode
  4. 4.4Pipelining and hazardsPlanned
  5. 4.5Caches and the memory hierarchyPlanned
  6. 4.6Branch prediction and out-of-order executionPlanned
  7. 4.7Multicore and parallel machinesPlanned