The previous chapter stored one bit in a latch or a flip-flop. A computer needs billions of bits, organized so that any group of them can be found by number. This chapter goes from a single register to a memory you can read and write by address.
Registers
An n-bit register is n D flip-flops that share a clock: on every rising edge, all of them capture their D inputs at once. Real registers add two controls:
- load (or write enable): a multiplexer in front of each D input chooses between the new value and the flip-flop's own output, so the register only changes when load is 1;
- clear: forces every flip-flop to 0, for example at reset.
Tanenbaum's 8-bit register also has an inverter on its clock input that seems pointless, since each flip-flop inverts it again. Its job is electrical: one input can't drive eight flip-flops on its own, so the inverter serves as an amplifier. Real circuits are full of such buffers.
The register file
The CPU's registers — the 16 of x86-64 that the assembly level names — sit together in a register file, and it has ports:
- a read port is a big multiplexer: the register number (4 bits) selects which register's 64 bits appear at the output. Two read ports let an instruction read both operands in the same cycle;
- a write port is a decoder: the destination register number enables the load input of exactly one register, which captures the result at the clock edge.
This is the register file of the datapath chapter, seen from the inside. Out-of-order cores have register files with hundreds of physical registers and a dozen or more ports, which makes them among the most expensive structures on the chip.
A memory you can address
Registers don't scale: each one needs its own wires in and out. A memory shares its wires between all its words. You give it an address and a read or write command, and it reads or writes just the word at that address. Here is a complete one, with 2 words of 2 bits, built from the same D latches:
Each box is a D latch holding one bit. The address A picks a word (NOT A selects word 0, A selects word 1). With WE = 1, the selected word's latches open and store D1·D0; the other word is untouched. With WE = 0 nothing changes, and O1·O0 always shows the selected word, read through the AND gates and the ORs. Real memory chips are this, with thousands of rows and columns.
Unit-delay model: every gate takes one step to react. With auto off, toggle switches and press step to watch the change travel gate by gate.
Try it:
- Write word 0. Set A = 0, D1·D0 = 1·0, then turn WE on and off again. Only the word 0 latches open: the enable gates combine WE with the decoded address.
- Write word 1. Set A = 1, D1·D0 = 0·1, and pulse WE again.
- Read. Change D to anything — nothing is stored while WE = 0. Flip A between 0 and 1: the outputs O1·O0 show 10 for word 0 and 01 for word 1.
Every part of a real memory is here:
- an address decoder turns the address into one active word select line (here just NOT A and A);
- write enable combines with the word select so that only the addressed word's latches open;
- on a read, the word select opens the read gates of the addressed word, and an OR per column (or, in real chips, a shared bit line) carries its bits to the outputs. The other words contribute 0s.
Grow it by adding rows (words) and columns (bits). k address bits select one of 2ᵏ words, and a word can be any width.
Rows and columns
A real memory chip is a square grid: each row is a word line, each column a bit line, and one cell sits at each crossing. The address is split in two:
- the row half drives a row decoder, which activates one word line: every cell in that row puts its bit on its column's bit line;
- the column half drives a multiplexer that picks the bits you actually want from that row.
A square layout keeps both decoders small. A 1-million-bit array arranged as 1024 × 1024 needs two 10-bit decoders instead of one 20-bit decoder with a million outputs.
DRAM chips also save pins by sending the row and column halves of the address one after the other on the same pins. The row goes first, latched when the RAS (row address strobe) signal is asserted, then the column, latched by CAS. Once a row is open, reading more columns from it is fast, which is one reason caches fetch whole lines.
The pins of a memory chip
| Signal | Role |
|---|---|
| address | which word (or, for DRAM, row then column) |
| data | in on a write, out on a read — usually the same pins |
| CS̅ (chip select) | this chip should respond; the other chips ignore the bus |
| WE̅ (write enable) | write rather than read |
| OE̅ (output enable) | drive the data pins; otherwise they float |
The bar marks signals that are asserted low: the action happens when the pin is at 0. Datasheets are careful to say a signal is asserted or negated rather than high or low, because which voltage means "yes" depends on the pin. OE controls a tri-state output buffer, as described in the buses section: when the chip isn't selected, or is being written to, its data pins let go of the shared data bus.
Real memory cells
The latch-based memory above uses more than a dozen transistors per bit — four NAND gates are already 16 — far too many for gigabytes. Real memories use much smaller cells:
- SRAM (static RAM): two cross-coupled inverters — the same feedback idea as a latch — plus two access transistors, 6 transistors per bit. It holds its value as long as power is on, and it's fast. It's used for registers and caches.
- DRAM (dynamic RAM): 1 transistor and 1 capacitor per bit. The bit is a tiny charge that leaks away, so every row has to be refreshed — read and rewritten — every few tens of milliseconds. It's much denser and cheaper than SRAM but slower, and it's used for main memory.
Because memory is such a regular grid, it scales well. The number of bits per chip has followed Moore's law: Gordon Moore's observation, revised in 1975, that the number of transistors on a chip doubles about every two years. (The "18 months" often quoted, including in Tanenbaum's text, is a popular variant, not Moore's own figure.) How SRAM, DRAM, ROM and flash cells work is the subject of a later chapter on memory chips.
Takeaways
- A register is n flip-flops sharing a clock, usually with load and clear controls. The register file adds read ports (multiplexers) and write ports (decoders).
- A memory shares its wires between all its words: an address decoder picks the word, write enable controls storing, and read gates or bit lines deliver the selected word.
- Real chips are grids of word lines and bit lines, with row and column decoders. DRAM sends the row and column addresses one after the other (RAS, CAS).
- Chips have chip select, write enable and output enable pins, often asserted low. Tri-state outputs let many chips share one data bus.
- SRAM cells use 6 transistors and hold their value; DRAM cells use 1 transistor and 1 capacitor and need refreshing.