Skip to content

Level 8 · Chapter 8.2

Building gates from CMOS

How every gate on a modern chip is wired from transistors: the CMOS inverter, NAND and NOR as pull-up and pull-down networks, why AND costs more than NAND, where the power goes (P = αCV²f and leakage), fan-out, and propagation delay measured with a ring oscillator.

The previous chapter described the transistor as a switch: an nMOS transistor closes when its gate is 1, a pMOS transistor closes when its gate is 0. This chapter wires those switches into the gates of the digital logic level using CMOS, the design every processor, memory and flash chip has used for decades. The method is short enough to fit on one page, and it explains several things the logic level took on faith: why NAND and NOR are the "natural" gates, why a gate has a delay, and where a chip's power goes.

The rule: a pull-up and a pull-down network

A static CMOS gate has one output and two networks of switches:

  • a pull-down network of nMOS transistors between the output and ground, which conducts exactly when the output should be 0;
  • a pull-up network of pMOS transistors between the supply and the output, which conducts exactly when the output should be 1.

The two networks are complementary: for every combination of inputs, exactly one of them conducts. So the output is always firmly connected to either the supply or ground, never to both, and never to neither.

The pull-up network is the dual of the pull-down network: wherever the pull-down has transistors in series, the pull-up has them in parallel, and vice versa. That's De Morgan's law made of transistors. And since an nMOS switch closes on a 1, the pull-down conducts when some function f of the inputs is 1, which pulls the output to 0. A single CMOS gate therefore always computes an inverted function, ¬f. It can compute NOT, NAND and NOR, but not AND or OR.

The inverter

The smallest gate has one transistor in each network:

ApMOS (supply → out)nMOS (out → ground)OUT
0onoff1
1offon0

Two transistors. When the input is 0, the pMOS switch connects the output to the supply; when it's 1, the nMOS switch connects it to ground.

NAND: series down, parallel up

For a 2-input NAND, the output must be 0 only when A and B are both 1. Two nMOS transistors in series between the output and ground do exactly that: the path conducts only if both switches are closed. The pull-up is the dual: two pMOS transistors in parallel between the supply and the output, so that if either input is 0, its pMOS switch pulls the output to 1.

ABnMOS AnMOS Bpull-down (series)pMOS ApMOS Bpull-up (parallel)OUT
00offoffopenononclosed1
01offonopenonoffclosed1
10onoffopenoffonclosed1
11ononclosedoffoffopen0

Four transistors. In every row, exactly one of the two networks is closed.

NOR: parallel down, series up

NOR is the mirror image. The output must be 0 if either input is 1, so the pull-down is two nMOS transistors in parallel. The pull-up is its dual, two pMOS transistors in series: the output is connected to the supply only when both inputs are 0.

ABpull-down (parallel)pull-up (series)OUT
00openclosed1
01closedopen0
10closedopen0
11closedopen0

Four transistors again. The pattern extends to more inputs: a 3-input NAND has three nMOS in series and three pMOS in parallel, 6 transistors in all. Long series chains get slow, because every switch in the chain adds resistance, so real gate libraries rarely go beyond four inputs; wider functions are built as trees of smaller gates.

Why NAND is preferred over NOR

NAND and NOR cost the same number of transistors, but chip designers prefer NAND. The reason is in the transistors themselves: at the same size, a pMOS transistor conducts less current than an nMOS transistor, typically by a factor of about two, because the charge carriers in its channel are less mobile (semiconductors and doping, at the physics level, explains why). To make a pMOS switch as strong as an nMOS one, it's drawn wider.

A NOR puts the weak pMOS transistors in series, which makes the pull-up slower still, so they must be made even wider. A NAND puts them in parallel and the strong nMOS transistors in series. For the same speed, a NAND is smaller.

AND and OR cost more

Since one CMOS stage always inverts, AND must be a NAND followed by an inverter, and OR a NOR followed by an inverter:

GateBuilt asTransistors
NOT1 pMOS + 1 nMOS2
NAND (2 inputs)series nMOS, parallel pMOS4
NOR (2 inputs)parallel nMOS, series pMOS4
AND (2 inputs)NAND + NOT6
OR (2 inputs)NOR + NOT6
SRAM bittwo cross-coupled inverters + 2 access transistors6
full adderclassic "mirror adder"28

The book counts 2 transistors for NAND and NOR and 3 for AND and OR, because its bipolar gates use a resistor as the pull-up. In CMOS the resistor becomes a pMOS network, which doubles the count but eliminates the current the resistor would draw whenever the output is low.

The universality of NAND, shown at the logic level, still holds: you can build everything from NANDs. But it's not how chips minimize transistors. Here are NOT, AND and OR built from NAND gates only:

Logic · NOT, AND and OR from NAND gates only

Try it: Click an input switch in the circuit (or its button above) to toggle it — the gates and the truth table follow.

auto
0
gate delays
stable after 0
stable
state
2
critical path
gate delays, worst case
6
gates
ABNAND gate: output 11NOT ANAND gate: output 1NAND gate: output 00A AND BNAND gate: output 1NAND gate: output 1NAND gate: output 00A OR B
1 0 inputs changed, output switches next delayclick a switch to toggle it
ABNOT AA AND BA OR B
00100
01101
10001
11011

NAND is universal: tie both inputs together and it's NOT; follow it with a NOT and it's AND; feed it two inverted inputs and it's OR (De Morgan). So any circuit can be built from NAND gates alone — and NOR works the same way.

Unit-delay model: every gate takes one step to react. With auto off, toggle switches and press step to watch the change travel gate by gate.

This circuit uses six NAND gates, which is 24 transistors in CMOS. Built directly — an inverter (2), an AND (6) and an OR (6) — the same three outputs take 14. Real chips are built from standard cell libraries containing hundreds of gates — NAND, NOR, inverters, but also compound gates such as AND-OR-INVERT, which compute ¬(AB + C) in a single stage of 6 transistors — and synthesis tools pick whatever cells make the circuit smallest and fastest.

Where the power goes

A CMOS gate that isn't switching has no path from the supply to ground, so in the ideal case it draws no current at all. Power is spent in two ways.

Dynamic power is spent every time an output changes. Each gate's output drives a small capacitance: the gates of the transistors it feeds, plus the wire. Raising the output to 1 charges that capacitance from the supply; lowering it to 0 dumps the charge to ground. Each full charge–discharge cycle takes an energy of C·V² from the supply, half of it lost as heat in the pull-up and half in the pull-down. Over a whole chip:

dynamic power = α · C · V² · f

where C is the total switched capacitance, V the supply voltage, f the clock frequency, and α the activity factor, the fraction of the capacitance that actually switches in a cycle — often well below 0.5, since most signals don't change every cycle.

The V² term is the one that matters. Halving the frequency halves the power but doubles the time a task takes, so the energy for the task stays the same. Lowering the voltage by 30% cuts the energy per operation to 0.7² = 49% — but transistors switch more slowly at lower voltage, so the frequency must drop too. That's why processors run dynamic voltage and frequency scaling (DVFS): under light load, they lower the clock and the voltage together, and the energy per instruction falls much faster than the speed. It's also why phones and laptops use many cores at moderate speed rather than one very fast core.

A small extra short-circuit power is spent during each transition: while the input passes through the middle voltage, both networks conduct briefly.

Static power is spent even when nothing switches, because a transistor that is "off" isn't perfectly off. A small leakage current flows through the channel and through the ultra-thin gate insulator. For one transistor it's tiny; for billions, it's a large share of the total. Leakage grows as transistors shrink and as the threshold voltage is lowered to keep them fast. It's what ended Dennard scaling, as the previous chapter described. Chips fight it with power gating: switching off the supply to whole blocks — an idle core, an unused GPU cluster — so they leak nothing. The physics level's chapter on the limits of computing covers the physics of leakage and heat.

Fan-out and propagation delay

A gate's output doesn't change instantly. It has to charge or discharge the capacitance it drives through the resistance of its transistors, and that takes time: the propagation delay, the time from an input change to the output change. The more capacitance a gate drives, the longer it takes.

The number of gate inputs one output drives is its fan-out. Each input adds capacitance, so a gate with a large fan-out is slow. Designers handle this in several ways: they make the driving gate's transistors wider (stronger), or insert buffers — pairs of inverters — to split the load into a tree. A standard benchmark for a process is the FO4 delay: the delay of an inverter driving four copies of itself. Circuit speeds are often expressed in FO4s, which makes designs comparable across processes.

The logic level uses a simple model in which every gate has the same delay. The ripple-carry adder from the adders chapter, for example, has a longest path of 9 gates from input to output; in the simulator, adding 1 to 1111 takes 8 gate delays to settle. At the device level, the delay depends on transistor sizes, fan-out and wire length, and timing tools compute it for every path on the chip. The slowest path through a pipeline stage — the critical path — sets the maximum clock frequency.

Measuring delay with a ring oscillator

How do you measure a delay of a few picoseconds? You make the gate measure itself. Connect an odd number of inverting gates in a loop: there's no stable state, so the signal chases itself around the loop forever:

Logic · A ring oscillator: three inverting gates in a loop

Try it: Click an input switch in the circuit (or its button above) to toggle it — the gates and the truth table follow.

auto
0
gate delays
stable after 0
stable
state
loop
critical path
feedback: sequential
3
gates
ENNAND gate: output 1NOT gate: output 0NOT gate: output 11OUT
1 0 inputs changed, output switches next delayclick a switch to toggle it
ENOUT
0stable
1oscillates, period 6 delays

An odd number of inverting gates in a loop has no stable state. With EN = 1 the signal chases its own tail and the output toggles every 3 gate delays: the simulator detects the repeating state instead of settling. Real chips use this to measure gate speed.

Unit-delay model: every gate takes one step to react. With auto off, toggle switches and press step to watch the change travel gate by gate.

Set EN to 1 and the output toggles every 3 gate delays, for a period of 6: in the simulator, the output reads 1, 1, 1, 0, 0, 0 and repeats. In general, a ring of N stages oscillates with a period of 2·N·d, where d is the propagation delay of one stage: the signal must go around the loop twice, once as a rising edge and once as a falling edge, to return to its starting state. So measuring the frequency of a ring oscillator gives the gate delay directly: d = 1 / (2·N·f).

A ring of hundreds of stages oscillates at a frequency that's easy to count with ordinary instruments, and it divides out to a per-gate delay far too short to measure directly. Chip factories put ring oscillators on every wafer, often in the gaps between the dies, and measure them to check that each batch of transistors is as fast as it should be. The flip-flops chapter met the same circuit as a clock source.

Takeaways

  • A static CMOS gate has an nMOS pull-down network and a dual pMOS pull-up network; for every input, exactly one conducts. Series in one network means parallel in the other.
  • One CMOS stage always inverts: NOT (2 transistors), NAND and NOR (4) are natural; AND and OR need an extra inverter (6). The book's 2-and-3 counts are for resistor-based bipolar gates.
  • NAND is preferred over NOR because it puts the weaker pMOS transistors in parallel.
  • Building everything from NANDs is universal but not economical: NOT + AND + OR take 24 transistors from NANDs, 14 directly.
  • Dynamic power is α·C·V²·f; voltage counts squared, hence DVFS. Static power is leakage, fought with power gating.
  • Propagation delay grows with the capacitance a gate drives, i.e. with its fan-out. A ring oscillator of N stages has a period of 2·N·d, which is how gate delays are measured.

In this level

  1. 8.1Transistors as switches
  2. 8.2Building gates from CMOS
  3. 8.3From sand to chips
  4. 8.4Hard disks, SSDs and RAID
  5. 8.5Optical discs and tape
  6. 8.6Keyboards, displays, printers and cameras
  7. 8.7From modems to Ethernet: sending bits over a wire