Every gate in a processor is built from switches, and the switch is almost always a MOSFET: a metal-oxide-semiconductor field-effect transistor. It's the most manufactured object in history — a single high-end chip has tens of billions of them, and the Apple M2 Ultra this site was written on has 134 billion. It has no moving parts and draws almost no current at its control input: it's switched by an electric field acting through an insulator, which is where the name comes from.
The device level above treats the MOSFET as an ideal switch and builds CMOS gates from it (see Transistors as switches and Building gates from CMOS). This chapter looks inside: how a voltage on one electrode turns a strip of silicon from insulator to conductor, why it never quite turns off, and how the limits of that physics ended thirty years of easy speed-ups.
The structure
Take the n-channel MOSFET (NMOS). It's built on p-type silicon, the doped silicon of the previous chapter, with three terminals on top:
- two heavily doped n-type regions, the source and the drain, a short distance apart;
- between them, on top of the silicon, a very thin insulating layer, the gate oxide;
- on top of the oxide, a conducting electrode, the gate. It was originally aluminium (the "metal" of the name), then doped polysilicon for decades, and is metal again since 2007.
The distance from source to drain under the gate is the channel length L; the width of the device across is W.
With no voltage on the gate, source and drain are n-type islands in a p-type sea. The path from one to the other crosses two p-n junctions back to back, and whichever way the voltage goes, one of them is reverse-biased. No current flows: the switch is off.
The gate is a capacitor
The gate, the oxide and the silicon beneath form a parallel-plate capacitor, as in the first chapter. Its capacitance per unit area is:
C_ox = ε_ox / t_ox
where tox is the oxide thickness and ε_ox = 3.9 ε₀ for SiO₂. Oxides in the 2000s got down to about 1.2 nm:
eps0 = 8.8541878128e-12
def cox(t, er=3.9):
return er * eps0 / t # F/m²
for t in (100e-9, 10e-9, 2e-9, 1.2e-9):
print(t * 1e9, "nm:", cox(t) * 100, "µF/cm²")
# 100 nm: 0.035 10 nm: 0.35 2 nm: 1.73 1.2 nm: 2.88 µF/cm²
Thinner oxide means more capacitance, which means more charge for the same gate voltage. That charge is what makes the channel, so thinner oxide means a stronger transistor. It also means a lot of engineering, as we'll see.
Inversion: making a channel
Now raise the gate voltage VGS (measured from the source) above zero:
- Depletion. The positive gate repels the holes, the p-type majority carriers, from the surface just under the oxide. What's left there are the fixed, negative boron ions: a depletion layer, as in a p-n junction. Still no conduction.
- Inversion. Raise the voltage further, and the surface starts attracting the few electrons that exist in p-type silicon (the minority carriers), plus electrons from the source and drain next door. Past a certain voltage, electrons outnumber holes at the surface: that thin layer, a few nanometres deep, has turned n-type. It's called an inversion layer, and it connects source to drain with a continuous n-type path — the channel. The switch is on.
The gate voltage at which the channel forms is the threshold voltage VT, a few hundred millivolts in modern logic transistors. Above threshold, the channel charge grows in proportion to the extra voltage, the overdrive VGS − VT:
Q_channel ≈ C_ox × (V_GS − V_T) per unit area
How many electrons is that? Take an oxide with the capacitance of 1 nm of SiO₂, 0.5 V of overdrive, and a small gate 20 nm × 20 nm:
e = 1.602176634e-19
Q = cox(1e-9) * 0.5 # C/m²
print(Q / e / 1e4) # 1.08e+13 electrons per cm²
print(Q * (20e-9 * 20e-9) / e) # 43 electrons
About 10¹³ electrons per cm² — and, under a 20 × 20 nm gate, only about 40 of them. A modern transistor switches a current carried by a few dozen electrons at a time.
The PMOS transistor is the mirror image: p-type source and drain in n-type silicon, a channel of holes, turned on by a gate voltage below the source's. CMOS logic pairs the two so that one is off whenever the other is on; that's the device level's subject.
On current
With the channel formed, a voltage between drain and source drives a current along it. Roughly:
- the current is proportional to W/L: a wider channel is more parallel lanes, a shorter one a shorter trip;
- it grows with the overdrive. For long channels, the classic model gives I ∝ (VGS − VT)². In today's short channels the electrons hit a maximum drift speed — about 10⁷ cm/s in silicon — and the dependence becomes closer to linear;
- it's proportional to the carrier mobility, which is why PMOS, with holes about three times less mobile than electrons, needs to be wider for the same current. Since the 90 nm generation (2003), chips deliberately strain the silicon lattice to raise mobility.
To switch a gate, this on current charges or discharges the capacitance of the next gates and wires. More current, less capacitance, and less voltage to swing: that's the whole of transistor speed.
Off isn't off: subthreshold leakage
Below the threshold, the channel doesn't vanish all at once. There's still a small energy barrier between source and channel, and a few electrons from the source are energetic enough to get over it. As with the diode, the number that do follows a Boltzmann factor, so the subthreshold current falls exponentially as the gate voltage drops:
I ∝ exp(q(V_GS − V_T) / nkT)
The gate controls the barrier only partly — through the oxide capacitance, in competition with the depletion capacitance underneath — and n ≥ 1 expresses that loss. The steepness is quoted as the subthreshold swing S: how many millivolts of gate voltage it takes to change the current by a factor of ten. Even with perfect gate control (n = 1):
S = (kT/q) × ln 10
import math
k = 1.380649e-23
for T in (77, 300, 358):
print(T, "K:", k * T / e * math.log(10) * 1000, "mV/decade")
# 77 K: 15.3 300 K: 59.5 358 K (85 °C): 71.0 mV/decade
At room temperature, no MOSFET can turn off faster than 59.5 mV per decade of current; real ones are somewhat worse, and a hot chip is worse still. This is not an engineering shortfall. It comes from the thermal energy distribution of the electrons themselves, and the only knob in the formula is temperature.
The limit has a brutal consequence. The current at VGS = 0 (the "off" current) is the current at threshold divided by 10 for each S millivolts of VT:
for VT in (0.2, 0.3, 0.4):
print(VT, [10 ** (-VT / S) for S in (0.060, 0.070)])
# 0.2 V: 4.6e-04, 1.4e-03 0.3 V: 1.0e-05, 5.2e-05 0.4 V: 2.2e-07, 1.9e-06
print(10 ** (0.1 / 0.07)) # 27: lowering VT by 100 mV at 70 mV/dec
With S = 70 mV/decade, each 100 mV cut from the threshold multiplies the off current by about 27. Multiply a small off current by tens of billions of transistors and it becomes a large fraction of a chip's power, spent while doing nothing. So VT can't go much below about 0.2–0.3 V. And the supply voltage has to stay well above VT for the transistor to turn on strongly. That pins the supply voltage near 1 V — and that, as the next section shows, ended an era.
Dennard scaling
In 1974, Robert Dennard and his colleagues at IBM published the recipe that drove the industry for thirty years. Shrink every dimension of a transistor by a factor κ — length, width, oxide thickness — and reduce the voltage by the same κ, and the electric fields inside stay the same (constant-field scaling). Then:
| Quantity | Scales as | For κ = 1.4 (one generation) |
|---|---|---|
| dimensions, voltage | 1/κ | × 0.71 |
| capacitance | 1/κ | × 0.71 |
| gate delay | 1/κ | × 0.71 (40 % faster) |
| transistors per area | κ² | × 1.96 |
| power per transistor (C V² f) | 1/κ² | × 0.51 |
| power per area | 1 | unchanged |
Each generation, roughly every two years, gave twice the transistors, each 40 % faster, for the same power per square millimetre. Moore's law provided the transistors; Dennard scaling made them free to run. Clock frequencies climbed from a few MHz to 3.8 GHz for the Pentium 4 in 2004.
It ended around 2005, because two things stopped shrinking:
- the threshold voltage, pinned by the 60 mV-per-decade limit above, and with it the supply voltage;
- the gate oxide, which at about 1.2 nm — a handful of atomic layers — leaked current straight through by quantum tunnelling, as the last chapter calculates.
Keep shrinking dimensions but not voltage, and the table changes. Capacitance still drops by 1/κ and frequency could rise by κ, but V² stays the same: power per transistor stays the same while density doubles, so power per area doubles every generation. Chips hit the limit of what a heat sink can remove — the power wall. Intel cancelled its 4 GHz Pentium 4 in 2004, and the industry turned to multiple cores per chip instead of faster ones. Two decades later the fastest desktop chips boost to about 6 GHz, where the Dennard trend alone would have taken them far beyond.
Tanenbaum describes this moment in his first chapter, but frames it as higher clock speeds "requiring" higher voltage. The deeper cause is the one above: the voltage stopped going down when everything else kept shrinking. The book also gives CMOS as running near 1.5 V. That's now the high end: modern cores run well under 1 V when efficient, and approach 1.4 V only at a desktop chip's highest boost clocks.
Keeping the gate in control
Shrinking the channel brought another problem. When source and drain get very close, the drain's electric field reaches into the channel and lowers the source barrier itself (drain-induced barrier lowering). The gate loses control, and the transistor leaks even when the gate says off. The cure is always the same: put more gate around less channel.
- High-k metal gate (2007). Replace SiO₂ with a hafnium-based oxide whose permittivity is about 20–25 instead of 3.9. The same capacitance as 1 nm of SiO₂ — an equivalent oxide thickness of 1 nm — then takes 5–6 nm of physical material, thick enough to block tunnelling. Intel introduced it at 45 nm, with a metal gate replacing polysilicon.
- FinFET (2011). Stand the channel up as a thin vertical fin and wrap the gate over its top and both sides. Three sides of control instead of one. Intel shipped it first at 22 nm (calling it tri-gate); the rest of the industry followed within a few years.
- Gate-all-around (2022–2025). Slice the channel into horizontal nanosheets stacked on top of each other, and surround each one with gate on all four sides. Samsung shipped it at its 3 nm node in 2022; TSMC's N2 and Intel's 18A (which calls it RibbonFET) followed in 2025.
| Planar (to ~2011) | FinFET (2011–) | Gate-all-around (2022–) | |
|---|---|---|---|
| channel | flat surface of the wafer | vertical fin | stacked horizontal sheets |
| gate covers | top | top and two sides | all around |
| main benefit | simple | much better off-state control | best control, adjustable sheet width |
Research continues with complementary FETs that stack the NMOS and PMOS on top of each other, and with transistors that could beat 60 mV/decade by using different physics — tunnelling instead of thermal emission over a barrier. None is in production.
Node names lost their literal meaning long ago. Up to the 1990s, a "0.35 µm" process had gates about 0.35 µm long. Today's "3 nm" and "2 nm" are generation labels: no feature on the chip measures 3 nm. What still roughly holds is the trend: each named step packs more transistors per square millimetre.
Takeaways
- A MOSFET is controlled by a field through the gate oxide, which with the silicon forms a capacitor: C_ox = ε/tox, 2.9 µF/cm² for 1.2 nm of SiO₂.
- A gate voltage above the threshold VT inverts the surface under the oxide into a conducting channel between source and drain. Under a 20 × 20 nm gate, that channel holds only a few dozen electrons.
- Below threshold, the current falls exponentially but never stops. The subthreshold swing can't be steeper than (kT/q) ln 10 = 59.5 mV per decade at 300 K: cutting VT by 100 mV multiplies leakage by roughly 25 to 50.
- Dennard scaling shrank dimensions and voltage together, giving more and faster transistors at constant power density. Around 2005, voltage and oxide thickness stopped scaling, power density rose, and clock speeds stalled: the power wall.
- Gate control was restored by high-k metal gates (2007), FinFETs (2011) and gate-all-around nanosheets (2022–2025). Node names like "3 nm" no longer measure anything.