Skip to content

Level 9 · Chapter 9.6

The limits: heat, tunneling, Landauer and quantum computing

Where physics pushes back: the power wall and power density, quantum tunnelling through gate oxides a few atoms thick, Landauer's minimum energy to erase a bit compared with a real CMOS switch, reversible computing, the speed of light across a chip, and what quantum computers can and cannot speed up.

For fifty years, computers got faster by making transistors smaller. Each step of this level has shown something that resists: wires whose delay grows with the square of their length, a subthreshold swing that can't beat 60 mV per decade, bits that leak away. This last chapter collects the hard limits — heat, quantum tunnelling, the thermodynamics of erasing information, the speed of light — and ends with quantum computing, which is often presented as the way around all of them. It isn't, but it does change the rules for a few specific problems.

Where the power goes

A CMOS chip spends energy in two ways.

Dynamic power comes from switching. As the first chapter showed, charging a capacitance C to voltage V and discharging it again turns CV² into heat. A chip with total switchable capacitance C, running at frequency f, where a fraction α of the nodes switch each cycle (the activity factor), dissipates:

P_dynamic = α × C × V² × f

Static power is leakage: the subthreshold current of the MOSFET chapter flowing through billions of transistors that are supposed to be off, plus tunnelling through their gates, which we'll get to. It's paid whether the chip computes or not, and it grows quickly with temperature.

To get a feel for the numbers, take an illustrative chip — these are assumptions, not a specific product: a billion gates, each driving 1 fF of gate and wire capacitance, 10 % switching per cycle, 0.7 V, 3 GHz:

print(1e9 * 0.1 * 1e-15 * 0.7**2 * 3e9, "W")   # 147 W

About 150 W: the power of a high-end desktop processor, with the right order of magnitude from four plausible numbers. Every one of those watts leaves the chip as heat.

The formula also explains why chips lower their voltage and clock together when they don't need full speed (dynamic voltage and frequency scaling). Near the top of the range, the maximum frequency a transistor supports is roughly proportional to the voltage, so power grows roughly as f³: running 20 % slower lets the voltage drop about 20 % as well, and power falls to about 0.8³ ≈ 51 %. That's why laptop and phone processors spend most of their time well below their peak clock.

Power density and the power wall

The problem isn't total power so much as power density: the heat has to leave through a die of a few square centimetres. Some comparisons, with stated assumptions:

import math
sigma = 5.670374419e-8                       # Stefan–Boltzmann constant
print(2000 / (math.pi * 10**2), "W/cm²")     # 2 kW cooking plate, 20 cm across: 6.4
print(100 / 1, "W/cm²")                       # 100 W on 1 cm² of die: 100
print(sigma * 5772**4 / 1e4, "W/cm²")         # radiated by the Sun's surface: 6,294

A processor dissipating 100 W on 1 cm² runs at more than ten times the power density of a kitchen hot plate. In a famous 2001 keynote, Intel's Pat Gelsinger extrapolated the trend of the time and warned that chip power density was heading toward that of a nuclear reactor, then a rocket nozzle, then the Sun's surface. It didn't happen, because the industry stopped following the trend.

As the MOSFET chapter explained, Dennard scaling kept power density constant while transistors shrank, as long as voltage shrank too. When voltage stopped falling around 2005, each new generation doubled the power per square millimetre at the same frequency. Chips hit the power wall: the limit of what air or liquid cooling can remove from a small area while keeping the silicon below about 100 °C. Tanenbaum describes the result in his first chapter: Intel's cancelled 4 GHz Pentium 4, and the turn to several cores per chip.

The consequence today is sometimes called dark silicon: a chip has more transistors than it can power at full speed at the same time. Designers spend them on things that are idle most of the time — large caches, specialized accelerators for video, encryption or neural networks, extra cores that run in turn — rather than on making one core faster.

Quantum tunnelling

In classical physics, a particle without enough energy to climb over a barrier stays on its side. In quantum mechanics, an electron is described by a wave, and the wave doesn't stop abruptly at the barrier: it decays exponentially inside it. If the barrier is thin enough, some of the wave comes out the other side, and the electron has a probability of being found there. That's tunnelling.

For a simple rectangular barrier of height φ and thickness d, the transmission probability is roughly:

T ≈ exp(−2κd), with κ = √(2 m* φ) / ħ

where m* is the electron's effective mass in the barrier and ħ is the reduced Planck constant. Take a gate oxide of SiO₂: a barrier of about 3.1 eV for electrons coming from silicon, and an effective mass of about 0.4 times the free electron mass (both approximate; this is an order-of-magnitude model):

hbar = 6.62607015e-34 / (2 * math.pi)
m0, e = 9.1093837015e-31, 1.602176634e-19
kappa = math.sqrt(2 * 0.4 * m0 * 3.1 * e) / hbar
print(kappa * 1e-9, "per nm")                    # 5.7
print(math.exp(2 * kappa * 0.2e-9))              # 9.8: x10 per 0.2 nm thinner
for d in (3, 2, 1.5, 1.2):
    print(d, "nm:", math.exp(-2 * kappa * d * 1e-9))
# 3 nm: 1.4e-15   2 nm: 1.2e-10   1.5 nm: 3.7e-08   1.2 nm: 1.1e-06

Every 0.2 nm of oxide removed — less than one atomic layer — multiplies the tunnelling current by about ten. Going from 3 nm to 1.2 nm multiplies it by nearly a billion. At 1.2 nm, the oxide in the mid-2000s was about five atomic layers thick, and gate leakage had become a significant part of the power budget. That's why oxides stopped thinning, and why the industry moved to high-k dielectrics that give the same capacitance with a physically thicker, tunnelling-proof layer.

Tunnelling is also used on purpose. Flash memory, as the chapter on storing a bit showed, writes and erases by forcing electrons through an oxide with a high field. And it sets a floor on channel length: when source and drain are only a few nanometres apart, electrons tunnel straight through the barrier under the gate, and the gate can no longer turn the transistor off.

Tanenbaum names the Heisenberg uncertainty principle as the quantum effect that may eventually trouble small transistors. The effect that actually arrived first, and shaped every transistor since the mid-2000s, is tunnelling. He also states Moore's law as a doubling every 18 months. Moore's 1965 paper predicted a doubling every year, and in 1975 he revised it to every two years; the 18-month figure is usually attributed to Intel's David House, who was talking about performance, not transistor count.

Landauer's limit

Is there a minimum energy for computing at all? In 1961 Rolf Landauer, at IBM, argued that there is — not for computing as such, but for erasing information. Resetting a bit that might be 0 or 1 to a known 0 reduces the number of possible states of the memory by half. The second law of thermodynamics doesn't allow that entropy to disappear, so it must be pushed into the environment as heat, at least:

E_min = kT ln 2

k = 1.380649e-23
E_L = k * 300 * math.log(2)
print(E_L, "J", E_L / e, "eV")                  # 2.87e-21 J = 0.0179 eV at 300 K

That's 2.87 × 10⁻²¹ J per erased bit at room temperature. The principle was confirmed experimentally in 2012, by measuring the heat released when erasing a one-bit memory made of a single glass bead held in optical traps.

How far is a real CMOS gate from it? Each transition of a gate output dissipates ½CV². Assume 0.1 to 1 fF of load at 0.7 V:

for C in (0.1e-15, 1e-15):
    E = 0.5 * C * 0.7**2
    print(C * 1e15, "fF:", E, "J,", E / E_L, "x Landauer")
# 0.1 fF: 2.45e-17 J, 8,534x    1 fF: 2.45e-16 J, 85,337x

A single logic transition costs roughly 10⁴ to 10⁵ times the Landauer limit, before counting the wires that carry the result any distance, the clock tree, or leakage. Landauer's limit is not what stops us today; CV², the 60 mV-per-decade floor on voltage and the wires are. But it's a real floor, and it has a surprising corollary.

Reversible computing

Landauer's limit applies only to operations that lose information. An AND gate does: from an output of 0 you can't tell which of three input combinations produced it. A NOT gate doesn't. In 1973, Charles Bennett showed that any computation can be done reversibly — with gates like the Toffoli gate, whose inputs can always be recovered from its outputs — by keeping the intermediate results and then running the computation backwards to "uncompute" them. In principle, a reversible computer has no minimum energy per operation.

Reversibility is a necessary condition, not a sufficient one: an ordinary CMOS gate still burns ½CV² because it charges its load abruptly through a resistance. The electronic counterpart is adiabatic charging: ramp the supply slowly over a time T much longer than the circuit's RC, and the energy lost becomes about (RC/T) × CV²:

R, C, V = 1e3, 1e-15, 1.0                    # assumptions: 1 kΩ, 1 fF, 1 V
print(0.5 * C * V**2)                         # conventional: 5e-16 J
print((R * C / 1e-9) * C * V**2)              # 1 ns ramp:    1e-18 J (500x less)

The price is speed — the energy saved is paid for in time — plus more transistors and a power supply that recovers the energy from each ramp. Reversible and adiabatic logic remain research topics, with small test chips, rather than a practical replacement for CMOS.

The speed of light

The signals chapter computed how far a signal goes in a nanosecond. At the clock rates of today's fastest cores, the numbers are sobering:

c = 299792458
f = 5e9
print(1 / f * 1e12, "ps")                       # 200 ps per cycle
print(c / f * 100, "cm in vacuum")              # 6.0 cm
print(c / math.sqrt(3) / f * 100, "cm in an insulator with εr = 3")  # 3.5 cm
rt = 2 * 0.10 / (0.5 * c)                       # to memory 10 cm away and back
print(rt * 1e9, "ns =", rt * f, "cycles")       # 1.33 ns = 6.7 cycles

Even light in a vacuum covers only 6 cm per cycle at 5 GHz; real on-chip wires, limited by RC delay, cover far less. A signal can't cross a large die in one cycle, and the round trip to memory 10 cm away costs almost 7 cycles before the DRAM itself does anything. Tanenbaum's version of the argument, with memory a foot away and 1 ns each way, concludes that faster computers must be smaller. That's what happened: memory moved onto the processor package (the stacked HBM of GPUs, the on-package memory of Apple's chips), chips are assembled from chiplets a few millimetres apart, and designs keep communication local — the root of the cache hierarchy of the caches chapter.

Quantum computing

Quantum computers are often described as the successor to all of the above. They aren't a faster version of ordinary computers; they're a different model of computation, dramatically better for a few problems and no better for most.

A qubit is a two-state quantum system — the spin of an electron, two energy levels of an ion, a tiny superconducting circuit. Unlike a bit, its state can be a superposition α|0⟩ + β|1⟩, where α and β are complex amplitudes with |α|² + |β|² = 1. Measuring it gives 0 with probability |α|² or 1 with probability |β|², and leaves it in that state.

n qubits have a state described by 2ⁿ amplitudes, one for each n-bit value. That's why classical simulation gets hard fast:

for n in (10, 50, 100):
    print(n, 2**n, 16 * 2**n, "bytes")          # 16 bytes per complex amplitude
# 10: 1,024 amplitudes (16 KB)   50: 1.1e15 (18 PB)   100: 1.3e30

Storing the state of 50 qubits takes 18 petabytes; 100 qubits, more than all storage ever built. But this is not 2ⁿ parallel computations you can read out. A measurement returns only n bits, sampled at random according to the amplitudes. A quantum algorithm has to arrange the amplitudes, through interference, so that wrong answers cancel and right ones reinforce before measuring. Only some problems have the structure that allows it.

What quantum computers are known to speed up:

  • Factoring and discrete logarithms (Shor's algorithm, 1994): exponentially faster than the best known classical algorithms. That breaks RSA and elliptic-curve cryptography, the public-key systems protecting today's internet traffic.
  • Unstructured search (Grover's algorithm): √N steps instead of N — a quadratic speed-up, not an exponential one. Against a 128-bit key it means about 2⁶⁴ sequential quantum steps, still enormous; the usual response is simply to use 256-bit symmetric keys.
  • Simulating quantum systems — molecules, materials — which was Feynman's original motivation in 1981, and probably the most useful near-term application.

What they're not expected to do: speed up ordinary programs, databases, web servers or games; solve NP-complete problems efficiently (no quantum algorithm is known to, and most researchers believe none exists); or process large data sets quickly, since loading the data is itself a bottleneck.

Why it's hard

Qubits must be isolated from their environment, because any interaction that reveals their state destroys the superposition (decoherence). For superconducting qubits, with transition frequencies around 5 GHz, that means operating at about 10–20 millikelvin, in a dilution refrigerator. The reason is the ratio of the qubit's energy quantum hf to the thermal energy kT:

h = 6.62607015e-34
for T in (0.01, 0.02, 300):
    print(T, "K:", h * 5e9 / (k * T))
# 0.01 K: 24   0.02 K: 12   300 K: 0.0008

At 20 mK the qubit's energy step is 12 times the thermal energy, so heat rarely flips it; at room temperature it would be drowned 1,000 times over.

Even so, physical qubits make errors far more often than transistors do, so useful machines need quantum error correction: encoding each reliable logical qubit in many physical ones and correcting errors continuously. In December 2024 Google reported, with its 105-qubit Willow chip, that making a surface-code logical qubit larger reduced its error rate — the threshold behaviour error correction depends on. The largest machines today have at most a few thousand physical qubits. A 2025 estimate by Google's Craig Gidney put factoring a 2048-bit RSA key at under a million noisy physical qubits running for under a week — far beyond current hardware, but close enough that cryptographers aren't waiting: NIST published its first post-quantum cryptography standards (FIPS 203, 204 and 205) in August 2024, and systems are migrating to them, since encrypted traffic recorded today could be decrypted later.

Tanenbaum mentions quantum computing as one of the "hopeful technologies" that might extend computing beyond silicon. A more accurate description today: a new tool for specific problems, built on the same physics — superconductors, lasers, semiconductors — that this level has been about, and running alongside classical computers, not replacing them.

Takeaways

  • CMOS power is dynamic (αCV²f) plus static leakage. Near peak, power grows roughly as f³: 20 % slower saves about half the power.
  • A processor's power density (~100 W/cm²) is over ten times a hot plate's. When voltage stopped scaling, chips hit the power wall; today's answer is multicore, accelerators and dark silicon.
  • Electrons tunnel through thin barriers: ~×10 more gate leakage for every 0.2 nm of SiO₂ removed. That ended oxide thinning around 1.2 nm and brought high-k dielectrics; it also makes flash possible.
  • Landauer: erasing a bit costs at least kT ln 2 = 2.87 × 10⁻²¹ J at 300 K. A CMOS transition (½CV² for 0.1–1 fF at 0.7 V) costs 10⁴–10⁵ times more. Reversible and adiabatic logic could in principle go below it, at the cost of speed.
  • At 5 GHz light travels 6 cm per cycle, and memory 10 cm away is ~7 cycles round trip: computers get faster by getting smaller and more local.
  • Quantum computers manipulate 2ⁿ amplitudes but read out only n bits. They speed up factoring (Shor), search quadratically (Grover) and quantum simulation — not everyday computing. Error correction is the hurdle, and post-quantum cryptography is already being deployed.

In this level

  1. 9.1Electrons, conductors and insulators
  2. 9.2Semiconductors, doping and the p-n junction
  3. 9.3How a MOSFET switches: the field effect
  4. 9.4Storing a bit: charge, magnetism and light
  5. 9.5Signals on wires and fibers
  6. 9.6The limits: heat, tunneling, Landauer and quantum computing