The previous chapter wired transistors into gates. A modern chip has billions of them, and nobody places them one at a time. They're all made together, in layers, by printing patterns onto a disc of silicon — hundreds of chips per disc — the way a photograph is printed from a negative. This chapter follows that process from raw material to finished package, and then does the arithmetic that decides what a chip costs.
Silicon, very pure
Silicon is the second most common element in the Earth's crust, found in sand and quartz as silicon dioxide. Turning it into chip material takes two purification steps:
- Quartz is reduced with carbon in an electric furnace to metallurgical-grade silicon, about 98–99% pure — good enough for steel and aluminium alloys, far too dirty for chips.
- That silicon is converted into a gas (trichlorosilane), distilled, and decomposed back onto heated silicon rods. The result is electronic-grade polysilicon, typically 99.9999999% pure or better — "nine nines" — so that fewer than one atom in a billion is something other than silicon.
Purity matters because a chip's behavior is set by deliberately added impurities, the dopants that make regions of silicon n-type or p-type (semiconductors and doping, at the physics level, explains what they do). Stray atoms would add dopants nobody asked for.
Crystals and wafers
A transistor needs a perfect single crystal, with every atom in its place in the lattice. It's grown by the Czochralski method: the polysilicon is melted in a quartz crucible, a small seed crystal is dipped into the melt and slowly pulled up while rotating, and the silicon freezes onto it in the same crystal orientation. The result is a cylindrical ingot up to about two meters long.
The ingot is sliced with a wire saw into wafers, which are then polished to a mirror finish, flat to within nanometers. Leading-edge fabs use wafers 300 mm in diameter and 0.775 mm thick. The industry planned a move to 450 mm in the 2010s and abandoned it: re-equipping every fab was too expensive for the gain.
Printing with light: photolithography
A chip is built on the wafer in layers: first the transistors, in the silicon itself, then more than ten layers of copper wiring above them, separated by insulators. Each layer is defined by photolithography, repeated for every pattern:
- Coat the wafer with a light-sensitive photoresist.
- Expose it: a machine called a scanner shines ultraviolet light through a mask (or reticle), a quartz plate carrying the pattern of one layer, and a lens system shrinks the image four times onto the wafer. The scanner steps across the wafer, exposing one field at a time.
- Develop: the exposed resist (or the unexposed, depending on its type) dissolves, leaving the pattern in resist.
- Process the uncovered areas: etch away material, implant dopant ions, or deposit new material.
- Strip the remaining resist, clean, and start the next layer.
A leading-edge chip goes through dozens of masks and many hundreds of process steps, which takes months from bare wafer to finished dies. It all happens in cleanrooms, where the air holds far fewer particles than an operating theatre, because one dust particle landing on a wafer can ruin a die.
DUV and EUV
The smallest feature a lens can print is roughly proportional to the wavelength of the light divided by the lens's numerical aperture (how widely it collects light). So the industry has moved to ever shorter wavelengths:
| Light source | Wavelength | Notes |
|---|---|---|
| mercury lamps | 436, 365 nm | 1980s–1990s |
| DUV (deep ultraviolet), KrF laser | 248 nm | 1990s |
| DUV, ArF laser | 193 nm | since ~2000; "immersion" versions put water between lens and wafer to raise the numerical aperture |
| EUV (extreme ultraviolet) | 13.5 nm | in production since ~2019 |
With 193 nm light, fabs printed features far smaller than the wavelength by tricks such as multi-patterning: splitting one dense pattern over two to four separate masks and exposures. It works, but every extra exposure costs time, money and yield.
EUV light is absorbed by almost everything, including air and glass. So an EUV scanner works in a vacuum, uses mirrors instead of lenses, and makes its light by hitting tiny droplets of molten tin with a powerful laser, tens of thousands of times a second, turning each droplet into a glowing plasma. A single Dutch company, ASML, makes all of the world's EUV scanners; each costs well over $100 million. The next generation, High-NA EUV, raises the numerical aperture from 0.33 to 0.55 to print still finer lines.
What "3 nm" means: nothing you can measure
Chip generations are called process nodes and named by a length: 90 nm, 45 nm, 7 nm, 5 nm, 3 nm. Tanenbaum describes the 45 nm and 32 nm Core i7 generations by their line width, "how wide the wires between transistors are". That was roughly true for decades: the node name tracked the length of the transistor's gate.
From the late 1990s on, node names and gate lengths drifted apart, and by now the names are pure marketing labels. No feature on a "3 nm" chip is 3 nm wide. The distance between neighbouring transistor gates, the gate pitch, is still tens of nanometers, and the tightest metal wires are still a couple of dozen nanometers apart. What a new node name really promises is more transistors per area, and better speed or power, than the previous node of the same company. Different companies' nodes aren't directly comparable either: Intel renamed its "10 nm" process "Intel 7" in 2021 because it was roughly as dense as competitors' 7 nm processes.
Density now also improves by changing the shape of the transistor. Flat transistors gave way to FinFETs, where the channel is a vertical fin with the gate wrapped around three sides (Intel's 22 nm process, announced in 2011), and then to gate-all-around transistors, where the gate surrounds horizontal stacked sheets of silicon on all four sides — Samsung's 3 nm process first, then TSMC's N2 and Intel's 18A. Better control of the channel means less leakage, the enemy described in the previous chapter.
Yield: why big chips are expensive
A finished wafer is covered with identical rectangles, the dies. Some have defects: a particle, a scratch, a misprinted line. A die with a defect in the wrong place doesn't work. The fraction of good dies is the yield, and it depends heavily on die area.
The simplest model assumes defects land randomly, at an average density D per unit area. Then the probability that a die of area A has zero defects is given by the Poisson distribution:
yield = e^(−D·A)
Take a defect density of 0.1 per cm² (a plausible figure for a mature process; real numbers are closely guarded) and a 300 mm wafer. The number of dies that fit is approximately π·r²/A minus the partial dies lost around the edge, π·d/√(2A):
| Die area | Dies per wafer | Yield | Good dies |
|---|---|---|---|
| 100 mm² | 640 | 90.5% | 579 |
| 150 mm² | 416 | 86.1% | 358 |
| 600 mm² | 90 | 54.9% | 49 |
A die six times larger gives seven times fewer candidates, and only half of them work: nearly 12 times fewer good dies. Since the wafer costs the same either way, each good large die costs nearly twelve times as much as a small one, for six times the silicon.
This corrects a claim in the book. Discussing putting two CPUs on one chip, Tanenbaum says doubling the chip area halves the number of chips per wafer, which "essentially doubles the unit manufacturing cost". With yield taken into account, it more than doubles it. In the table, doubling the area from 100 to 200 mm² would drop the yield from 90.5% to 81.9%, and the cost per good die would rise about 2.3 times.
There's also a hard ceiling. A scanner exposes at most one field of 26 × 33 mm, about 858 mm², so no ordinary die can be larger. The book says large dies can reach 18 × 18 mm (324 mm²); today's biggest GPUs are over 800 mm², right at that limit.
Chipmakers soften the yield problem in two ways. Designs include redundancy: spare rows in memory arrays, which are switched in to replace defective ones. And dies with a defect in one unit are binned: that unit is disabled and the chip is sold as a cheaper model. The M2 Ultra this chapter was written on reports a 60-core GPU; the same chip is also sold with 76 GPU cores. Manufacturers don't say which lower configurations come from partly defective dies, but disabling units is a standard way to sell them.
Testing and packaging
Before the wafer is cut, a probe card touches the pads of every die and runs electrical tests, and bad dies are marked. The wafer is thinned, cut apart (diced), and the good dies are packaged. The package connects the die's microscopic pads to connections large enough to solder to a circuit board, spreads the heat, and protects the silicon. What all those connections carry — power, clocks, memory and I/O — is the subject of the chapter on CPU chips and pins.
Tanenbaum describes DIP, PGA and LGA packages, which plug into sockets. DIPs are now used only for simple parts and hobby kits. Desktop processors use LGA sockets — AMD moved its desktop processors from PGA to LGA in 2022 — but most chips today, in phones, laptops and embedded boards, are BGA (ball grid array) packages: an array of solder balls underneath the package, soldered directly to the board, with no socket. Inside, the die itself is usually mounted flip-chip: upside down, its pads connected to the package through tiny solder bumps, instead of with fine bond wires around the edge.
Chiplets
The yield table suggests an idea: instead of one large die, make several small ones and connect them inside the package. Four 150 mm² dies instead of one 600 mm² die give 358 good pieces per wafer, enough for 89 complete sets, against 49 good large dies — almost twice as many products per wafer, before paying for the more complex package.
That's the chiplet approach, and it's now common:
- AMD's Ryzen and EPYC processors, since 2019, combine small dies of CPU cores with a separate I/O die, which can even be made on an older, cheaper process.
- The M2 Ultra is two M2 Max dies joined by Apple's UltraFusion connection, a silicon bridge carrying thousands of signals between them. To software it looks like one chip with 24 cores.
- GPUs for AI sit next to stacks of HBM memory: DRAM dies stacked vertically and connected through holes in the silicon, placed beside the processor on a silicon interposer that carries thousands of short connections — TSMC's CoWoS packaging is the best-known example.
The idea isn't entirely new. Tanenbaum mentions the Pentium Pro of 1995, whose second-level cache was a separate die "in the same cavity" of the package. What's new is the density of the connections, which lets several dies behave almost like one.
Who can still do it
A leading-edge fab costs tens of billions of dollars, and each new node costs more to develop than the last. Companies that once made their own chips have left the race one after another, so the industry has split in two:
- Foundries manufacture chips designed by others. TSMC, in Taiwan, is the largest by far and makes the chips of Apple, AMD, NVIDIA, Qualcomm and many others — including the M2 Ultra, on what Apple calls a "second-generation 5-nanometer" process. Samsung is the other foundry at the leading edge.
- Intel has long designed and manufactured its own chips, and now also offers manufacturing to others as a foundry.
Most chip companies — Apple, AMD, NVIDIA — are fabless: they design chips and have foundries make them. And behind all of them sits a supply chain with its own bottlenecks, such as ASML's EUV scanners.
Takeaways
- Silicon is purified to about nine nines, grown into single-crystal ingots by the Czochralski method, and sliced into 300 mm wafers.
- Each layer of a chip is defined by photolithography: resist, exposure through a mask, development, then etching, implanting or deposition. Leading-edge chips need dozens of masks and months of processing.
- DUV (193 nm) needs multi-patterning for today's features; EUV (13.5 nm) prints them directly, with mirrors in a vacuum, from machines only ASML makes.
- Node names like "3 nm" are marketing labels, not dimensions — unlike the book's "line width". Transistors moved from flat to FinFET to gate-all-around.
- Yield falls exponentially with die area (e^(−D·A)): at 0.1 defects/cm², a 600 mm² die yields nearly 12 times fewer good dies than a 100 mm² one. Doubling area more than doubles cost, contrary to the book.
- Dies are tested, diced and packaged — mostly in BGA, flip-chip today — and increasingly combined as chiplets, like the two dies of the M2 Ultra.