Storage devices talk to the computer; input/output devices talk to people. A keyboard turns finger presses into bits, a display turns bits into light, a camera turns light into bits, a printer turns bits into ink. Each of them is a small computer system of its own, with sensors or emitters, a controller, and a protocol to the host. This chapter looks at how the common ones work at the device level. How the operating system drives them is covered by the chapter on files, devices and I/O, and the buses they plug into by the chapter on PCI Express and USB.
Tanenbaum's section on I/O devices dates from the early 2010s, and some of its devices — CRTs, ball mice, the Wiimote and Kinect — have since disappeared. The principles mostly haven't.
Keyboards: a matrix of switches
Under each key is a switch: a rubber dome that pushes a conductive pad onto the circuit board, or, in "mechanical" keyboards, a separate spring-loaded switch per key. A keyboard with over a hundred keys doesn't give each switch its own wire to the controller. The switches sit at the crossings of a matrix of rows and columns: pressing a key connects its row to its column. An 8 × 16 matrix serves 128 keys with only 24 wires.
The keyboard's microcontroller scans the matrix continuously. It drives one row at a time and reads all the columns: a column that sees the driven row has a closed switch at that crossing. Scanning every row takes well under a millisecond, and it's repeated constantly.
Matrices have a flaw. If three keys at three corners of a rectangle are pressed — say row 1 column 1, row 1 column 2, and row 2 column 1 — current can flow from row 2 through those three switches into column 2, and the fourth corner looks pressed too: ghosting. It's the same ambiguity Tanenbaum describes for early touch screens. Cheap keyboards detect the pattern and ignore the extra key, which limits how many keys can be held at once; keyboards with a diode in series with every switch block the sneak path and support n-key rollover, any number of simultaneous keys.
Debouncing
A mechanical contact doesn't close cleanly. The metal bounces, and for a few milliseconds the switch connects and disconnects repeatedly before it settles. Read naively, one press looks like several. The controller must debounce: accept a new state only after it has been stable for a while.
Here is a debouncer running on a recorded sequence of samples, 1 meaning "contact closed". The key is pressed once, but the samples bounce on the way down and on the way up:
Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.
- /* 1 = contact closed. The key is pressed once, but the contacts bounce. */
- char raw[] = "0001011011111111111110100100000000";
- int main() {
- int stable = 0, count = 0, presses = 0, edges = 0;
- for (int t = 1; raw[t]; t++) {
- int now = raw[t] - '0';
- if (now != raw[t - 1] - '0') edges++; /* a naive reader counts edges */
- if (now == stable) { count = 0; continue; }
- if (++count == 5) { /* 5 equal samples in a row: believe it */
- stable = now;
- count = 0;
- if (stable) presses++;
- }
- }
- printf("rising+falling edges: %d, debounced presses: %d\n", edges, presses);
- return presses;
- }
The raw signal has 10 edges — five apparent presses — but the debouncer reports 1 press. It changes its mind only after five consecutive samples disagree with the current state; the bounces never last that long. Real keyboards do the same with a few milliseconds of stable readings, sometimes with a separate timer per key.
From key to computer: USB HID
The book describes the PC keyboard of the PS/2 era: every press and every release raises an interrupt, and the handler reads the key's number from a register in the keyboard controller. Today's keyboards are USB or Bluetooth devices — the one on the Mac this chapter was written on is a Bluetooth keyboard — and they follow the HID (human interface device) class, which both buses share.
A USB keyboard doesn't interrupt anything. The host controller polls it at a fixed interval the keyboard requests, up to once per millisecond for a full-speed device, and the keyboard answers with a report describing its current state. In the standard 8-byte "boot" report:
| Byte | Contents |
|---|---|
| 0 | modifier keys, one bit each: left Ctrl, Shift, Alt, GUI (bits 0–3), right Ctrl, Shift, Alt, GUI (bits 4–7) |
| 1 | reserved |
| 2–7 | up to six keys currently held, as usage codes (a = 0x04, b = 0x05, …) |
Holding left Shift and M gives 02 00 10 00 00 00 00 00: bit 1 of the modifier byte is Shift, and 0x10 is M. There's no "press" or "release" message: the host compares each report to the previous one to find what changed. The six key slots limit the boot report to six keys at once besides modifiers; keyboards with n-key rollover use a longer report format. As Tanenbaum says, turning key codes into characters — Shift + M into "M", on a French or US layout — is entirely the software's job.
Mice work the same way. The book describes ball mice and a 3-byte serial protocol. A modern mouse has an optical sensor — a tiny camera that photographs the desk surface thousands of times per second and computes how the texture moved between images — and sends HID reports with its button states and the relative movement in x and y since the last report.
Touch screens
The touch screen of every phone and tablet is projected capacitive. Two layers of transparent electrodes, usually made of indium tin oxide, run in rows and columns, separated by a thin insulator; each crossing forms a small capacitor. A finger, which conducts and is coupled to ground, draws away part of the field at the crossings under it and changes their capacitance.
The controller measures the mutual capacitance at each crossing individually: it drives one row at a time and measures the charge coupled into every column — the same scanning idea as a keyboard matrix. Because each crossing is measured separately, two fingers produce two distinct spots, and the ghosting that plagued resistive and infrared screens disappears. The controller turns the grid of measurements into coordinates many times per second; recognizing taps, swipes and pinches is the operating system's job.
Tanenbaum notes that a pen or a gloved finger does nothing on a capacitive screen. That's still true of ordinary objects, but active styluses such as the Apple Pencil emit their own signal that the screen detects, and some gloves are made with conductive fingertips.
Displays: LCD and OLED
The cathode-ray tubes the book starts from are gone. Flat panels are made of two main technologies.
An LCD (liquid crystal display) doesn't produce light; it filters it. A backlight shines through a polarizer, a layer of liquid crystal, color filters and a second polarizer. An electric field across each subpixel changes how the liquid crystal twists the light's polarization, and so how much passes the second polarizer. Each subpixel has its own thin-film transistor that holds its voltage between refreshes: an active matrix, as Tanenbaum describes. The twisted-nematic (TN) design the book explains is now mostly found in cheap panels; IPS and VA designs give better viewing angles and contrast. Since the backlight is always on behind the crystal, blacks are never completely black; mini-LED backlights divide it into thousands of zones that dim independently.
An OLED (organic light-emitting diode) display, which the book lists as "on the horizon", has become standard in phones, and is common in laptops, tablets, TVs and monitors. Each subpixel is its own light source. A black pixel is simply off, so contrast is essentially infinite, and there's no backlight to power. The main weaknesses: the organic materials age with use, blue fastest, so a static image shown for a very long time can leave a faint trace (burn-in), and very high brightness is harder than with a backlight.
Displays also refresh faster. The book gives 60–100 refreshes per second; phones and laptops now commonly run at up to 120 Hz, and gaming monitors much faster.
The book also says that "any color can be constructed" by mixing red, green and blue. That's not quite true. Three primaries can only produce colors inside the triangle they span on a chromaticity diagram, the display's gamut; very saturated colors outside it can't be shown. Standard gamuts such as sRGB and the wider Display P3 used by Apple devices say exactly which triangle a display covers.
The framebuffer
Whatever the panel technology, the image comes from a framebuffer: memory holding one value per pixel, which the display controller reads out, line by line, every refresh. Tanenbaum calls it video RAM, a dedicated memory on the graphics card. On Apple silicon there's no separate video memory: the framebuffers live in the same unified DRAM as everything else, and a display engine in the chip reads them out.
This Mac has two identical monitors. system_profiler SPDisplaysDataType reports each as Resolution: 5120 x 1440, @ 60.00Hz. With 4 bytes per pixel — 8 bits each for red, green and blue plus 8 unused or alpha bits, the usual layout:
| Quantity | Computation | Result |
|---|---|---|
| pixels per display | 5,120 × 1,440 | 7,372,800 |
| bytes per frame | × 4 | 29,491,200 bytes = 28.1 MiB |
| bytes per second, one display | × 60 | 1.77 GB/s |
| both displays | × 2 | 3.54 GB/s |
So just scanning the two screens out reads 3.5 GB/s of memory, continuously. The real traffic is higher: the window server composes the final image from every window's own buffers, and each buffer is usually double or triple buffered so that the display never shows a half-drawn frame. The book's example — 1920 × 1080 at 3 bytes per pixel, 6.2 MB per frame, 155 MB/s at 25 frames per second — shows how much displays have grown: one of these frames is almost five times larger, and it's refreshed more than twice as often.
The indexed color the book describes, with an 8-bit index into a 256-entry palette to save memory, has disappeared from computers. Memory is cheap enough to store every pixel's color directly, and HDR and wide-gamut displays use 10 or more bits per channel.
Cameras
A digital camera's sensor is an array of photosites, each collecting the charge freed by the light falling on it during the exposure (the photoelectric effect is physics-level material). An analog-to-digital converter turns each charge into a number, typically 10 to 14 bits in current sensors.
The book says most cameras use CCD sensors and only "some" use CMOS. That has reversed completely: practically every camera sensor now, from phones to professional cameras, is a CMOS image sensor, which can be made on processes similar to logic chips, puts the amplifiers and converters on the sensor itself, and reads out faster with less power. CCDs survive in some scientific instruments.
A photosite measures only brightness, not color. To see color, a Bayer filter (patented by Bryce Bayer at Kodak in 1976) covers the photosites with a mosaic of red, green and blue filters in 2 × 2 groups: one red, one blue and two green, because the eye is most sensitive to green. The camera then demosaics the image: for each photosite, it estimates the two missing colors from its neighbours.
Tanenbaum says that a camera claiming 6 million pixels is "lying", since four photosites make one pixel and the rest is interpolation, so it really has 1.5 million. That overstates it. Demosaicing produces one full-color pixel per photosite, and because brightness detail is measured at every photosite, the real resolution is much closer to the photosite count than to a quarter of it. Color detail is lower, which is why fine colored patterns can show artifacts.
The book describes JPEG compression as applying a "two-dimensional spatial Fourier transform". It's actually a discrete cosine transform, a close relative, applied to 8 × 8 blocks of pixels; the high-frequency coefficients are then quantized coarsely, which is where the loss comes from. Since 2017, iPhones store photos as HEIF by default, using the more efficient HEVC video codec's image coding. And phone cameras now rely on computational photography: they capture several frames in quick succession and merge them, reducing noise and extending dynamic range far beyond what one exposure could record.
Printers
Printing puts bits onto paper by very different means, as the book describes, and its account still holds broadly:
- A laser printer charges a photosensitive drum, discharges it with a scanned laser where the page should stay white or black (depending on the design), dusts it with charged toner that sticks to the right areas, transfers the toner to paper and fuses it with heat.
- An inkjet printer fires picoliter droplets from hundreds of nozzles, either by boiling a tiny bubble of ink with a resistor (thermal) or by flexing a piezoelectric crystal.
- Thermal printers, which darken heat-sensitive paper, print most receipts.
All of them are binary: a dot is printed or not. Grays and colors are produced by halftoning — patterns of dots whose density the eye averages — and color printing uses CMYK inks: cyan, magenta and yellow subtract colors from white light, and black (the K) is added because the three don't mix into a true black. Most printers accept page description languages such as PostScript, PDF or PCL and render the page themselves. The solid-ink and wax printers the book lists have become rare; dye-sublimation survives in photo printers.
Game controllers and motion
The book's two examples, Nintendo's Wiimote (2006) and Microsoft's Kinect, are both discontinued; Microsoft stopped making the Kinect in 2017. The Wiimote's 3-axis accelerometer, though, is now in every phone, watch and controller, usually paired with a gyroscope that measures rotation: both are MEMS devices, microscopic mechanical structures etched in silicon, whose movement changes a measurable capacitance — the same principle the book describes. Kinect's idea of projecting a pattern of infrared dots to measure depth survives in the face-recognition cameras of phones.
Takeaways
- A keyboard is a matrix of switches that its controller scans row by row; diodes prevent ghosting. Contacts bounce, so presses must be debounced: 10 raw edges became 1 press in the demo.
- Keyboards and mice today are USB or Bluetooth HID devices, polled by the host, sending reports (an 8-byte boot report: modifiers plus up to six usage codes) rather than raising an interrupt per key as in the book.
- Projected capacitive touch screens measure each row–column crossing, so multi-touch has no ghosting.
- LCDs filter a backlight through liquid crystal; OLEDs, "on the horizon" in the book, emit light per subpixel and are now standard in phones. RGB can't make every color: a display has a gamut.
- The framebuffer for one 5120 × 1440 display at 4 bytes per pixel is 28.1 MiB; scanning two of them at 60 Hz reads 3.54 GB/s — versus the book's 155 MB/s example. On Apple silicon it lives in unified memory, not separate video RAM.
- Cameras use CMOS, not CCD, sensors, behind a Bayer filter; demosaicing loses far less resolution than the book's "four photosites per pixel" suggests; JPEG uses a DCT, not a Fourier transform.