You double-click an icon, and a fraction of a second later a window appears. It feels like one action. Underneath, nearly every level of this site takes part: an input device raises an interrupt, a program decides what to start, the kernel builds a process, a loader stitches together dozens of files, the memory system pulls code in from storage one page at a time, and finally the CPU runs the program's first instruction.
This chapter follows that path from the top down. Each step is a door into a deeper chapter; here we only walk past them, in order, to see how they fit together.
One click, every level
| Level | What it does when you open an app |
|---|---|
| Applications | Finder (or Explorer, or your Linux desktop) sees a double-click on an icon |
| Operating system | a new process is created, and the program file is read |
| Assembly and ISA | the loader sets up registers and the stack and jumps to the entry point |
| Microarchitecture | the first instructions miss in every cache and wait for memory |
| Digital logic and devices | the storage controller and memory chips deliver the bytes |
A computer can be described as a stack of levels (the levels of abstraction chapter tells that story), each one a machine whose language is implemented by the level below, either by translation (a compiler turns the whole program into the lower language first) or by interpretation (a program at the lower level carries it out step by step). The classic six levels stop at the problem-oriented languages programmers write in. This site adds two on top: the algorithms those programs implement, and the applications people actually use. Opening an app is the moment all of them get involved at once.
Before looking at each step, follow the whole trip once. Press Play: the ladder on the left shows which level is at work, and the chip under the text shows what is being passed along at that moment.
Try it: Press Play to follow the action down through the levels, or use the arrows and the dots to go at your own pace.
- 0Apps
- 1Algorithms
- 2C
- 3Assembly
- 4OS
- 5Machine code
- 6µops & buses
- 7Logic gates
- 8Transistors
- 9Physics
The mouse button goes down
The mouse notices the press and sends a small report to the computer over USB or Bluetooth.
1. The click becomes an event
The mouse reports that a button went down. That report travels over USB or Bluetooth, an interrupt tells the kernel something arrived, a driver decodes it, and the window server (the process that owns the screen, WindowServer on macOS) works out which window was under the pointer and sends that application an event. Two clicks close enough in time and space become a double-click. The keypress chapter follows this path in detail for a key; a mouse button takes the same road.
Here the window under the pointer belongs to Finder, which is itself just an application. Its event loop receives the double-click, sees that the item is an application, and asks for it to be started.
2. Someone decides to start a program
Starting a program is always a request from one process to the kernel. Who makes it depends on the system:
- On Linux, a desktop shell or a terminal shell uses the classic UNIX pair:
forkcopies the current process, and the child callsexecto replace itself with the new program. The processes chapter walks through both. - On Windows, Explorer calls
CreateProcess, which creates the process and loads the program in one call. - On macOS, Finder asks the LaunchServices framework, which asks
launchd(process 1, the ancestor of everything) to start the app. That's why a Mac app's parent is not Finder. On this Mac, every running Chrome browser process had parent process ID 1.
POSIX also offers posix_spawn, which does fork-then-exec in one call and is what macOS prefers internally. Whatever the route, the kernel ends up with the same job: make a new process and put a program in it.
3. The kernel creates a process
The kernel allocates a process control block (an ID, an owner, a table of open files, a place to save registers) and an empty address space. Then it opens the program file and reads its header. On a Mac that's a Mach-O file; on Linux an ELF file; on Windows a PE file. The executable files chapter opens all three byte by byte. The header says which CPU the code is for, which pieces of the file go where in memory, with which permissions, and which dynamic linker to run first.
Apple's Calculator shows the first check. Its executable is a universal binary with two complete programs inside, one for Intel Macs and one for Apple silicon:
$ lipo -archs /System/Applications/Calculator.app/Contents/MacOS/Calculator
x86_64 arm64e
The kernel picks the arm64e slice on this machine. On Apple silicon it also refuses to run code that doesn't carry a valid code signature, and the first time you open an app downloaded from the internet, macOS checks that the developer's signature and Apple's notarization are valid before anything runs.
4. The loader maps the program and its libraries
Very little of an app's code is in its own file. Calculator's executable is 3.4 MB, and it names 44 libraries and frameworks it needs: AppKit, SwiftUI, Foundation, CoreGraphics and others:
$ otool -L /System/Applications/Calculator.app/Contents/MacOS/Calculator
/System/Library/Frameworks/AppKit.framework/Versions/C/AppKit
/System/Library/Frameworks/Foundation.framework/Versions/C/Foundation
/System/Library/Frameworks/SwiftUI.framework/Versions/A/SwiftUI
/usr/lib/libSystem.B.dylib
… (44 in all)
Each of those libraries needs others. The dynamic linker (dyld on macOS, ld-linux on Linux) finds them all, maps each into the address space, and fills in the tables through which the program calls them, as described in the assembler, linker and loader chapter. Even the smallest C program pulls in a crowd. This one links against a single library, libSystem, and asks how many images the loader actually brought in:
#include <stdio.h>
#include <mach-o/dyld.h>
int main(void) {
printf("%u images\n", _dyld_image_count());
}
45 images
Forty-five: the program, libSystem, and the 43 system libraries libSystem is made of and depends on, among them libsystem_c (the C library proper), libsystem_malloc (malloc and free) and libsystem_kernel (the small wrappers that execute the actual system calls). On Linux, the system calls chapter traced the same work for a "hello" program: 32 system calls, 31 of them spent by the loader finding, mapping and protecting the C library.
If the loader opened and parsed thousands of library files for every launch, starting programs would be slow. So macOS prelinks all its system libraries into one huge file, the dyld shared cache. On this Mac it holds 3,649 libraries in about 5.8 GB of files for the arm64e architecture. It's mapped at the same place in every process, so its pages are shared by all of them, which is also why ls /usr/lib/libSystem.B.dylib finds no such file: the library exists only inside the cache.
5. Nothing is read until it's needed
A program is often pictured as being loaded into memory and then executed. The more accurate picture is demand paging, and that's what every modern system does. Nothing is read up front. The loader maps files: it tells the kernel "these addresses correspond to that part of that file", and the kernel records the promise without reading a byte. The first time the program touches one of those addresses, the hardware finds no valid translation and raises a page fault, an exception in the sense of the traps chapter. The kernel then finds the page, maps it, and restarts the instruction. The whole mechanism is the subject of the virtual memory chapter.
There are two kinds of page fault, and macOS's /usr/bin/time -l counts both. A page reclaim (a minor fault) finds the page already in memory (in the file cache, or in the shared cache used by every other process) and only has to map it. A page fault proper (a major fault) has to read it from storage. Here's what four programs cost to run once, measured on the Apple M2 Ultra this chapter was written on (two runs each):
| Command | Instructions retired | Page reclaims | Page faults | Memory used |
|---|---|---|---|---|
/usr/bin/true | 9.5 million | 243 | 0 | 1.4 MB |
ls / | 11.1–11.4 million | 255 | 1–2 | 1.5 MB |
python3 -c pass | 192–194 million | ~1,725 | 41–43 | 16 MB |
node -e 0 | 537–540 million | ~3,150 | 74–127 | 48 MB |
true does nothing at all (its whole job is to exit with status 0), yet running it takes 9.5 million instructions: the kernel creating the process, dyld binding 45 images, libSystem initializing, and the exit. Nearly all its 16 KiB pages were already in memory. Python and Node.js have to start an interpreter, a memory manager and a standard library before they can run a single line of your code, and it shows.
6. main runs, and then waits
When everything is mapped, the loader jumps to the program's entry point. A small piece of runtime code sets up the language environment and calls main. For a command-line tool like ls, main does its work, calls exit, and the kernel tears the process down.
An app with a window does something else: after setting up, it enters an event loop and waits for the next event. Waiting costs nothing: the process is blocked in a system call, and the scheduler gives its core to someone else. That's why a computer can have so many programs open. While this chapter was being written, top reported:
Processes: 1231 total, 12 running, 1219 sleeping, 10197 threads
1,231 processes on 24 cores, and 99% of them asleep, waiting for an event.
What it costs
Timing the whole round trip (posix_spawn, run, exit, waitpid) from a small C program, 20 to 50 times each (the minimum is the undisturbed case; the mean includes interference from everything else the machine was doing):
| Program | Minimum | Mean |
|---|---|---|
/usr/bin/true | 1.4 ms | 2.8–3.3 ms |
ls / | 1.4–1.6 ms | 3.4–4.1 ms |
python3 -c pass | 24–25 ms | 31–37 ms |
node -e 0 | 56–66 ms | 70–75 ms |
A millisecond or two is the floor for any process on this machine; a language runtime adds tens of milliseconds. A large graphical app, with its window, fonts, GPU context and saved state to restore, takes longer again, and the first launch after a reboot is the slowest of all, because its pages are not in memory yet and many more of its faults are major ones.
That's the reason for the tricks you see everywhere: apps that stay running in the background, browsers that keep processes ready for the next tab, phones that freeze apps instead of quitting them. Starting a program is expensive, and not starting one is the cheapest optimization there is.
Takeaways
- Opening an app crosses every level: an input event, a request to the kernel, a new process, the loader, page faults, and finally the CPU running
main. - Programs are started by other programs:
fork+execon UNIX,CreateProcesson Windows,posix_spawnorlaunchdon macOS, where every app's parent is process 1. - The executable is mostly a list of libraries: Calculator names 44, and even a trivial C program ends up with 45 images, most from macOS's prelinked dyld shared cache.
- Loading is lazy: files are mapped, and page faults bring pages in on first touch: as minor faults if they're already in memory, as major faults if they must be read from storage.
- Doing nothing costs 9.5 million instructions and 1.4 ms; starting Python or Node.js costs tens of milliseconds. Keeping programs running is cheaper than starting them.