Skip to content

Level 4 · Chapter 4.2

Processes and the address space

What a process is — an address space, a saved CPU state and a set of kernel resources — with a real Linux address space read from /proc/self/maps, address-space randomization, fork, exec and wait with copy-on-write, process states and scheduling, and what creating and switching processes costs on real machines.

A program is a file: the executable of the previous chapter, sitting on disk. A process is a program running: the operating system's unit of execution, with its own memory, its own saved CPU state, and its own view of the machine. Starting a program twice gives two processes running the same code, unaware of each other.

The book puts it in the terms of this site's levels: a process is characterized by its state and its address space, and its state includes at least the program counter, the flags, the stack pointer and the general registers. Everything the levels below provide — registers, memory, instructions — the OS gives each process a private copy of, or the illusion of one.

What a process is made of

  • An address space: the range of addresses the process can use, and what's at each of them. Each process has its own, from near 0 up to 128 TiB on x86-64 Linux (256 TiB on ARM64), and the same address means different memory in different processes. The chapter on virtual memory and paging shows how the hardware makes that work.
  • A CPU state: the program counter, stack pointer, flags and all the other registers. While the process runs, they're in the CPU; while it waits, the kernel keeps them in memory.
  • Kernel resources: a process ID, an owner and permissions, open files, current directory, signal handlers, limits. The kernel keeps all of this in a process control block — task_struct in Linux, several kilobytes of fields per process.

The process never sees the process control block. From its point of view, it has a CPU and a memory to itself.

The address space, for real

The address space isn't one block: it's a set of regions, each mapped from somewhere, with its own permissions. On Linux, /proc/self/maps lists them. Here's a small C program that prints where its code, globals, heap, a large malloc, a library function and a local variable are, followed by its own map, on 64-bit ARM Linux (lightly trimmed):

main (code)      0xaaaaad0b0914
initialized      0xaaaaad0d0060
zeroed           0xaaaaad0d0068
malloc(16)       0xaaaad7d2d2a0
malloc(1 MiB)    0xffff90a0f010
printf (libc)    0xffff90b5cfe0
local (stack)    0xfffff4aa3754

aaaaad0b0000-aaaaad0b1000 r-xp  /w/as                    ← code (.text)
aaaaad0cf000-aaaaad0d0000 r--p  /w/as                    ← read-only data
aaaaad0d0000-aaaaad0d1000 rw-p  /w/as                    ← .data and .bss
aaaad7d2d000-aaaad7d4e000 rw-p  [heap]                   ← small mallocs
ffff90a0f000-ffff90b10000 rw-p                           ← the 1 MiB malloc (its own mmap)
ffff90b10000-ffff90c9b000 r-xp  …/libc.so.6              ← the C library's code
ffff90cb0000-ffff90cb2000 rw-p  …/libc.so.6              ← …and its data
ffff90ccf000-ffff90cf6000 r-xp  …/ld-linux-aarch64.so.1  ← the dynamic loader
ffff90d0b000-ffff90d0d000 r-xp  [vdso]                   ← kernel code mapped into every process
fffff4a83000-fffff4aa4000 rw-p  [stack]

Each line is a range of addresses, its permissions (read, write, execute, private), and what it's mapped from. The executable's sections become separate regions: code is readable and executable but not writable, data writable but not executable — the W^X rule. The heap grows up from just above the program; libraries, large allocations and the stack sit near the top of the address space, and the stack grows down. The [vdso] region is the piece of kernel code that answers calls like clock_gettime without a system call, mentioned in the traps chapter. The kernel's own memory lives in the top half of the 64-bit address range, mapped in every process but inaccessible from user mode.

Run the same program again and every address changes:

main (code)      0xaaaaaf4f0914      (was 0xaaaaad0b0914)
malloc(16)       0xaaaab8ea92a0      (was 0xaaaad7d2d2a0)
printf (libc)    0xffffa6c9cfe0      (was 0xffff90b5cfe0)
local (stack)    0xffffdb996b64      (was 0xfffff4aa3754)

That's ASLR, address space layout randomization: the kernel places each region at a random offset every time a program starts, so an attacker exploiting a memory bug can't know where the code, the stack or the libraries are. Only the offsets within a region stay fixed: main always ends in 914.

The simulator shows the same regions for its own programs. Step through this one at the OS level and watch the address space fill in:

Live · One variable in each region

Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.

C source — click a line number for a breakpoint
  1. int initialized = 42;
  2. int zeroed;
  3. int main() {
  4. int local = 1;
  5. int *heap = malloc(sizeof(int));
  6. *heap = initialized + zeroed + local;
  7. return *heap;
  8. }
generated assembly— ▸ is the next instruction; it maps to the ▸ C line above
  1. int initialized = 42;
  2. int zeroed;
  3. int main() {
  4. int local = 1;
  5. int *heap = malloc(sizeof(int));
  6. *heap = initialized + zeroed + local;
  7. return *heap;
  8. }
step 0
Loading emulator…
The process as the operating system sees it: an address space of code, globals, heap and stack.

It returns 43. The simulator's layout is simpler than Linux's — no libraries, no randomization — but the regions are the same: code, data, heap growing up, stack growing down.

fork, exec and wait

UNIX creates processes in two steps, which is unusual but has lasted fifty years:

  • fork() makes an exact copy of the calling process: same code, same memory contents, same open files. It returns twice — in the parent, it returns the child's process ID; in the child, it returns 0. That's how each copy knows which one it is.
  • exec() replaces the calling process's program with a new one from an executable file: new code, new data, new stack, same process ID and open files.
  • wait() makes the parent sleep until a child exits, and collects its exit status.

When you type ls in a shell, the shell forks, the child execs /bin/ls, and the parent waits. Between the fork and the exec, the child can rearrange its open files — that's how ls > out.txt redirects output without ls knowing anything about it.

This program forks once. The child changes a global variable and exits with status 7:

int x = 1;

int main(void) {
    pid_t pid = fork();
    if (pid == 0) {
        x = 2;
        printf("child:  fork() returned 0, x = %d at %p\n", x, &x);
        exit(7);
    }
    int status;
    waitpid(pid, &status, 0);
    printf("parent: fork() returned %d, x = %d at %p, child exit status %d\n",
           pid, x, &x, WEXITSTATUS(status));
}

On Linux:

child:  fork() returned 0, x = 2 at 0xaaaabf030068
parent: fork() returned 14, x = 1 at 0xaaaabf030068, child exit status 7

The child set x to 2 at address 0xaaaabf030068, and the parent still reads 1 at the same address. Same virtual address, different memory: each process has its own address space.

Copying a whole address space on every fork would be slow, especially since the child usually calls exec right away and throws the copy out. So the kernel doesn't copy anything. Parent and child share every page, marked read-only, and a page is copied only when one of them writes to it — copy-on-write. Here, the child's x = 2 triggered a page fault, the kernel copied that one page, and the write went into the child's copy.

A process that has exited but hasn't been waited for yet is a zombie: its memory is gone, but the kernel keeps its process table entry so the parent can still read the exit status. When a parent dies first, its children are adopted by process 1, init (systemd on most Linux systems, launchd on macOS), which waits for them. Every process on the system descends from it, forming a process tree.

Windows has no fork: CreateProcess creates a new process and loads a program into it in a single call. POSIX added posix_spawn to do the same on UNIX, and macOS prefers it.

States and scheduling

A machine runs far more processes than it has cores: this Mac had 1,297 processes on 24 cores while this chapter was written. Nearly all of them are waiting at any moment. A process is in one of three main states:

  • running on a core;
  • ready: able to run, waiting for a core;
  • blocked (sleeping): waiting for something — a key press, a disk read, a network packet, a timer, a child to exit.

The scheduler picks which ready process runs on each core. A process leaves the CPU when it blocks, typically inside a system call, or when the timer interrupt says its time slice is used up, which is what makes multitasking preemptive: no process can hold a core forever. Each switch is a context switch: save the running process's registers into its control block, pick the next process, switch the address space, restore its registers, return to user mode. On Linux, the scheduler has been EEVDF since kernel 6.6; it gives each ready process a fair share of CPU time, weighted by priority.

What it costs

Measured on the M2 this chapter was written on, under macOS and in a Linux virtual machine:

macOSLinux VM
fork + exit in the child + wait0.7–1 msabout 0.12 ms
start /usr/bin/true with posix_spawn and wait1.3–2 msabout 0.22 ms
pass a byte back and forth between two processes through pipes (two context switches)7–28 µs24–36 µs

These numbers vary a lot from run to run, and from system to system: macOS's fork is notoriously slower than Linux's, and a virtual machine adds its own overhead to every switch. But the orders of magnitude are the point. A function call takes about a nanosecond, a system call about a hundred nanoseconds, a context switch microseconds, and creating a process from an executable hundreds of microseconds or more. That's why servers keep processes and threads around in pools instead of starting a new one per request, and why threads — several flows of execution sharing one address space, the subject of a later chapter — exist.

Takeaways

  • A process is a running program: an address space, a saved CPU state, and kernel resources, recorded in its process control block.
  • The address space is a set of regions with their own permissions: code (r-x), data (rw-), heap, memory-mapped files and libraries, stack. /proc/self/maps lists them.
  • ASLR places every region at a random address on each run.
  • UNIX creates processes with fork (copy the caller, return twice), exec (load a new program) and wait (collect the exit status). Copy-on-write makes fork cheap: pages are only copied when written.
  • Processes are running, ready or blocked; the scheduler and the timer interrupt share the cores between them, with a context switch each time.
  • Costs span orders of magnitude: a context switch takes microseconds, starting a program hundreds of microseconds to milliseconds.

In this level

  1. 4.1Executable files: ELF, PE and Mach-O
  2. 4.2Processes and the address space
  3. 4.3System calls and privilege levels
  4. 4.4Virtual memory and pagingPlanned
  5. 4.5Files, devices and I/OPlanned
  6. 4.6Threads and synchronizationPlanned
  7. 4.7Hardware virtualization and hypervisorsPlanned
  8. 4.8Inside UNIX and WindowsPlanned