Tanenbaum describes the operating system as a level of its own: the operating system machine. It offers everything the ISA offers, plus a set of new instructions — open a file, read from it, create a process, allocate memory, send a packet. Ordinary instructions still run directly on the hardware. The new ones, the system calls, are carried out by the kernel on the program's behalf. From the program's point of view, read is just one more instruction, a slow and very powerful one.
The traps chapter showed the hardware side: the syscall instruction, the switch to kernel mode, the cost. This chapter looks at the same boundary from the operating system's side.
Why programs can't just do it themselves
A program in user mode can compute anything, but it can't touch the machine. The instructions that control hardware — reading and writing I/O ports, loading page tables, disabling interrupts, halting the CPU, changing privilege — fault if user code tries them. The page tables mark the kernel's own memory as off-limits to user mode. So everything that touches a device, another process or the kernel's data has to be asked for.
That's the point: isolation. The kernel can check every request — does this process have permission to open that file? is that buffer inside its own memory? — and one program can't read another's memory, crash the disk driver, or keep the CPU forever. The same design runs on every modern OS, with the same two levels in practice: x86's rings 0 (kernel) and 3 (user), with rings 1 and 2 unused; ARM64's EL1 and EL0; RISC-V's S-mode and U-mode.
The protection works in the other direction too. Since the Meltdown and similar attacks of 2018, and even before, kernels don't want to touch user memory by accident: x86's SMEP and SMAP and ARM's PAN make the CPU fault if the kernel executes user code or reads user memory outside of the few routines designed for it. And after Meltdown, Linux on affected x86 chips even switches to separate page tables on every kernel entry (KPTI), so that user mode can't speculatively read kernel memory — one more reason system calls got slower.
Into the kernel and back
On x86-64 Linux, a system call goes like this:
- The C library function —
write,getpid, … — puts the call number inraxand up to six arguments inrdi,rsi,rdx,r10,r8,r9, and executessyscall. - The CPU saves the return address in
rcxand the flags inr11, switches to ring 0, and jumps to the kernel's entry point, which the kernel stored in a model-specific register at boot. - The kernel's entry code switches to the kernel stack, saves the user registers, and uses
raxas an index into its system call table, an array of function pointers. - The handler checks its arguments. Pointers get special care: a user program could pass the address of kernel memory, or of nothing at all, so the kernel never dereferences them directly. It copies data in and out with
copy_from_userandcopy_to_user, which verify the addresses and recover cleanly from page faults. - The result goes back in
rax, the user registers are restored, andsysretreturns to the instruction aftersyscall, in ring 3.
Errors come back as negative numbers: −2 is ENOENT (no such file), −9 is EBADF (bad file descriptor), −13 is EACCES (permission denied), −38 is ENOSYS (no such system call). The C library's wrapper turns them into the convention C programs see: it stores the positive code in the global errno and returns −1.
The simulator implements a tiny kernel with four of Linux's system calls — write, exit, getpid and getppid — and returns ENOSYS for the rest:
Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.
- .data
- msg: .asciz "hi\n"
- .text
- mov eax, 39 ; getpid()
- syscall
- mov ebx, eax ; keep the result: 1000
- mov eax, 1 ; write(
- mov edi, 9 ; fd 9, which isn't open,
- lea rsi, [rip+msg] ; buffer,
- mov edx, 3 ; 3 bytes)
- syscall ; rax = -9: EBADF
- mov eax, 1 ; write(fd 1 = standard output, ...)
- mov edi, 1
- syscall ; rax = 3
- mov eax, 60 ; exit(0)
- xor edi, edi
- syscall
getpid returns 1000, the write to file descriptor 9 returns −9, and the write to standard output prints hi and returns 3, the number of bytes written. Watch rcx after each syscall: the instruction itself overwrites it with the return address, and r11 with the flags, which is why the calling convention treats both as destroyed by a system call. Every other register comes back unchanged.
What a real program asks for
To see the system calls a real program makes, Linux has strace. It's built on ptrace, the same system call debuggers use: the tracer asks the kernel to stop the traced process at every system call entry and exit, and reads its registers each time. A minimal version is about forty lines of C. Here is what it printed for the C program printf("hello\n"), compiled normally, on 64-bit ARM Linux (runs of similar calls grouped):
brk () = 0xaaab17737000 ← where is the heap?
mmap () = 0xffffbd4d4000
faccessat ("/etc/ld.so.preload") = -2 ← ENOENT: no such file
openat ("/etc/ld.so.cache") = 3 ← the loader's index of libraries
fstatat, mmap, close ← map the cache, close it
openat ("/lib/aarch64-linux-gnu/libc.so.6") = 3
read () = 0x340 ← libc's ELF headers
fstatat, mmap ×4, munmap ×2, mprotect, close ← map libc's segments
set_tid_address, set_robust_list, rseq ← set up the main thread
mprotect ×3 ← make relocated data read-only
prlimit64, munmap, fstatat
getrandom () = 8 ← a random key for malloc's checks
brk ×2 ← grow the heap for stdout's buffer
write ("hello\n") = 6
exit_group (0)
Thirty-two system calls, one of which does the work the program asked for. The others are the dynamic loader finding and mapping the C library, marking pages read-only after relocation, and setting up thread state; then malloc fetches a random key it uses to detect double frees (the check from the heap chapter), and printf grows the heap for its output buffer. Built with -static, the same program makes 15 system calls; ls / makes 74, the largest groups being mmap (14), mprotect and close (8 each).
The traced faccessat also shows an error in real life: the loader checks for /etc/ld.so.preload, gets −2 (ENOENT) because the file doesn't exist, and carries on — errors are an ordinary answer, not a crash.
How many, and how stable
Linux on 64-bit ARM numbers its system calls from 0 to 450, and x86-64 has a similar number. Linux treats the system call numbers and their behavior as a stable promise: a binary from twenty years ago still runs, and programs may make system calls directly, without the C library — Go's runtime does.
Other systems make a different promise. Windows's system call numbers are undocumented and change between versions; programs must go through the functions of ntdll.dll, which knows the numbers for the running version. macOS also reserves the right to change its system calls and supports only calls made through its system library, libSystem — which is why Go switched to calling libSystem on macOS in 2018, after a kernel update broke programs that made system calls directly. OpenBSD goes further and blocks system calls made from anywhere but its C library. The boundary is the same everywhere; where the stable interface lies is a choice.
What it costs, and how to avoid it
In a Linux virtual machine on an Apple M2, a getppid system call — about the cheapest there is — took 154 ns, and clock_gettime made as a real system call 186 ns. The same clock_gettime through the vDSO, the page of kernel code mapped into every process (visible as [vdso] in the address space), took 17 ns: it reads the clock from shared memory without entering the kernel at all.
That's the general strategy for making system calls cheaper: make fewer of them. Programs buffer their output — printf collects text and makes one write for many calls. Servers read and write many buffers per call. And Linux's io_uring goes further: the program and the kernel share queues in memory, the program writes requests into one and reads completions from the other, and a single system call — or none, with a kernel thread polling the queue — can submit hundreds of operations.
Takeaways
- The operating system machine is the ISA plus new instructions — system calls — carried out by the kernel.
- User mode can't touch hardware, page tables or kernel memory; everything else is asked for, so the kernel can check it: that's isolation. Real systems use two privilege levels (x86 rings 0 and 3, ARM EL1 and EL0).
- A system call passes a number and arguments in registers, traps into the kernel, which dispatches through its system call table, validates pointers with
copy_from_user, and returns the result inrax. Errors are negative errno values, which the C library turns into −1 anderrno. syscalldestroysrcxandr11; every other register is preserved.- A dynamically linked "hello" makes 32 system calls, 31 of them to load and set up the C library; traced with a minimal
ptrace-based strace. - Linux keeps its system call numbers stable; Windows and macOS only support calls through their system libraries.
- System calls cost around 150–190 ns here; the vDSO answers
clock_gettimein 17 ns without entering the kernel, and buffering and io_uring make fewer calls do more work.