Skip to content

Level 4 · Chapter 4.3

System calls and privilege levels

The operating system as a machine with extra instructions: why programs can't touch the hardware, what happens on the way into the kernel and back, how errors come back, which system calls a real program makes — traced with a home-made strace — and what they cost.

Tanenbaum describes the operating system as a level of its own: the operating system machine. It offers everything the ISA offers, plus a set of new instructions — open a file, read from it, create a process, allocate memory, send a packet. Ordinary instructions still run directly on the hardware. The new ones, the system calls, are carried out by the kernel on the program's behalf. From the program's point of view, read is just one more instruction, a slow and very powerful one.

The traps chapter showed the hardware side: the syscall instruction, the switch to kernel mode, the cost. This chapter looks at the same boundary from the operating system's side.

Why programs can't just do it themselves

A program in user mode can compute anything, but it can't touch the machine. The instructions that control hardware — reading and writing I/O ports, loading page tables, disabling interrupts, halting the CPU, changing privilege — fault if user code tries them. The page tables mark the kernel's own memory as off-limits to user mode. So everything that touches a device, another process or the kernel's data has to be asked for.

That's the point: isolation. The kernel can check every request — does this process have permission to open that file? is that buffer inside its own memory? — and one program can't read another's memory, crash the disk driver, or keep the CPU forever. The same design runs on every modern OS, with the same two levels in practice: x86's rings 0 (kernel) and 3 (user), with rings 1 and 2 unused; ARM64's EL1 and EL0; RISC-V's S-mode and U-mode.

The protection works in the other direction too. Since the Meltdown and similar attacks of 2018, and even before, kernels don't want to touch user memory by accident: x86's SMEP and SMAP and ARM's PAN make the CPU fault if the kernel executes user code or reads user memory outside of the few routines designed for it. And after Meltdown, Linux on affected x86 chips even switches to separate page tables on every kernel entry (KPTI), so that user mode can't speculatively read kernel memory — one more reason system calls got slower.

Into the kernel and back

On x86-64 Linux, a system call goes like this:

  1. The C library function — write, getpid, … — puts the call number in rax and up to six arguments in rdi, rsi, rdx, r10, r8, r9, and executes syscall.
  2. The CPU saves the return address in rcx and the flags in r11, switches to ring 0, and jumps to the kernel's entry point, which the kernel stored in a model-specific register at boot.
  3. The kernel's entry code switches to the kernel stack, saves the user registers, and uses rax as an index into its system call table, an array of function pointers.
  4. The handler checks its arguments. Pointers get special care: a user program could pass the address of kernel memory, or of nothing at all, so the kernel never dereferences them directly. It copies data in and out with copy_from_user and copy_to_user, which verify the addresses and recover cleanly from page faults.
  5. The result goes back in rax, the user registers are restored, and sysret returns to the instruction after syscall, in ring 3.

Errors come back as negative numbers: −2 is ENOENT (no such file), −9 is EBADF (bad file descriptor), −13 is EACCES (permission denied), −38 is ENOSYS (no such system call). The C library's wrapper turns them into the convention C programs see: it stores the positive code in the global errno and returns −1.

The simulator implements a tiny kernel with four of Linux's system calls — write, exit, getpid and getppid — and returns ENOSYS for the rest:

Live · Three system calls, one of them failing

Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.

program— ▸ is the next instruction
  1. .data
  2. msg: .asciz "hi\n"
  3. .text
  4. mov eax, 39 ; getpid()
  5. syscall
  6. mov ebx, eax ; keep the result: 1000
  7. mov eax, 1 ; write(
  8. mov edi, 9 ; fd 9, which isn't open,
  9. lea rsi, [rip+msg] ; buffer,
  10. mov edx, 3 ; 3 bytes)
  11. syscall ; rax = -9: EBADF
  12. mov eax, 1 ; write(fd 1 = standard output, ...)
  13. mov edi, 1
  14. syscall ; rax = 3
  15. mov eax, 60 ; exit(0)
  16. xor edi, edi
  17. syscall
step 0
Loading emulator…
The process as the operating system sees it: an address space of code, globals, heap and stack.

getpid returns 1000, the write to file descriptor 9 returns −9, and the write to standard output prints hi and returns 3, the number of bytes written. Watch rcx after each syscall: the instruction itself overwrites it with the return address, and r11 with the flags, which is why the calling convention treats both as destroyed by a system call. Every other register comes back unchanged.

What a real program asks for

To see the system calls a real program makes, Linux has strace. It's built on ptrace, the same system call debuggers use: the tracer asks the kernel to stop the traced process at every system call entry and exit, and reads its registers each time. A minimal version is about forty lines of C. Here is what it printed for the C program printf("hello\n"), compiled normally, on 64-bit ARM Linux (runs of similar calls grouped):

brk             ()                                    = 0xaaab17737000   ← where is the heap?
mmap            ()                                    = 0xffffbd4d4000
faccessat       ("/etc/ld.so.preload")                = -2               ← ENOENT: no such file
openat          ("/etc/ld.so.cache")                  = 3                ← the loader's index of libraries
fstatat, mmap, close                                                     ← map the cache, close it
openat          ("/lib/aarch64-linux-gnu/libc.so.6")  = 3
read            ()                                    = 0x340            ← libc's ELF headers
fstatat, mmap ×4, munmap ×2, mprotect, close                             ← map libc's segments
set_tid_address, set_robust_list, rseq                                   ← set up the main thread
mprotect ×3                                                              ← make relocated data read-only
prlimit64, munmap, fstatat
getrandom       ()                                    = 8                ← a random key for malloc's checks
brk ×2                                                                   ← grow the heap for stdout's buffer
write           ("hello\n")                           = 6
exit_group      (0)

Thirty-two system calls, one of which does the work the program asked for. The others are the dynamic loader finding and mapping the C library, marking pages read-only after relocation, and setting up thread state; then malloc fetches a random key it uses to detect double frees (the check from the heap chapter), and printf grows the heap for its output buffer. Built with -static, the same program makes 15 system calls; ls / makes 74, the largest groups being mmap (14), mprotect and close (8 each).

The traced faccessat also shows an error in real life: the loader checks for /etc/ld.so.preload, gets −2 (ENOENT) because the file doesn't exist, and carries on — errors are an ordinary answer, not a crash.

How many, and how stable

Linux on 64-bit ARM numbers its system calls from 0 to 450, and x86-64 has a similar number. Linux treats the system call numbers and their behavior as a stable promise: a binary from twenty years ago still runs, and programs may make system calls directly, without the C library — Go's runtime does.

Other systems make a different promise. Windows's system call numbers are undocumented and change between versions; programs must go through the functions of ntdll.dll, which knows the numbers for the running version. macOS also reserves the right to change its system calls and supports only calls made through its system library, libSystem — which is why Go switched to calling libSystem on macOS in 2018, after a kernel update broke programs that made system calls directly. OpenBSD goes further and blocks system calls made from anywhere but its C library. The boundary is the same everywhere; where the stable interface lies is a choice.

What it costs, and how to avoid it

In a Linux virtual machine on an Apple M2, a getppid system call — about the cheapest there is — took 154 ns, and clock_gettime made as a real system call 186 ns. The same clock_gettime through the vDSO, the page of kernel code mapped into every process (visible as [vdso] in the address space), took 17 ns: it reads the clock from shared memory without entering the kernel at all.

That's the general strategy for making system calls cheaper: make fewer of them. Programs buffer their output — printf collects text and makes one write for many calls. Servers read and write many buffers per call. And Linux's io_uring goes further: the program and the kernel share queues in memory, the program writes requests into one and reads completions from the other, and a single system call — or none, with a kernel thread polling the queue — can submit hundreds of operations.

Takeaways

  • The operating system machine is the ISA plus new instructions — system calls — carried out by the kernel.
  • User mode can't touch hardware, page tables or kernel memory; everything else is asked for, so the kernel can check it: that's isolation. Real systems use two privilege levels (x86 rings 0 and 3, ARM EL1 and EL0).
  • A system call passes a number and arguments in registers, traps into the kernel, which dispatches through its system call table, validates pointers with copy_from_user, and returns the result in rax. Errors are negative errno values, which the C library turns into −1 and errno.
  • syscall destroys rcx and r11; every other register is preserved.
  • A dynamically linked "hello" makes 32 system calls, 31 of them to load and set up the C library; traced with a minimal ptrace-based strace.
  • Linux keeps its system call numbers stable; Windows and macOS only support calls through their system libraries.
  • System calls cost around 150–190 ns here; the vDSO answers clock_gettime in 17 ns without entering the kernel, and buffering and io_uring make fewer calls do more work.

In this level

  1. 4.1Executable files: ELF, PE and Mach-O
  2. 4.2Processes and the address space
  3. 4.3System calls and privilege levels
  4. 4.4Virtual memory and pagingPlanned
  5. 4.5Files, devices and I/OPlanned
  6. 4.6Threads and synchronizationPlanned
  7. 4.7Hardware virtualization and hypervisorsPlanned
  8. 4.8Inside UNIX and WindowsPlanned