Skip to content

Level 4 · Chapter 4.7

Hardware virtualization and hypervisors

How one machine runs several operating systems: type 1 and type 2 hypervisors, trap-and-emulate and the Popek–Goldberg requirements x86 failed, binary translation and paravirtualization, VT-x, AMD-V and ARM's EL2, nested page tables, virtio devices, what virtualization costs, and why containers are something else, with the Linux VM behind Docker on a Mac as the running example.

Every Linux measurement in the previous chapters ran inside a virtual machine. On the Mac this chapter was written on, uname -a says:

Darwin studio.local 25.6.0 Darwin Kernel Version 25.6.0: … RELEASE_ARM64_T6020 arm64

and the same command inside a Docker container says:

Linux 84ef2fddeaae 6.12.76-linuxkit #1 SMP Thu Apr 30 11:19:05 UTC 2026 aarch64 GNU/Linux

Two different kernels, both running on the same M2 at the same time. The Linux one believes it has 24 CPUs, 8 GiB of RAM, disks and a network card. What it actually has is a slice of the Mac, handed out by a hypervisor: the software that creates and runs virtual machines, each of which looks to its operating system like a complete computer.

Why virtualize

The classic motivation is the hosting company: rather than buy a server per customer, run many customers' systems on one server and add hardware only when the existing machines are full. That's today's cloud. Each virtual machine is isolated from the others as if on separate hardware, can be given a precise share of CPU, memory and I/O, can be snapshotted, and can be moved to another physical server while it runs. For individuals, the other reason is running several operating systems at once. That's how Docker runs Linux containers on a Mac, whose own kernel isn't Linux.

Type 1 and type 2 hypervisors

Hypervisors are traditionally classified by where they sit:

  • A type 1 hypervisor runs directly on the hardware, in place of an operating system, and every OS runs above it as a guest: VMware ESXi, Xen, Microsoft Hyper-V. It's what cloud servers run.
  • A type 2 hypervisor runs as a program on top of a normal host OS: VirtualBox, VMware Workstation and Fusion, Parallels.

The line has blurred. Linux's KVM turns the Linux kernel itself into a hypervisor: each VM is an ordinary Linux process, with its virtual CPUs as threads. macOS provides the same kind of kernel service to applications through its Hypervisor framework, and a higher-level Virtualization framework on top. That's what runs the Docker VM here: the Mac's process list shows com.apple.Virtualization.VirtualMachine, a process of Apple's framework, and sysctl kern.hv_support reports 1.

Trap and emulate

The guest's kernel expects to run in kernel mode, in control of the page tables, the interrupts and the devices. The hypervisor can't let it: it would take over the real machine. The classic solution, used by IBM's CP/CMS and VM/370 from the late 1960s and early 1970s, is trap and emulate. The guest kernel runs deprivileged, in user mode. Its ordinary instructions (additions, loads, branches) run directly on the hardware at full speed. When it executes a privileged instruction, such as loading a page table register or disabling interrupts, the CPU traps, exactly as for a user program that tries one, and the trap goes to the hypervisor. The hypervisor performs the operation on the guest's virtual state and returns. The guest never notices.

The same mechanism can be seen, one level down, in the Linux VM here. MIDR_EL1, the ARM register that identifies the CPU model, is readable only in kernel mode. A user program that executes mrs x0, midr_el1 takes an exception to the kernel, and Linux, instead of killing the program, emulates the instruction: it puts the value in x0 and resumes after the instruction.

unsigned long v;
__asm__ volatile("mrs %0, midr_el1" : "=r"(v));   // privileged, but Linux emulates it
printf("read %#lx\n", v);

In the Linux VM it printed read 0x610f0000 (0x61 is Apple's implementer code), and each read cost about 118 ns, a little less than a real system call in the same VM (154 ns). Reading CurrentEL, which Linux doesn't emulate, killed the program with SIGILL, as did both reads under macOS, which emulates neither. A hypervisor does to the guest kernel what this kernel does to its program.

When the architecture doesn't cooperate

In 1974, Popek and Goldberg stated when trap and emulate works. Call sensitive the instructions that read or change the machine's control state (privilege level, interrupt mask, memory mapping), and privileged those that trap when executed in user mode. An architecture is efficiently virtualizable if every sensitive instruction is privileged: then every operation the hypervisor must intercept traps, and everything else can run directly.

32-bit x86 failed the test. A 2000 study counted 17 instructions that were sensitive but didn't trap. POPF is the classic: in kernel mode it can change the interrupt-enable flag, in user mode the same instruction silently leaves that flag unchanged. There's no trap, so a deprivileged guest kernel that disables interrupts with it just doesn't, and the hypervisor never hears about it. SGDT, SIDT and SMSW let user code read descriptor-table addresses and control bits, so a guest could see the hypervisor's real values instead of its own.

Two software workarounds made x86 virtualization practical anyway:

  • Binary translation, VMware's approach from 1999: the hypervisor scans guest kernel code before it runs and rewrites the problematic instructions into safe sequences that call the hypervisor, caching the translated code. User-mode guest code runs untouched.
  • Paravirtualization, Xen's approach from 2003: modify the guest OS so it doesn't use those instructions at all and calls the hypervisor directly, through hypercalls, when it needs a privileged operation. Fast, but it needs a modified kernel.

Hardware support

Intel and AMD then fixed the problem in hardware: VT-x in 2005 and AMD-V in 2006. Rather than making every sensitive instruction trap in ring 3, they added a new dimension. The CPU runs either in VMX root mode, for the hypervisor, or in VMX non-root mode, for guests, and each mode has its own rings 0 to 3. A guest kernel really runs in ring 0, but of non-root mode. Configurable events cause a VM exit to the hypervisor: sensitive instructions, chosen I/O accesses, certain interrupts, page faults in the second-level tables. The hypervisor describes all this in a VMCS (virtual machine control structure), which also holds the guest's saved state, and resumes the guest with a VM entry.

ARM designed virtualization into its privilege levels. Since the ARMv7 virtualization extensions (where it was called Hyp mode), and in practically every ARMv8 and ARMv9 application core, there's an EL2 between the guest kernel's EL1 and the secure firmware's EL3: the hypervisor runs at EL2, and registers there decide which guest operations trap to it. The MIDR_EL1 value above is an example of what EL2 controls: when a guest kernel reads that register, the hardware returns a value the hypervisor chose. ARMv8.1 added VHE (virtualization host extensions) so that a host kernel like Linux with KVM can run entirely at EL2. RISC-V has its H extension for the same purpose.

Memory: two translations

A guest OS manages page tables that map its processes' virtual addresses to what it believes are physical addresses. But those "guest-physical" addresses are just a range of the hypervisor's memory, which must be translated again to host-physical addresses. Before hardware help, hypervisors maintained shadow page tables: a combined guest-virtual to host-physical table, kept in sync by trapping every change the guest made to its own tables: correct, but costly.

Hardware now does both translations. Intel's EPT (extended page tables, 2008) and AMD's NPT (nested page tables) add a second tree, owned by the hypervisor, that the MMU applies to every guest-physical address, including the addresses of the guest's own page tables. On a TLB miss, each of the four levels of the guest's walk needs its address translated through the four levels of the host's tree, plus the final data address: up to 24 memory accesses instead of 4. The TLB caches the combined result, so most accesses pay nothing. But it's part of why the pointer chase in the virtual memory chapter cost 30 ns per load with 4 KiB pages in this VM, and why huge pages help VMs even more than native systems. ARM calls its version stage-2 translation.

I/O: emulated, paravirtual, or direct

The simplest hypervisor intercepts every I/O access by the guest and carries it out in software. That's the most compatible approach (the guest sees a familiar device, like an old network card, and its unmodified driver works) and the slowest, since every register access is a VM exit. Two alternatives dominate today.

Paravirtual devices give up on imitating real hardware. The guest runs drivers written for a virtual device whose interface is designed to be cheap: requests go into rings of buffers in shared memory, and the guest notifies the hypervisor once per batch. The standard for this is virtio. The Docker VM's devices, listed from /sys/bus/virtio/devices, are all virtio: a network device, two block devices (disks), a console, an entropy source, a memory balloon (which lets the host reclaim memory from the guest), a socket device for host-guest communication, and six virtio-fs devices, which share Mac directories into the VM: the -v "$PWD":/w of every Docker command in these chapters.

Device passthrough gives a VM a real device directly, protected by the IOMMU, which translates and restricts the device's DMA addresses the way the MMU does for the CPU's. With SR-IOV, one network card or SSD presents itself as many virtual functions, each assigned to a different VM, so the guest's driver talks to the hardware without the hypervisor in the path. Cloud providers rely on this for near-native I/O.

What it costs

With hardware support, a guest's unprivileged code runs directly on the CPU. A function call measured 0.9 ns under macOS and 0.9 ns in the Linux VM, and the single-thread atomic additions of the threads chapter took the same 2 and 4 ns on both. The costs appear where the hypervisor gets involved, or where hardware does double work:

  • System calls don't involve the hypervisor: they go straight to the guest kernel. getppid took 154–155 ns in the Linux VM against 95 ns under macOS, but that compares two different kernels, not a VM and bare metal.
  • TLB misses walk two sets of tables, as above.
  • Memory the guest touches for the first time may need the hypervisor too: a guest page fault, then a stage-2 fault if the host hasn't backed that guest-physical page yet.
  • I/O crosses into the hypervisor or a helper process. An fsync of 4 KiB took 740 µs in the VM, against 40 µs for fsync natively; the virtual disk is a file on the Mac, and the flush has to go all the way through.

For most workloads the result is within a few percent of native. For I/O-heavy ones, it depends on how the devices are virtualized.

Containers are not virtual machines

A Docker container looks like a small machine too, but it isn't virtualized hardware. It's a group of ordinary processes, on the same kernel as everything else, that the kernel shows a restricted view of the system through namespaces, and limits through cgroups. Inside a container:

$ echo $$
1
$ ls /proc/self/ns
cgroup  ipc  mnt  net  pid  pid_for_children  time  time_for_children  user  uts

The shell believes it's process 1, because it's the first process in its own PID namespace; the VM's kernel knows it by another number. The other namespaces give it its own mount table, network stack, hostname, IPC objects and cgroup view. Started with --memory 256m --cpus 2, the container's cgroup files read 268435456 for memory.max and 200000 100000 for cpu.max: 200 ms of CPU per 100 ms period, two cores' worth. The thrashing experiment in the virtual memory chapter used exactly that memory limit.

Because containers share the kernel, uname inside any container on this Mac prints the same 6.12.76-linuxkit: the one kernel in the VM, shared by every container. That's why containers start in milliseconds and cost almost nothing, and also why they isolate less than VMs (a kernel vulnerability exposes every container on the machine), and why Linux containers on macOS or Windows need a Linux VM underneath. Where isolation matters, the two are combined: AWS's Firecracker runs each serverless function or container in its own tiny VM, booting in a fraction of a second.

Takeaways

  • A hypervisor runs virtual machines, each with its own OS. Type 1 runs on the hardware (ESXi, Xen, Hyper-V), type 2 on a host OS; KVM and macOS's Hypervisor framework make the host kernel the hypervisor.
  • Trap and emulate: guest code runs directly, and privileged operations trap to the hypervisor. It works only if every sensitive instruction is privileged (Popek and Goldberg, 1974). 32-bit x86 had 17 exceptions, such as POPF.
  • Before hardware support, binary translation (VMware) and paravirtualization (Xen) worked around x86. VT-x and AMD-V added root and non-root modes with VM exits; ARM has EL2.
  • EPT/NPT (ARM's stage 2) translate guest-physical to host-physical addresses in hardware; a TLB miss can take up to 24 memory accesses.
  • I/O is emulated, paravirtual (virtio: the Docker VM here has 13 virtio devices) or passed through with an IOMMU and SR-IOV.
  • Unprivileged code runs at native speed; costs come from exits, nested page walks and I/O.
  • Containers are processes isolated by namespaces and limited by cgroups, sharing one kernel: every container on this Mac runs on the VM's Linux 6.12.

In this level

  1. 4.1Executable files: ELF, PE and Mach-O
  2. 4.2Processes and the address space
  3. 4.3System calls and privilege levels
  4. 4.4Virtual memory and paging
  5. 4.5Files, devices and I/O
  6. 4.6Threads and synchronization
  7. 4.7Hardware virtualization and hypervisors
  8. 4.8Inside UNIX and Windows