Skip to content

Level 0 · Chapter 0.5

Why is my computer slow?

The usual reasons a computer feels slow, mapped to the levels that cause them: too much work for the cores, waiting on memory, running out of RAM, waiting on storage or the network, heat and contention, with real measurements of each and the tools to tell them apart.

"The computer is slow" can mean a dozen different things. A program might be doing too much work, or waiting for memory, or for the disk, or for a server on the other side of the world. The machine might be out of memory and shuffling pages around, too hot and running at a lower clock, or simply shared among too many programs. Each of these has a different cause at a different level, and a different fix.

This chapter goes through the usual suspects, measures each one on the Mac it was written on (an Apple M2 Ultra with 16 performance cores, 8 efficiency cores and 64 GB of memory), and shows how to tell them apart.

Four resources

Almost every slowdown comes down to a program waiting for one of four things:

ResourceThe program is…Seen inWhere it's explained
CPUcomputing, and there's too much to computehigh CPU %pipelining, out-of-order execution
memorywaiting for data to arrive from RAMhigh CPU %, but little work donecaches
RAM capacitywaiting for pages to be decompressed or read back from diskmemory pressure, swap in usevirtual memory
storage and networkblocked in a system call, waiting for I/Olow CPU %, the app still stallssystem calls, loading a web page

The first step is always to look. On macOS, Activity Monitor shows CPU, memory, energy, disk and network per process; top shows the same in a terminal. Its header, taken while this chapter was being written:

Processes: 1231 total, 12 running, 1219 sleeping, 10197 threads
Load Avg: 11.54, 14.44, 13.97
CPU usage: 30.48% user, 9.64% sys, 59.86% idle
PhysMem: 58G used (6043M wired, 8978M compressor), 5373M unused.
VM: 19165(0) swapins, 461956(0) swapouts.

The load average is the average number of threads running or ready to run over the last 1, 5 and 15 minutes. Here about 12 to 14, on 24 cores: busy, but with room to spare. On Linux the same tools exist (top, htop, vmstat, free), and on Windows, Task Manager and Resource Monitor.

CPU-bound: too much work

A program is CPU-bound when its speed is set by how fast the cores execute its instructions. The fix is less work (a better algorithm, as in the complexity chapter) or more cores working on it at once.

That second fix has limits. A single-threaded program uses one core, whatever the machine has: in Activity Monitor it shows at 100%, meaning one core's worth, while the other 23 cores sit idle. And when there are more busy threads than cores, they take turns. This test runs a pure arithmetic loop (no memory traffic at all) in several processes at once, and times how long until all of them finish:

Copies runningWall time (two runs)
10.75 s
160.86–0.89 s
241.14–1.19 s
482.03–2.07 s

Up to 16 copies, each gets its own performance core and the time barely moves. From 17 to 24, the extra copies land on the slower efficiency cores, and everyone waits for the stragglers. At 48 copies, twice the number of cores, the scheduler splits each core between two processes, and everything takes 2.7 times longer than alone. That's contention: a background job that keeps every core busy (a backup, an indexer, a build, a browser tab running a heavy script) makes everything else slower, even though nothing is "wrong" with the program you're using.

Memory-bound: busy, but waiting

A program can show 100% CPU and still spend most of its time waiting, for memory. A core can execute several instructions per nanosecond, but a load that misses every cache has to go all the way to DRAM. This test follows a chain of pointers through a block of memory in random order, so each load's address depends on the previous one and nothing can be predicted or overlapped. The time per load, by size of the block:

Working setTime per loadServed by
16–64 KiB1.2 nsL1 cache (128 KiB per core)
1 MiB6.2–6.3 nsL2 cache (16 MiB, shared by 4 cores)
8 MiB10–12 nsstill L2, but slower as it fills
32 MiB79–81 nspartly beyond the caches
128 MiB – 1 GiB129–137 nsDRAM

From the smallest cache to main memory, the same instruction becomes about 100 times slower. The caches chapter explains why; here's what it looks like from the application's side. This program sums the same 64 million integers (256 MiB) twice, with identical code: once reading them in memory order, once through a shuffled index.

for (size_t i = 0; i < n; i++) s1 += a[seq[i]];   /* seq = 0, 1, 2, 3, … */
for (size_t i = 0; i < n; i++) s2 += a[idx[i]];   /* idx = a shuffled 0 … n-1 */
Access orderTimePer element
in order14 ms0.2 ns
shuffled258–269 ms3.9–4.0 ns

Same additions, same result, about 19 times slower. The shuffled loop is still much faster than 130 ns per element because its loads don't depend on each other, so an out-of-order core keeps many misses in flight at once. But both loops look identical in Activity Monitor: one core at 100%. CPU usage doesn't tell you whether the CPU is computing or waiting. Profilers that read the hardware's performance counters (Instruments on macOS, perf on Linux, VTune on Intel) can tell the two apart by counting cache misses and stall cycles.

What matters is the working set: the data a program uses over a short stretch of time. If it fits in a cache, access is fast; if it's even slightly too big, the whole working set can fall out of the cache on every pass. The demo below makes that visible with a cache that has room for 16 items, keeping the most recently used ones, and a program that goes over its data four times:

The cache cliff

Try it: Pick how many items the program uses, then watch it go over them four times. The cache (the shelf) holds 16: green is a hit, red a trip to slow memory.

Items used:
The data, in main memory
0123456789101112131415
The cache: 16 places
pass 1/40 hits · 0 misses

With 16 items, the first pass misses 16 times and the next three passes hit every time: a hit rate of 75%. With 17, one item more than fits, the hit rate drops to zero: by the time the program comes back to an item, the cache has just thrown it out to make room for the others. Performance doesn't degrade gently when a working set outgrows a cache; it falls off a cliff.

Out of memory: compression and swapping

The same cliff exists one level down, between RAM and storage. When the programs' combined working sets no longer fit in physical memory, the operating system has to make room. Early timesharing systems swapped whole processes out every hundred milliseconds or so, and Denning's working set model decided what to keep. The model survives; whole-process swapping doesn't. Modern systems page out individual pages, least recently used first, and before writing anything to disk, macOS, Windows 10 and later, and many Linux configurations compress them in memory.

vm_stat on this Mac shows the compressor at work (pages are 16 KiB):

Pages stored in compressor:                  2436936.
Pages occupied by compressor:                 574584.
Swapins:                                       19165.
Swapouts:                                     461956.

37.2 GiB of memory contents squeezed into 8.8 GiB, a ratio of 4.2 to 1. Decompressing a page costs microseconds; reading it back from storage costs more; recomputing or redownloading it costs far more. Beyond that, nearly 7 GB of swap space on the SSD was in use:

$ sysctl vm.swapusage
vm.swapusage: total = 8192.00M  used = 6953.88M  free = 1238.12M  (encrypted)

None of that is a problem while the pages that were swapped out stay unused: memory holding a browser tab you haven't looked at in a week. It becomes a problem when the working sets themselves don't fit: every access to a paged-out page is a page fault that waits for the disk, and the pages brought in push out others that are needed a moment later. That's thrashing: the disk is busy, the CPU is idle, and everything crawls. Activity Monitor's memory pressure graph turns yellow, then red, and on the command line vm_stat shows swap-ins climbing. The fix is to use less memory at once (close something) or to have more.

Waiting on storage and the network

A program waiting for a disk or a server isn't using the CPU at all. It's blocked inside a system call (read, recv, connect), and the scheduler has given its core to someone else. CPU usage is low, and yet the app doesn't respond, especially if the wait happens on its main thread, the one running its event loop.

The costs span orders of magnitude. Reading one random 4 KiB block from a 2 GiB file on this Mac:

Where the data came fromTime per read
the SSD (file cache bypassed with F_NOCACHE), median80–140 µs
the operating system's file cache in RAM2.2 µs
for comparison: one load from DRAM0.13 µs

The SSD figure moves with the drive's state and the machine's load: the disks and SSDs chapter measured a median of 80–85 µs, and repeated runs for this chapter, on a busier machine, gave 101–140 µs. That chapter also shows how an SSD gets much more done when many reads are in flight at once.

A spinning hard disk needs several milliseconds to move its head and wait for the platter, about a hundred times more than this SSD, which is why an old laptop's hard disk was so often "the" reason it was slow. And a network round trip, as the web page chapter measured, costs about 6 ms to a nearby server and far more to a distant one. An app that makes fifty sequential requests to a server 100 ms away needs five seconds, whatever the CPU.

Activity Monitor's Disk and Network tabs show which processes are moving data. A process that uses little CPU but lots of disk or network, or a system-wide spike in either, points here.

Put side by side, these waits span more than seven orders of magnitude. Stretch them to human time and the gaps become obvious: if reading from the L1 cache took one second, a trip to main memory would take almost two minutes, a read from the SSD about a day, and loading a web page most of a year.

How long things take

Try it: Switch to “If L1 took 1 second” to stretch every wait to human time, and to “Linear” to see how the slow ones dwarf the rest. Click a row to compare it with a cache hit.

Green bars are memory reads, orange ones file reads, pink ones the waits you notice on screen.

Heat: thermal throttling

A chip's power grows with its clock frequency and, faster still, with its supply voltage, and higher frequencies need higher voltages; the limits of computing chapter gives the formula. All that power becomes heat. When the chip gets too hot, the power management firmware lowers the frequency and voltage until the heat can be carried away: thermal throttling. A thin, fanless laptop under a long, heavy load runs slower after a few minutes than it did at the start; a desktop with big fans, like this one, rarely throttles.

macOS records thermal events, and here there were none:

$ pmset -g therm
Note: No thermal warning level has been recorded
Note: No performance warning level has been recorded
Note: No CPU power status has been recorded

Laptops also slow down on purpose to save battery: low-power modes cap the clocks, and background work is steered to the efficiency cores. A benchmark run on battery, hot, or right after another heavy job can differ from one run cold on mains power.

A checklist

SymptomLikely causeLook at
one process at 100% (one core), the rest idlesingle-threaded, CPU-bound workActivity Monitor CPU tab; a profiler
all cores busytoo many jobs at once: contentionload average; which processes use CPU
high CPU, but slower than expectedmemory-bound: cache missesa profiler with hardware counters
memory pressure high, swap growingworking sets bigger than RAMMemory tab; vm_stat
low CPU, app frozen or spinningblocked on disk or network, on the main threadDisk and Network tabs; a sample of the process
fast at first, slower after minutesthermal throttlingpmset -g therm; temperatures

Takeaways

  • Slowness is almost always waiting: for the cores, for memory, for RAM to be freed, for storage or for the network.
  • CPU-bound work is limited by cores: a single-threaded program uses one of 24; 48 busy processes on 24 cores take 2.7 times as long.
  • Memory-bound work shows as 100% CPU but is waiting: a load costs 1.2 ns from L1 and about 130 ns from DRAM, and summing 256 MiB in shuffled order took 19 times longer than in order.
  • Performance falls off a cliff when a working set outgrows its cache or RAM: in the demo, one item too many drops the hit rate from 75% to 0%.
  • When RAM runs short, modern systems compress first (4.2 to 1 here), then swap; thrashing starts when the working sets themselves don't fit.
  • I/O waits don't show as CPU use: a random SSD read took about 100 µs, a cached one 2.2 µs, a network round trip about 6 ms.
  • Heat and power limits lower the clock; a machine can get slower under sustained load.

In this level

  1. 0.1What happens when you open an app
  2. 0.2From keypress to pixel
  3. 0.3What happens when a web page loads
  4. 0.4Apps, processes and files
  5. 0.5Why is my computer slow?