Skip to content

Level 2 · Chapter 2.1

Compiled, interpreted and JIT-compiled languages

Three ways to run a program on a CPU that only understands machine code: compiling ahead of time, interpreting bytecode, and compiling just in time — with real CPython and V8 bytecode, V8's JIT tiers caught in the act, and the same loop timed in C, JavaScript and Python.

A CPU executes only its own machine code. A program written in C, Python or JavaScript has to get there somehow, and languages take three main routes:

  1. Compile ahead of time (AOT): translate the whole program to machine code before it runs. C, C++, Rust, Go and Swift work this way.
  2. Interpret: run a program that reads your program and carries it out, step by step. CPython works this way, via bytecode.
  3. Compile just in time (JIT): start by interpreting, watch what the program actually does, and compile the hot parts to machine code while it runs. JavaScript engines, the Java and .NET virtual machines, PyPy and LuaJIT work this way.

The route is a property of the implementation, not of the language: there are C interpreters and ahead-of-time Java compilers. But each language has a usual route, and it shapes how fast programs start, how fast they run, and what's left for someone reading the binary.

Ahead-of-time compilation

A compiler reads the whole program, checks it, optimizes it, and writes machine code. This site's simulator has a small C compiler built in. The demo below compiles a C function live and opens at the C level. Zoom in to see the assembly it produced, and further down to the bytes and micro-operations:

Live · A C function, compiled in your browser
C source — click a line number for a breakpoint
  1. int total(int n) {
  2. int s = 0;
  3. for (int i = 0; i < n; i++)
  4. s = s + i * 2;
  5. return s;
  6. }
  7. int main() {
  8. return total(5);
  9. }
step 0
Loading emulator…
Your program as you wrote it: the current line, its variables by name, and its output.

The program returns 20 after 81 instructions. The simulator's compiler is deliberately simple, like gcc -O0: every variable lives on the stack, so the code maps line by line to the C.

A production compiler with optimizations does far more. Here is the loop s = (s + i * 2) % 1000003 compiled by clang with -O2 for x86-64, trimmed to the core of the loop:

.LBB0_8:
    add   rsi, rcx            ; s += i*2   (rcx holds i*2)
    mov   rax, rsi
    imul  r8                  ; multiply by a "magic" constant…
    add   rdx, rsi
    mov   rax, rdx
    shr   rax, 63
    sar   rdx, 19             ; …and shift: rdx = s / 1000003
    add   rdx, rax
    imul  rax, rdx, 1000003
    sub   rsi, rax            ; s -= (s / 1000003) * 1000003, i.e. s %= 1000003
    …                         ; the same again for the next i: the loop is unrolled twice
    add   rcx, 4
    add   r9, -2
    jne   .LBB0_8

There is no division instruction left. The % became a multiplication by a precomputed constant and a few shifts, which is much faster, and the loop handles two iterations per pass. An ahead-of-time compiler can spend seconds finding tricks like this, because it runs only once, before the program ships.

Ahead-of-time compilation gives fast startup, predictable speed, and a program that needs no runtime to run. The costs: one binary per platform, a compile step between each edit and each run, and no knowledge of how the program will actually be used.

Interpretation and bytecode

An interpreter doesn't translate your program into machine code. It is a machine-code program that carries yours out. Most modern interpreters don't work on the source text directly: they first compile it to bytecode, a compact instruction set for an imaginary machine, then run a loop that fetches and executes one bytecode instruction at a time.

This is what CPython 3.14 compiles the same loop to (from dis.dis):

  L1:     FOR_ITER                18 (to L2)
          STORE_FAST               2 (i)
          LOAD_FAST_BORROW_LOAD_FAST_BORROW 18 (s, i)
          LOAD_SMALL_INT           2
          BINARY_OP                5 (*)
          BINARY_OP               13 (+=)
          STORE_FAST               1 (s)
          JUMP_BACKWARD           20 (to L1)

It's a stack machine: LOAD pushes values, BINARY_OP pops two and pushes the result. Each bytecode instruction is handled by a piece of C code inside the interpreter. That C code decodes the instruction, checks the types of its operands — * means something different for integers, floats and strings — works with integers stored as heap objects, and updates reference counts. A single BINARY_OP can cost dozens of machine instructions where the compiled C loop needed one or two.

Interpreters start instantly, run the same bytecode on every platform, and make dynamic features — changing types, eval, redefining functions at run time — easy. The price is speed.

Just-in-time compilation

A JIT combines both. V8, the JavaScript engine in Chrome and Node.js, first compiles a function to bytecode for its Ignition interpreter. The same loop, written in JavaScript:

LdaZero
Star0                       ; s = 0
LdaZero
Star1                       ; i = 0
Ldar a0
TestLessThan r1, [0]        ; i < n ?
JumpIfFalse [23]
Ldar r1
MulSmi [2], [2]             ; i * 2
Add r0, [1]                 ; + s
…
Inc [3]                     ; i++
JumpLoop [24], [0], [4]

The numbers in brackets are feedback slots. As the interpreter runs, it records what it sees at each operation — "this Add has always been two small integers" — and counts how often each function and loop runs. When code gets hot, V8 compiles it to machine code in the background. With node --trace-opt, you can watch it happen to this loop:

[marking … <JSFunction total> for optimization to MAGLEV, … reason: hot and stable]
[completed compiling … <JSFunction total> (target MAGLEV) OSR - took … 0.094 … ms]
[completed compiling … <JSFunction total> (target TURBOFAN_JS) OSR - took … 0.725 … ms]

It uses two tiers: Maglev, a fast mid-level compiler, then Turbofan, the full optimizing compiler. OSR (on-stack replacement) means V8 swapped the loop that was already running in the interpreter for the compiled version, without waiting for the function to be called again.

The optimized code is built on speculation. Because the feedback said s and i were always small integers, Turbofan generates plain integer machine instructions, with a quick check that the assumption still holds. If a string or a huge number ever shows up, the check fails and the engine deoptimizes: it throws the compiled code away and returns to the interpreter, which handles every case. An ahead-of-time compiler can't make such bets, but a JIT can, because it can always fall back.

The same loop, three ways

The same computation — 50 million iterations of s = (s + i * 2) % 1000003 — on the Apple M2 Ultra this chapter was written on. All four versions print the same result, 22650:

ImplementationTime
C, compiled ahead of time (clang -O2)0.17 s
JavaScript, V8 with its JIT (Node.js 25)0.25 s
JavaScript, V8 interpreter only (node --jitless)0.67 s
Python, CPython 3.14 bytecode interpreter3.3–3.7 s

The JIT brings JavaScript within about 1.5× of C on this loop. The same engine limited to its interpreter is almost 3× slower. CPython, which runs its bytecode without a JIT (3.13 added an experimental one, off by default), takes about 20× as long as C. These numbers are for one tight numeric loop on one machine: real programs spend time in I/O, libraries and memory, and the gaps are usually much smaller. Numeric Python code avoids the interpreter by calling libraries like NumPy, whose inner loops are compiled C.

Choosing, and blurring the lines

Ahead of timeInterpreterJIT
Startupinstantinstantfast, then warms up
Peak speedhighestlowestclose to AOT on hot code
Portabilityone binary per platformsame code everywheresame code everywhere
Memorysmallestsmalllarger (compiler, code, profiles)
What a reverse engineer seesstripped machine codebytecode, easy to decompilebytecode plus code generated at run time

Real systems mix the three:

  • Java compiles to bytecode ahead of time (.class files), then the JVM interprets and JIT-compiles it. GraalVM's native-image can instead compile it all ahead of time.
  • Android compiles app bytecode partly ahead of time at installation and JIT-compiles the rest.
  • WebAssembly is a portable, low-level bytecode that browsers compile to machine code as it loads. This site's CPU emulator is written in Rust, compiled to WebAssembly, and turned into native code by your browser.

The line between compiling and interpreting is less a wall than a dial: how much translation happens before the program runs, and how much while it runs. The next chapter, from source code to a running binary, follows the ahead-of-time path step by step.

Takeaways

  • The CPU runs only machine code. Languages get there by compiling ahead of time, interpreting, or compiling just in time, a property of the implementation, not the language.
  • Optimizing compilers transform code deeply before it ships: a % becomes a multiply and shifts, and loops get unrolled.
  • Interpreters usually run bytecode for a virtual stack machine. Each instruction goes through type checks and dispatch, so they're flexible, portable and slow.
  • JITs profile the program and compile hot code while it runs, speculating on types and deoptimizing when a guess fails. V8 goes through Ignition → Maglev → Turbofan, even swapping running loops (OSR).
  • Measured on one loop: C 0.17 s, JavaScript with its JIT 0.25 s, the same engine interpreting 0.67 s, CPython 3.3–3.7 s.

In this level

  1. 2.1Compiled, interpreted and JIT-compiled languages
  2. 2.2From source code to a running binary
  3. 2.3Variables, types and memory layoutPlanned
  4. 2.4Pointers and arraysPlanned
  5. 2.5Functions, calls and the stackPlanned
  6. 2.6Structs, malloc and the heapPlanned
  7. 2.7Control flow: if, loops, switchPlanned
  8. 2.8Bytecode virtual machines and garbage collectionPlanned