A CPU executes only its own machine code. A program written in C, Python or JavaScript has to get there somehow, and languages take three main routes:
- Compile ahead of time (AOT): translate the whole program to machine code before it runs. C, C++, Rust, Go and Swift work this way.
- Interpret: run a program that reads your program and carries it out, step by step. CPython works this way, via bytecode.
- Compile just in time (JIT): start by interpreting, watch what the program actually does, and compile the hot parts to machine code while it runs. JavaScript engines, the Java and .NET virtual machines, PyPy and LuaJIT work this way.
The route is a property of the implementation, not of the language: there are C interpreters and ahead-of-time Java compilers. But each language has a usual route, and it shapes how fast programs start, how fast they run, and what's left for someone reading the binary.
Ahead-of-time compilation
A compiler reads the whole program, checks it, optimizes it, and writes machine code. This site's simulator has a small C compiler built in. The demo below compiles a C function live and opens at the C level. Zoom in to see the assembly it produced, and further down to the bytes and micro-operations:
- int total(int n) {
- int s = 0;
- for (int i = 0; i < n; i++)
- s = s + i * 2;
- return s;
- }
- int main() {
- return total(5);
- }
The program returns 20 after 81 instructions. The simulator's compiler is deliberately simple, like gcc -O0: every variable lives on the stack, so the code maps line by line to the C.
A production compiler with optimizations does far more. Here is the loop s = (s + i * 2) % 1000003 compiled by clang with -O2 for x86-64, trimmed to the core of the loop:
.LBB0_8:
add rsi, rcx ; s += i*2 (rcx holds i*2)
mov rax, rsi
imul r8 ; multiply by a "magic" constant…
add rdx, rsi
mov rax, rdx
shr rax, 63
sar rdx, 19 ; …and shift: rdx = s / 1000003
add rdx, rax
imul rax, rdx, 1000003
sub rsi, rax ; s -= (s / 1000003) * 1000003, i.e. s %= 1000003
… ; the same again for the next i: the loop is unrolled twice
add rcx, 4
add r9, -2
jne .LBB0_8
There is no division instruction left. The % became a multiplication by a precomputed constant and a few shifts, which is much faster, and the loop handles two iterations per pass. An ahead-of-time compiler can spend seconds finding tricks like this, because it runs only once, before the program ships.
Ahead-of-time compilation gives fast startup, predictable speed, and a program that needs no runtime to run. The costs: one binary per platform, a compile step between each edit and each run, and no knowledge of how the program will actually be used.
Interpretation and bytecode
An interpreter doesn't translate your program into machine code. It is a machine-code program that carries yours out. Most modern interpreters don't work on the source text directly: they first compile it to bytecode, a compact instruction set for an imaginary machine, then run a loop that fetches and executes one bytecode instruction at a time.
This is what CPython 3.14 compiles the same loop to (from dis.dis):
L1: FOR_ITER 18 (to L2)
STORE_FAST 2 (i)
LOAD_FAST_BORROW_LOAD_FAST_BORROW 18 (s, i)
LOAD_SMALL_INT 2
BINARY_OP 5 (*)
BINARY_OP 13 (+=)
STORE_FAST 1 (s)
JUMP_BACKWARD 20 (to L1)
It's a stack machine: LOAD pushes values, BINARY_OP pops two and pushes the result. Each bytecode instruction is handled by a piece of C code inside the interpreter. That C code decodes the instruction, checks the types of its operands — * means something different for integers, floats and strings — works with integers stored as heap objects, and updates reference counts. A single BINARY_OP can cost dozens of machine instructions where the compiled C loop needed one or two.
Interpreters start instantly, run the same bytecode on every platform, and make dynamic features — changing types, eval, redefining functions at run time — easy. The price is speed.
Just-in-time compilation
A JIT combines both. V8, the JavaScript engine in Chrome and Node.js, first compiles a function to bytecode for its Ignition interpreter. The same loop, written in JavaScript:
LdaZero
Star0 ; s = 0
LdaZero
Star1 ; i = 0
Ldar a0
TestLessThan r1, [0] ; i < n ?
JumpIfFalse [23]
Ldar r1
MulSmi [2], [2] ; i * 2
Add r0, [1] ; + s
…
Inc [3] ; i++
JumpLoop [24], [0], [4]
The numbers in brackets are feedback slots. As the interpreter runs, it records what it sees at each operation — "this Add has always been two small integers" — and counts how often each function and loop runs. When code gets hot, V8 compiles it to machine code in the background. With node --trace-opt, you can watch it happen to this loop:
[marking … <JSFunction total> for optimization to MAGLEV, … reason: hot and stable]
[completed compiling … <JSFunction total> (target MAGLEV) OSR - took … 0.094 … ms]
[completed compiling … <JSFunction total> (target TURBOFAN_JS) OSR - took … 0.725 … ms]
It uses two tiers: Maglev, a fast mid-level compiler, then Turbofan, the full optimizing compiler. OSR (on-stack replacement) means V8 swapped the loop that was already running in the interpreter for the compiled version, without waiting for the function to be called again.
The optimized code is built on speculation. Because the feedback said s and i were always small integers, Turbofan generates plain integer machine instructions, with a quick check that the assumption still holds. If a string or a huge number ever shows up, the check fails and the engine deoptimizes: it throws the compiled code away and returns to the interpreter, which handles every case. An ahead-of-time compiler can't make such bets, but a JIT can, because it can always fall back.
The same loop, three ways
The same computation — 50 million iterations of s = (s + i * 2) % 1000003 — on the Apple M2 Ultra this chapter was written on. All four versions print the same result, 22650:
| Implementation | Time |
|---|---|
C, compiled ahead of time (clang -O2) | 0.17 s |
| JavaScript, V8 with its JIT (Node.js 25) | 0.25 s |
JavaScript, V8 interpreter only (node --jitless) | 0.67 s |
| Python, CPython 3.14 bytecode interpreter | 3.3–3.7 s |
The JIT brings JavaScript within about 1.5× of C on this loop. The same engine limited to its interpreter is almost 3× slower. CPython, which runs its bytecode without a JIT (3.13 added an experimental one, off by default), takes about 20× as long as C. These numbers are for one tight numeric loop on one machine: real programs spend time in I/O, libraries and memory, and the gaps are usually much smaller. Numeric Python code avoids the interpreter by calling libraries like NumPy, whose inner loops are compiled C.
Choosing, and blurring the lines
| Ahead of time | Interpreter | JIT | |
|---|---|---|---|
| Startup | instant | instant | fast, then warms up |
| Peak speed | highest | lowest | close to AOT on hot code |
| Portability | one binary per platform | same code everywhere | same code everywhere |
| Memory | smallest | small | larger (compiler, code, profiles) |
| What a reverse engineer sees | stripped machine code | bytecode, easy to decompile | bytecode plus code generated at run time |
Real systems mix the three:
- Java compiles to bytecode ahead of time (
.classfiles), then the JVM interprets and JIT-compiles it. GraalVM'snative-imagecan instead compile it all ahead of time. - Android compiles app bytecode partly ahead of time at installation and JIT-compiles the rest.
- WebAssembly is a portable, low-level bytecode that browsers compile to machine code as it loads. This site's CPU emulator is written in Rust, compiled to WebAssembly, and turned into native code by your browser.
The line between compiling and interpreting is less a wall than a dial: how much translation happens before the program runs, and how much while it runs. The next chapter, from source code to a running binary, follows the ahead-of-time path step by step.
Takeaways
- The CPU runs only machine code. Languages get there by compiling ahead of time, interpreting, or compiling just in time, a property of the implementation, not the language.
- Optimizing compilers transform code deeply before it ships: a
%becomes a multiply and shifts, and loops get unrolled. - Interpreters usually run bytecode for a virtual stack machine. Each instruction goes through type checks and dispatch, so they're flexible, portable and slow.
- JITs profile the program and compile hot code while it runs, speculating on types and deoptimizing when a guess fails. V8 goes through Ignition → Maglev → Turbofan, even swapping running loops (OSR).
- Measured on one loop: C 0.17 s, JavaScript with its JIT 0.25 s, the same engine interpreting 0.67 s, CPython 3.3–3.7 s.