Skip to content

Level 5 · Chapter 5.7

VLIW, EPIC and the Itanium

The road not taken: instruction sets that expose the machine's parallelism to the compiler — VLIW, and Intel and HP's EPIC architecture, IA-64 — with bundles, predication and speculative loads, a live comparison of compiler scheduling against out-of-order hardware when a load misses the cache, and why the Itanium was discontinued while its ideas survived elsewhere.

The microarchitecture level shows how a modern core finds parallelism on its own: it looks dozens of instructions ahead, renames registers, starts whatever is ready, and retires everything in order. That hardware is large and power-hungry. Since the 1980s, a recurring idea has been to do that work once, in the compiler, instead of billions of times a second in silicon. An ISA built on that idea has to say more than "run these instructions in order": it has to let the compiler state which instructions can run together.

That's VLIW (very long instruction word), and its most ambitious version, EPIC, was Intel and HP's bet to replace x86. Tanenbaum's book presents it at length and with enthusiasm. It's worth studying both for the ideas, and for what happened next.

VLIW: the compiler schedules

A VLIW instruction is a bundle of several operations, one per functional unit, all issued in the same cycle. The compiler guarantees they're independent. The hardware doesn't check dependences, doesn't reorder, doesn't rename: it just sends each slot to its unit. If the compiler finds nothing useful for a slot, it fills it with a no-op.

The payoff is simple, wide hardware. The costs are code size (all those no-ops), and a tight coupling between the binary and the chip: a program scheduled for a machine with two multipliers and a 3-cycle load is wrong, or at least slow, on a machine with three multipliers and a 4-cycle load. Early VLIW machines, like Multiflow's in the 1980s, required recompiling for every new model.

EPIC and IA-64

Intel and HP's IA-64, first shipped as the Itanium in 2001, called its version EPIC: explicitly parallel instruction computing. It kept the VLIW idea — the compiler finds the parallelism — but tried to fix VLIW's weaknesses:

  • Bundles of 128 bits hold three 41-bit instructions and a 5-bit template. The template says which kinds of units the three need, and where instruction groups end: runs of instructions, possibly spanning several bundles, that the compiler promises are independent. Because groups aren't tied to the machine's width, the same binary runs on wider or narrower implementations.
  • Lots of registers: 128 integer and 128 floating-point registers, so the compiler never runs short and never needs renaming. 96 of the integer registers form a register stack: each function allocates exactly the registers it needs, and calls shift the window instead of saving registers to memory.
  • Predication: 64 one-bit predicate registers, and every instruction names one. A comparison sets a predicate and its complement; the "then" and "else" instructions then run side by side, each guarded by its predicate, and only the true side writes its result. The branch disappears, and so does the risk of mispredicting it.
  • Speculative loads: the compiler can move a load above the branch that guards it (ld.s). If the load would fault — say, because the pointer is null on the path not taken — it doesn't: it marks its destination register as invalid (a "NaT" bit, the poison bit of the out-of-order chapter), and a later check instruction (chk.s) raises the fault only if the value is actually used. Advanced loads (ld.a) go further and move a load above a store that might write the same address; a hardware table watches the stores, and a check reloads the value if one of them hit.

The book's section on IA-64 argues that x86 had run out of road, and concludes that moving work from run time to compile time "is always a win". History went the other way.

Where compile-time scheduling breaks

The two demos below run the same work — three loads and the arithmetic on them — on the out-of-order simulator, whose in order setting behaves like a statically scheduled machine: it never starts an instruction before the ones in front of it.

In the first, the instructions are in source order: each load is followed by the instruction that uses it.

Out of order · Loads in source order

Try it: Switch between in order and out of order, change the width, window, load latency (4 = L1 hit, 12 = L2, 40 = L3) and renaming below, and compare the cycle counts — or press Edit to run your own code.

renaming

Loading emulator…

In order, each load waits behind the use of the previous one: 19 cycles when loads hit the L1 cache (4-cycle latency). Now set the load latency to 40, a trip to the L3 cache: 127 cycles in order, because the three loads wait one after another. Out of order, the hardware starts all three loads at once: 12 and 48 cycles.

In the second, the compiler has done the out-of-order core's job in advance, moving the three independent loads to the top:

Out of order · The same work, scheduled by the compiler

Try it: Switch between in order and out of order, change the width, window, load latency (4 = L1 hit, 12 = L2, 40 = L3) and renaming below, and compare the cycle counts — or press Edit to run your own code.

renaming

Loading emulator…

Now in-order and out-of-order take exactly the same time at every latency: 12, 20 and 48 cycles for 4, 12 and 40. That's the EPIC premise, and here it holds.

It holds because this code is easy: three loads from fixed addresses, no branches, no pointers. Real programs rarely let the compiler do this:

  • Pointers may alias. If a store through p sits between the loads and the compiler can't prove p points elsewhere, the loads can't move above it. Advanced loads help, at the cost of check instructions and recovery code.
  • Branches block motion. Loads behind an if can only move up speculatively, with the same check-and-recover machinery.
  • Latency is unknown. Whether a load takes 4 cycles or 300 depends on the cache, which depends on the data and on what else is running. The compiler must pick one schedule for all cases; out-of-order hardware adapts on every execution, to the latency that actually happens.

The out-of-order core also runs old binaries at full speed, and it keeps working when the next generation changes every latency. The x86 cores that the book expected to hit a wall kept getting faster, and in 2003 AMD extended x86 to 64 bits: the same design Tanenbaum mentions as Intel's "EMT-64", adopted by Intel in 2004. With 64-bit x86 available, the Itanium's main selling point was gone. It stayed a niche in high-end servers, mostly HP's. Intel shipped its last Itanium processors in 2021, and Linux removed IA-64 support in 2024.

What survived

The ideas didn't die with the chip:

  • VLIW lives in DSPs and accelerators, where code is small, hand-tuned and runs on one known chip: Texas Instruments' C6000 DSPs, and Qualcomm's Hexagon, the VLIW DSP in its Snapdragon phone chips. AMD's Radeon GPUs used VLIW designs until 2012.
  • Predication is everywhere, in smaller doses. The cmov and csel of the previous chapter are one-instruction predication. AVX-512 has mask registers that predicate every vector lane, ARM's SVE has predicate registers for the same purpose, and GPUs run divergent branches by predicating each thread.
  • Speculation with deferred faults became a hardware feature of out-of-order cores, which execute past branches and cancel what turns out wrong.
  • Compile-once, translate-later returned with binary translation: Transmeta's VLIW processors ran x86 code by translating it at run time, as Apple's Rosetta does today from x86 to ARM64.

Takeaways

  • A VLIW ISA issues bundles of operations the compiler guarantees are independent: simple hardware, larger code, binaries tied to one machine.
  • EPIC / IA-64 (Itanium, 2001) added 128-bit bundles with templates and instruction groups, 128 registers with a register stack, full predication with 64 predicate registers, and speculative and advanced loads.
  • Compile-time scheduling works when the compiler can see the parallelism: moved to the top, three loads run as fast in order as out of order (12, 20, 48 cycles).
  • It fails on aliasing, branches and unpredictable latencies, which out-of-order hardware handles at run time, on any binary. In source order, the in-order machine took 127 cycles against 48 out of order at L3 latency.
  • AMD's 64-bit x86 removed the Itanium's reason to exist; it was discontinued in 2021. VLIW survives in DSPs, predication in SIMD and GPUs.

In this level

  1. 5.1What the ISA promises: memory model, alignment and ordering
  2. 5.2Instruction formats and encoding
  3. 5.3Addressing modes
  4. 5.4x86, ARM and RISC-V side by side
  5. 5.5Traps, interrupts and exceptions
  6. 5.6Data types the hardware understands
  7. 5.7VLIW, EPIC and the Itanium