Skip to content

Level 6 · Chapter 6.4

A complete machine: the Mic-1 running IJVM

A classic teaching microarchitecture, end to end: the Mic-1 datapath, the IJVM stack machine it interprets, how IADD, ILOAD, GOTO and INVOKEVIRTUAL become sequences of microinstructions (run live in C), and how the Mic-2, Mic-3 and Mic-4 speed it up.

The microprogramming chapter explained the idea: a control store full of microinstructions drives the datapath one cycle at a time. This chapter follows one complete machine from the registers up. The machine is the Mic-1, a classic teaching microarchitecture. The instruction set it runs is IJVM, a small integer subset of Java bytecode used for teaching. The whole interpreter is 112 microinstructions long.

It's worth studying in detail for one reason: it's small enough to hold in your head, yet it contains every mechanism a real microcoded CPU needs: instruction fetch, operand decoding, a stack, memory latency, branches and procedure calls.

The Mic-1 datapath

The Mic-1 has ten 32-bit registers (MBR is 8 bits), one ALU and a shifter, connected by two buses:

RegisterHolds
MAR, MDRaddress and data for the 32-bit word memory port
PC, MBRaddress and data for the 8-bit byte port, used only to fetch the program
SPaddress of the top word of the stack
LVaddress of the current method's local variables
CPPaddress of the constant pool
TOSa copy of the word at the top of the stack
OPCscratch; often the address of the current opcode
Hthe "holding" register: the ALU's only left input

On each cycle, one register drives the B bus, which is the ALU's right input. The left input is always H. The result goes through the shifter onto the C bus, and can be written into any number of registers at once. So MAR = SP = SP − 1 is one cycle: SP on the B bus, the ALU computes B − 1, and the result lands in both MAR and SP.

There is no way to add two arbitrary registers in one cycle. One of them must first be copied into H. That limitation shows up in almost every instruction below.

The two memory ports behave differently. MAR counts in words: MAR = 2 reads bytes 8–11. PC counts in bytes, because IJVM instructions are a stream of bytes of different lengths. The byte from the PC port lands in MBR, and MBR can go onto the B bus two ways: sign-extended (written MBR) or zero-extended (MBRU).

Memory is slow, even in this idealized machine. A read started at the end of cycle k delivers its data at the end of cycle k + 1, so the value can only be used in cycle k + 2. The microprogram has to fill that gap with useful work, or waste a cycle.

Each microinstruction is 36 bits: a 9-bit next address, 3 JAM bits for branching, 8 ALU and shifter bits, 9 bits choosing which registers the C bus writes, 3 memory bits (read, write, fetch) and a 4-bit B-bus selector. The control store holds 512 of them. The microprogramming chapter covers how the next address is formed; here we look at what the microinstructions actually do.

IJVM: the machine being interpreted

IJVM is a stack machine. Its instructions don't name registers: operands are pushed on a stack, and operations pop them and push the result. Memory is divided into four areas, each reached through its own pointer:

  • the constant pool, read-only, at CPP. It holds constants and the addresses of methods;
  • the local variable frame of the current method, at LV. Parameters come first, then local variables;
  • the operand stack, directly above the frame, whose top is at SP;
  • the method area, the program's bytecode, read byte by byte through PC.

A program never sees an absolute data address. It can only say "local variable 3" or "constant 0", and the microcode adds the offset to LV or CPP.

The instruction set has 20 instructions. Here they are, with the same opcode numbers as the real JVM:

HexMnemonicEffect
10BIPUSH bytepush a signed byte
13LDC_W indexpush a word from the constant pool
15ILOAD varnumpush a local variable
36ISTORE varnumpop into a local variable
84IINC varnum constadd a signed byte to a local variable
60 64IADD, ISUBpop two words, push the sum or difference
7E 80IAND, IORpop two words, push the AND or OR
59 57 5FDUP, POP, SWAPcopy, drop or swap stack words
A7GOTO offsetunconditional branch
99 9BIFEQ, IFLT offsetpop a word, branch if zero or negative
9FIF_ICMPEQ offsetpop two words, branch if equal
B6INVOKEVIRTUAL dispcall a method
ACIRETURNreturn an integer
C4WIDEprefix: the next ILOAD or ISTORE has a 16-bit index
00NOPnothing

Branch offsets are signed 16-bit numbers, big-endian, relative to the address of the branch opcode itself.

From Java to IJVM

Here is a Java method that multiplies by repeated addition:

int mul(int a, int b) {
    int r = 0;
    while (b != 0) {
        r = r + a;
        b = b - 1;
    }
    return r;
}

This is what javac from OpenJDK 26 produces for it, shown by javap -c:

 0: iconst_0
 1: istore_3
 2: iload_2
 3: ifeq          17
 6: iload_3
 7: iload_1
 8: iadd
 9: istore_3
10: iload_2
11: iconst_1
12: isub
13: istore_2
14: goto          2
17: iload_3
18: ireturn

The raw bytes in the .class file include 99 00 0e (ifeq +14), 60 (iadd), 64 (isub), a7 ff f4 (goto −12) and ac (ireturn): the IJVM opcodes exactly. The difference is that the real JVM has one-byte short forms for common cases: iconst_0 instead of BIPUSH 0, istore_3 instead of ISTORE 3. IJVM leaves them out and always uses the general form, so the IJVM version is a few bytes longer.

Local variable 0 is this, the object the method was called on; a, b and r are variables 1, 2 and 3. javap -v reports locals=4, args_size=3.

The main loop: one microinstruction

Every IJVM instruction starts with the same microinstruction, Main1:

Main1   PC = PC + 1; fetch; goto (MBR)

It does three things in one cycle. It advances PC past the opcode. It starts fetching the next byte, which will be either the instruction's first operand or the next opcode. And it jumps to the control-store address equal to the opcode sitting in MBR: the multiway branch the microprogramming chapter described. The routine for IADD starts at control-store address 0x60, the one for ILOAD at 0x15.

That's why each microinstruction names its successor explicitly. Addresses 0x00 to 0xFF are reserved for the first microinstruction of each opcode, so the rest of each routine has to live in the gaps, scattered across the control store.

The invariant that makes this work: when control returns to Main1, the next opcode must already be in MBR. Every routine is written to guarantee it.

IADD: three microinstructions

iadd1   MAR = SP = SP − 1; rd       read the word below the top
iadd2   H = TOS                     top of stack into H, while the read completes
iadd3   MDR = TOS = MDR + H; wr; goto Main1

TOS already holds the top of the stack, so only one memory read is needed. iadd1 points SP and MAR at the second word and starts the read. The data won't be usable until two cycles later, so iadd2 uses the gap to move TOS into H. In iadd3 the word has arrived in MDR. The sum goes to MDR (to be written back to memory) and to TOS (to keep the copy current), and the write starts.

With Main1, IADD costs 4 microinstructions. ISUB, IAND and IOR are the same routine with a different ALU operation.

ILOAD: waiting for memory twice

iload1  H = LV                      LV into H, to be added to the index
iload2  MAR = MBRU + H; rd          address of the local variable; read it
iload3  MAR = SP = SP + 1           new top of stack; wait for the read
iload4  PC = PC + 1; fetch; wr      fetch the next opcode; write the value on the stack
iload5  TOS = MDR; goto Main1

The index byte was fetched by Main1, and it's zero-extended (MBRU): a variable number is never negative. LV has to go through H first, because MBR and LV are both on the B bus side. Then it's read, pushed, and copied into TOS. 6 microinstructions, counting Main1.

GOTO: assembling a 16-bit offset

goto1   OPC = PC − 1                save the address of the opcode
goto2   PC = PC + 1; fetch          fetch the second offset byte
goto3   H = MBR << 8                high byte, sign-extended, shifted into place
goto4   H = MBRU OR H               low byte, zero-extended, OR-ed in
goto5   PC = OPC + H; fetch         jump, and fetch the opcode there
goto6   goto Main1                  wait for that fetch to arrive

The offset is relative to the opcode, but PC has already moved past it, so goto1 recovers the opcode's address into OPC. The two offset bytes arrive one at a time through the 8-bit port. The high byte is sign-extended, because offsets can be negative, and the low byte is not. goto6 does nothing: it only waits for the new opcode to reach MBR, as Main1 requires. 7 microinstructions in total.

The conditional branches reuse this code. IFEQ pops the top of stack, reads the new top into TOS, and passes the popped value through the ALU to set the Z flag. That microinstruction's JAMZ bit makes the next address depend on Z. If the branch is taken, it jumps into the middle of the GOTO routine; if not, three short microinstructions skip over the offset. IFEQ costs 11 microinstructions when taken and 8 when not.

INVOKEVIRTUAL: building a frame

The call instruction is by far the longest: 22 microinstructions after Main1. Its operand is an index into the constant pool, where the method's address is stored. The method itself starts with two 16-bit numbers, its number of parameters (counting the object reference) and its number of local variables, followed by its first opcode.

The microcode fetches all of that one byte at a time, then builds the new frame. Here is the stack in the demo below, right after INVOKEVIRTUAL calls mul(6, 7):

word 13   caller's LV = 8           ← SP
word 12   caller's PC = 9
word 11   r                         (local 3)
word 10   b = 7                     (local 2)
word  9   a = 6                     (local 1)
word  8   link pointer = 12         ← LV   (was the object reference)

The caller had pushed the object reference and the two arguments. The microcode points LV at the object reference and overwrites it with a link pointer: the address, above the local variables, where it saves the caller's PC and LV. The new method's operand stack starts just above them.

IRETURN undoes this in 8 microinstructions. It follows the link pointer to restore PC and LV, and writes the return value where the object reference used to be, so the caller finds the result on top of its stack.

Running the Mic-1

The demo below is the Mic-1 written in C. The global variables are the Mic-1's registers, and every line that ends in cycle() is one microinstruction of the Mic-1 microprogram, with its label in the comment. cycle() models the memory timing: a read started in one cycle only reaches MDR or MBR at the end of the next one. The IJVM program is the mul method above, hand-assembled, called as mul(6, 7). (IJVM has no instruction to stop the machine, so the demo adds a HALT opcode, 0xFF, to its main program.)

Live · A Mic-1 interpreting IJVM, one microinstruction per cycle()

Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.

C source · click a line number for a breakpoint
  1. // A Mic-1 running IJVM: every line ending in cycle() is one microinstruction
  2. // of the Mic-1 microprogram, acting on the Mic-1's registers.
  3. int MAR, MDR, PC, MBR, SP, LV, CPP, TOS, OPC, H;
  4. int N, Z; // ALU flags latched by the last cycle
  5. int mem[32]; // word memory: constant pool and stack
  6. unsigned char code[] = { // method area (byte addressed)
  7. 0x10, 0, // 0 BIPUSH 0 OBJREF (unused)
  8. 0x10, 6, // 2 BIPUSH 6 a
  9. 0x10, 7, // 4 BIPUSH 7 b
  10. 0xb6, 0, 0, // 6 INVOKEVIRTUAL 0 mul(a, b)
  11. 0xff, // 9 HALT (not in IJVM: stops the demo)
  12. 0, 3, 0, 1, // 10 mul: 3 parameters (OBJREF, a, b), 1 local (r)
  13. 0x10, 0, // 14 BIPUSH 0
  14. 0x36, 3, // 16 ISTORE 3 r = 0
  15. 0x15, 2, // 18 ILOAD 2 loop:
  16. 0x99, 0, 20, // 20 IFEQ +20 if b == 0 goto done
  17. 0x15, 3, // 23 ILOAD 3
  18. 0x15, 1, // 25 ILOAD 1
  19. 0x60, // 27 IADD
  20. 0x36, 3, // 28 ISTORE 3 r = r + a
  21. 0x15, 2, // 30 ILOAD 2
  22. 0x10, 1, // 32 BIPUSH 1
  23. 0x64, // 34 ISUB
  24. 0x36, 2, // 35 ISTORE 2 b = b - 1
  25. 0xa7, 0xff, 0xed, // 37 GOTO -19 goto loop
  26. 0x15, 3, // 40 ILOAD 3 done:
  27. 0xac // 42 IRETURN return r
  28. };
  29. int rd_now, rd_wait, rd_addr, fetch_now, fetch_wait, fetch_addr;
  30. int cycles, count, shown[256];
  31. void rd() { rd_now = 1; }
  32. void wr() { mem[MAR] = MDR; }
  33. void fetch() { fetch_now = 1; }
  34. int MBRU() { return MBR; } // zero-extended
  35. int MBRS() { return (signed char)MBR; } // sign-extended
  36. // End of a clock cycle: a read started in the previous cycle arrives now,
  37. // so its data can be used by the microinstruction after next.
  38. void cycle() {
  39. if (rd_wait) { MDR = mem[rd_addr]; rd_wait = 0; }
  40. if (fetch_wait) { MBR = code[fetch_addr]; fetch_wait = 0; }
  41. if (rd_now) { rd_addr = MAR; rd_wait = 1; rd_now = 0; }
  42. if (fetch_now) { fetch_addr = PC; fetch_wait = 1; fetch_now = 0; }
  43. cycles++;
  44. }
  45. char *name(int op) {
  46. switch (op) {
  47. case 0x10: return "BIPUSH"; case 0x15: return "ILOAD";
  48. case 0x36: return "ISTORE"; case 0x60: return "IADD";
  49. case 0x64: return "ISUB"; case 0x99: return "IFEQ";
  50. case 0xa7: return "GOTO"; case 0xac: return "IRETURN";
  51. case 0xb6: return "INVOKEVIRTUAL";
  52. }
  53. return "?";
  54. }
  55. void alu(int v) { N = v < 0; Z = v == 0; }
  56. int main() {
  57. CPP = 0; mem[0] = 10; // constant pool: address of mul
  58. LV = 8; SP = 7; // empty stack for the main program
  59. PC = 0; MBR = code[0]; // invariant: opcode already in MBR
  60. while (1) {
  61. int op = MBR, start = cycles;
  62. if (op == 0xff) break;
  63. PC = PC + 1; fetch(); cycle(); // Main1: goto (MBR)
  64. switch (op) {
  65. case 0x10: // BIPUSH
  66. SP = MAR = SP + 1; cycle(); // bipush1
  67. PC = PC + 1; fetch(); cycle(); // bipush2
  68. MDR = TOS = MBRS(); wr(); cycle(); // bipush3
  69. break;
  70. case 0x15: // ILOAD
  71. H = LV; cycle(); // iload1
  72. MAR = MBRU() + H; rd(); cycle(); // iload2
  73. MAR = SP = SP + 1; cycle(); // iload3
  74. PC = PC + 1; fetch(); wr(); cycle(); // iload4
  75. TOS = MDR; cycle(); // iload5
  76. break;
  77. case 0x36: // ISTORE
  78. H = LV; cycle(); // istore1
  79. MAR = MBRU() + H; cycle(); // istore2
  80. MDR = TOS; wr(); cycle(); // istore3
  81. SP = MAR = SP - 1; rd(); cycle(); // istore4
  82. PC = PC + 1; fetch(); cycle(); // istore5
  83. TOS = MDR; cycle(); // istore6
  84. break;
  85. case 0x60: // IADD
  86. case 0x64: // ISUB
  87. MAR = SP = SP - 1; rd(); cycle(); // iadd1
  88. H = TOS; cycle(); // iadd2
  89. if (op == 0x60) MDR = TOS = MDR + H; // iadd3
  90. else MDR = TOS = MDR - H; // (isub3)
  91. wr(); cycle();
  92. break;
  93. case 0x99: // IFEQ
  94. MAR = SP = SP - 1; rd(); cycle(); // ifeq1
  95. OPC = TOS; cycle(); // ifeq2
  96. TOS = MDR; cycle(); // ifeq3
  97. alu(OPC); cycle(); // ifeq4: Z = OPC
  98. if (!Z) {
  99. PC = PC + 1; cycle(); // F
  100. PC = PC + 1; fetch(); cycle(); // F2
  101. cycle(); // F3
  102. break;
  103. }
  104. OPC = PC - 1; cycle(); // T, then goto2
  105. PC = PC + 1; fetch(); cycle(); // goto2
  106. H = MBRS() << 8; cycle(); // goto3
  107. H = MBRU() | H; cycle(); // goto4
  108. PC = OPC + H; fetch(); cycle(); // goto5
  109. cycle(); // goto6
  110. break;
  111. case 0xa7: // GOTO
  112. OPC = PC - 1; cycle(); // goto1
  113. PC = PC + 1; fetch(); cycle(); // goto2
  114. H = MBRS() << 8; cycle(); // goto3
  115. H = MBRU() | H; cycle(); // goto4
  116. PC = OPC + H; fetch(); cycle(); // goto5
  117. cycle(); // goto6
  118. break;
  119. case 0xb6: // INVOKEVIRTUAL
  120. PC = PC + 1; fetch(); cycle(); // 1
  121. H = MBRU() << 8; cycle(); // 2
  122. H = MBRU() | H; cycle(); // 3
  123. MAR = CPP + H; rd(); cycle(); // 4
  124. OPC = PC + 1; cycle(); // 5
  125. PC = MDR; fetch(); cycle(); // 6
  126. PC = PC + 1; fetch(); cycle(); // 7
  127. H = MBRU() << 8; cycle(); // 8
  128. H = MBRU() | H; cycle(); // 9
  129. PC = PC + 1; fetch(); cycle(); // 10
  130. TOS = SP - H; cycle(); // 11
  131. TOS = MAR = TOS + 1; cycle(); // 12
  132. PC = PC + 1; fetch(); cycle(); // 13
  133. H = MBRU() << 8; cycle(); // 14
  134. H = MBRU() | H; cycle(); // 15
  135. MDR = SP + H + 1; wr(); cycle(); // 16
  136. MAR = SP = MDR; cycle(); // 17
  137. MDR = OPC; wr(); cycle(); // 18
  138. MAR = SP = SP + 1; cycle(); // 19
  139. MDR = LV; wr(); cycle(); // 20
  140. PC = PC + 1; fetch(); cycle(); // 21
  141. LV = TOS; cycle(); // 22
  142. break;
  143. case 0xac: // IRETURN
  144. MAR = SP = LV; rd(); cycle(); // ireturn1
  145. cycle(); // ireturn2
  146. LV = MAR = MDR; rd(); cycle(); // ireturn3
  147. MAR = LV + 1; cycle(); // ireturn4
  148. PC = MDR; rd(); fetch(); cycle(); // ireturn5
  149. MAR = SP; cycle(); // ireturn6
  150. LV = MDR; cycle(); // ireturn7
  151. MDR = TOS; wr(); cycle(); // ireturn8
  152. break;
  153. }
  154. count++;
  155. if (shown[op] != cycles - start) { // print each new case once
  156. shown[op] = cycles - start;
  157. printf("%s: %d microinstructions\n", name(op), shown[op]);
  158. }
  159. }
  160. printf("%d IJVM instructions, %d microinstructions\n", count, cycles);
  161. return TOS;
  162. }
step 0
Loading emulator…
Your program as you wrote it: the current line, its variables by name, and its output.

Press Continue. It returns 42 and prints each instruction's cost the first time it appears:

BIPUSH: 4 microinstructions
INVOKEVIRTUAL: 23 microinstructions
ISTORE: 7 microinstructions
ILOAD: 6 microinstructions
IFEQ: 8 microinstructions
IADD: 4 microinstructions
ISUB: 4 microinstructions
GOTO: 7 microinstructions
IFEQ: 11 microinstructions
IRETURN: 9 microinstructions
87 IJVM instructions, 533 microinstructions

These are exactly the lengths of the Mic-1 microprogram's routines, Main1 included. On average, one IJVM instruction took just over 6 cycles. Step through the while loop and watch MAR, MDR, PC, MBR, SP, LV, TOS and H change in the Variables panel.

The timing is real, not decoration. Delete the cycle(); after H = TOS; in the IADD case, so iadd2 and iadd3 become one cycle. The program now returns 12 instead of 42: IADD and ISUB read MDR before the stack word has arrived, and add a stale value. A real microprogrammer has to respect those delays in every routine.

The C code above doesn't pretend to be a fast interpreter. The bytecode chapter runs the same kind of JVM bytecode with an ordinary C switch loop, and measures what that costs; here the point is to see the hardware's registers and cycles.

Making it faster: Mic-2, Mic-3, Mic-4

The Mic-1 can be sped up in steps. There are three basic levers: fewer cycles per instruction (a shorter path length), a shorter cycle, or overlapping instructions.

Shorter paths. Two cheap changes come first. Main1 can be merged into the end of each routine, where a cycle is sometimes idle anyway. And a third bus lets the ALU take any register on its left input, so H = LV disappears from ILOAD, which drops from 6 to 5 cycles.

Mic-2: an instruction fetch unit. The big win is to stop using the ALU to fetch the program. The IFU has its own incrementer, fetches whole 4-byte words ahead of time into a shift register, and presents the next byte in MBR1 and the next two bytes, already joined into a 16-bit value, in MBR2. It advances PC itself as bytes are consumed. Main1 disappears entirely: each routine ends with goto (MBR1), dispatching straight to the next instruction. Counting the microinstructions of the Mic-1 and Mic-2 microprograms:

InstructionMic-1Mic-2
IADD43
ILOAD63
GOTO74
IFEQ taken / not taken11 / 88 / 6
INVOKEVIRTUAL2311
mul(6, 7) in the demo533326

The same run would take 326 cycles instead of 533.

Mic-3: pipelining. The Mic-2's cycle includes driving the buses, the ALU, and writing back, one after the other. The Mic-3 puts latches on the A, B and C buses, cutting the datapath into three shorter stages that work on three different microinstructions at once. A microinstruction now takes three shorter cycles, but a new one can start every cycle. When one needs a value that the previous one hasn't produced yet (a RAW dependence), it stalls. SWAP takes 11 short cycles instead of 6 long ones: about 11 versus 18 in the same time units, if each short cycle is a third of a long one. The pipelining chapter covers hazards in detail.

Mic-4: a seven-stage pipeline. The last design stops treating the microprogram as a program. A decoding unit splits the byte stream into instructions using a table of instruction lengths, and looks up where each one's micro-operations start in a ROM. A queueing unit copies those micro-operations into a queue. Four MIR registers, one per stage, carry each micro-operation down the pipeline: operands, execute, write back, memory. The stages are IFU, decoder, queue, operands, execute, write back and memory.

That's the shape of a real x86 front end: decoders turn instructions into µops, which wait in a queue to be executed. The real-cores chapter looks at Intel's Core i7 and today's versions.

The real JVM today

IJVM's opcodes are the JVM's, but no mainstream JVM runs on microcode. A few processors did execute Java bytecode in hardware, such as ARM's Jazelle extension on some 32-bit ARM chips of the 2000s, but just-in-time compilers made it pointless, and 64-bit ARM dropped it.

HotSpot, the JVM in OpenJDK, starts by interpreting bytecode with a template interpreter: a small piece of machine code generated at startup for each opcode, dispatched much like goto (MBR). When a method gets hot, it's compiled to native code, first quickly (the C1 compiler, tier 3) and then with full optimization (C2, tier 4). Running mul 20 million times on OpenJDK 26 with -XX:+PrintCompilation shows it happening within the first 20 ms:

17    8       3       Hot::mul (19 bytes)
17    9       4       Hot::mul (19 bytes)

The 19 bytes are the bytecode above. On the M2 Ultra this chapter was written on, the whole run took about 1.1 s with -Xint (interpreter only) and 0.05 s with the JIT enabled, JVM startup included. The compilers chapter explains how JIT tiers work.

Takeaways

  • The Mic-1 is a complete microprogrammed CPU: ten registers, one ALU fed by H and the B bus, a word port (MAR/MDR) for data and a byte port (PC/MBR) for code, and a 512 × 36-bit control store.
  • IJVM is a stack machine with the JVM's opcodes. Data is reached only through LV (locals), SP (stack) and CPP (constant pool).
  • Main1 fetches the next byte and dispatches on the opcode in one cycle. The routines then take from 4 cycles (IADD) to 23 (INVOKEVIRTUAL), about 6 on average in our run.
  • The microprogram hides memory latency by doing useful work while a read is in flight; get the timing wrong and the result is wrong.
  • Mic-2 adds an instruction fetch unit (533 → 326 cycles for our run), Mic-3 pipelines the datapath, and Mic-4 decodes into queued micro-operations, the organization of real CPUs.
  • The real JVM uses the same bytecode, but runs it with an interpreter and JIT compilers, not microcode.

In this level

  1. 6.1The fetch–decode–execute cycle
  2. 6.2Datapath, internal and system buses
  3. 6.3Control units and microcode
  4. 6.4A complete machine: the Mic-1 running IJVM
  5. 6.5Pipelining and hazards
  6. 6.6Caches and the memory hierarchy
  7. 6.7Branch prediction
  8. 6.8Out-of-order execution, register renaming and speculation
  9. 6.9Real cores: x86, ARM and AVR compared
  10. 6.10SIMD, GPUs and coprocessors
  11. 6.11Multicore, multithreading and cache coherence
  12. 6.12Shared-memory multiprocessors and NUMA
  13. 6.13Clusters, message passing and supercomputers