The microprogramming chapter explained the idea: a control store full of microinstructions drives the datapath one cycle at a time. This chapter follows one complete machine from the registers up. The machine is the Mic-1, a classic teaching microarchitecture. The instruction set it runs is IJVM, a small integer subset of Java bytecode used for teaching. The whole interpreter is 112 microinstructions long.
It's worth studying in detail for one reason: it's small enough to hold in your head, yet it contains every mechanism a real microcoded CPU needs: instruction fetch, operand decoding, a stack, memory latency, branches and procedure calls.
The Mic-1 datapath
The Mic-1 has ten 32-bit registers (MBR is 8 bits), one ALU and a shifter, connected by two buses:
| Register | Holds |
|---|---|
| MAR, MDR | address and data for the 32-bit word memory port |
| PC, MBR | address and data for the 8-bit byte port, used only to fetch the program |
| SP | address of the top word of the stack |
| LV | address of the current method's local variables |
| CPP | address of the constant pool |
| TOS | a copy of the word at the top of the stack |
| OPC | scratch; often the address of the current opcode |
| H | the "holding" register: the ALU's only left input |
On each cycle, one register drives the B bus, which is the ALU's right input. The left input is always H. The result goes through the shifter onto the C bus, and can be written into any number of registers at once. So MAR = SP = SP − 1 is one cycle: SP on the B bus, the ALU computes B − 1, and the result lands in both MAR and SP.
There is no way to add two arbitrary registers in one cycle. One of them must first be copied into H. That limitation shows up in almost every instruction below.
The two memory ports behave differently. MAR counts in words: MAR = 2 reads bytes 8–11. PC counts in bytes, because IJVM instructions are a stream of bytes of different lengths. The byte from the PC port lands in MBR, and MBR can go onto the B bus two ways: sign-extended (written MBR) or zero-extended (MBRU).
Memory is slow, even in this idealized machine. A read started at the end of cycle k delivers its data at the end of cycle k + 1, so the value can only be used in cycle k + 2. The microprogram has to fill that gap with useful work, or waste a cycle.
Each microinstruction is 36 bits: a 9-bit next address, 3 JAM bits for branching, 8 ALU and shifter bits, 9 bits choosing which registers the C bus writes, 3 memory bits (read, write, fetch) and a 4-bit B-bus selector. The control store holds 512 of them. The microprogramming chapter covers how the next address is formed; here we look at what the microinstructions actually do.
IJVM: the machine being interpreted
IJVM is a stack machine. Its instructions don't name registers: operands are pushed on a stack, and operations pop them and push the result. Memory is divided into four areas, each reached through its own pointer:
- the constant pool, read-only, at CPP. It holds constants and the addresses of methods;
- the local variable frame of the current method, at LV. Parameters come first, then local variables;
- the operand stack, directly above the frame, whose top is at SP;
- the method area, the program's bytecode, read byte by byte through PC.
A program never sees an absolute data address. It can only say "local variable 3" or "constant 0", and the microcode adds the offset to LV or CPP.
The instruction set has 20 instructions. Here they are, with the same opcode numbers as the real JVM:
| Hex | Mnemonic | Effect |
|---|---|---|
10 | BIPUSH byte | push a signed byte |
13 | LDC_W index | push a word from the constant pool |
15 | ILOAD varnum | push a local variable |
36 | ISTORE varnum | pop into a local variable |
84 | IINC varnum const | add a signed byte to a local variable |
60 64 | IADD, ISUB | pop two words, push the sum or difference |
7E 80 | IAND, IOR | pop two words, push the AND or OR |
59 57 5F | DUP, POP, SWAP | copy, drop or swap stack words |
A7 | GOTO offset | unconditional branch |
99 9B | IFEQ, IFLT offset | pop a word, branch if zero or negative |
9F | IF_ICMPEQ offset | pop two words, branch if equal |
B6 | INVOKEVIRTUAL disp | call a method |
AC | IRETURN | return an integer |
C4 | WIDE | prefix: the next ILOAD or ISTORE has a 16-bit index |
00 | NOP | nothing |
Branch offsets are signed 16-bit numbers, big-endian, relative to the address of the branch opcode itself.
From Java to IJVM
Here is a Java method that multiplies by repeated addition:
int mul(int a, int b) {
int r = 0;
while (b != 0) {
r = r + a;
b = b - 1;
}
return r;
}
This is what javac from OpenJDK 26 produces for it, shown by javap -c:
0: iconst_0
1: istore_3
2: iload_2
3: ifeq 17
6: iload_3
7: iload_1
8: iadd
9: istore_3
10: iload_2
11: iconst_1
12: isub
13: istore_2
14: goto 2
17: iload_3
18: ireturn
The raw bytes in the .class file include 99 00 0e (ifeq +14), 60 (iadd), 64 (isub), a7 ff f4 (goto −12) and ac (ireturn): the IJVM opcodes exactly. The difference is that the real JVM has one-byte short forms for common cases: iconst_0 instead of BIPUSH 0, istore_3 instead of ISTORE 3. IJVM leaves them out and always uses the general form, so the IJVM version is a few bytes longer.
Local variable 0 is this, the object the method was called on; a, b and r are variables 1, 2 and 3. javap -v reports locals=4, args_size=3.
The main loop: one microinstruction
Every IJVM instruction starts with the same microinstruction, Main1:
Main1 PC = PC + 1; fetch; goto (MBR)
It does three things in one cycle. It advances PC past the opcode. It starts fetching the next byte, which will be either the instruction's first operand or the next opcode. And it jumps to the control-store address equal to the opcode sitting in MBR: the multiway branch the microprogramming chapter described. The routine for IADD starts at control-store address 0x60, the one for ILOAD at 0x15.
That's why each microinstruction names its successor explicitly. Addresses 0x00 to 0xFF are reserved for the first microinstruction of each opcode, so the rest of each routine has to live in the gaps, scattered across the control store.
The invariant that makes this work: when control returns to Main1, the next opcode must already be in MBR. Every routine is written to guarantee it.
IADD: three microinstructions
iadd1 MAR = SP = SP − 1; rd read the word below the top
iadd2 H = TOS top of stack into H, while the read completes
iadd3 MDR = TOS = MDR + H; wr; goto Main1
TOS already holds the top of the stack, so only one memory read is needed. iadd1 points SP and MAR at the second word and starts the read. The data won't be usable until two cycles later, so iadd2 uses the gap to move TOS into H. In iadd3 the word has arrived in MDR. The sum goes to MDR (to be written back to memory) and to TOS (to keep the copy current), and the write starts.
With Main1, IADD costs 4 microinstructions. ISUB, IAND and IOR are the same routine with a different ALU operation.
ILOAD: waiting for memory twice
iload1 H = LV LV into H, to be added to the index
iload2 MAR = MBRU + H; rd address of the local variable; read it
iload3 MAR = SP = SP + 1 new top of stack; wait for the read
iload4 PC = PC + 1; fetch; wr fetch the next opcode; write the value on the stack
iload5 TOS = MDR; goto Main1
The index byte was fetched by Main1, and it's zero-extended (MBRU): a variable number is never negative. LV has to go through H first, because MBR and LV are both on the B bus side. Then it's read, pushed, and copied into TOS. 6 microinstructions, counting Main1.
GOTO: assembling a 16-bit offset
goto1 OPC = PC − 1 save the address of the opcode
goto2 PC = PC + 1; fetch fetch the second offset byte
goto3 H = MBR << 8 high byte, sign-extended, shifted into place
goto4 H = MBRU OR H low byte, zero-extended, OR-ed in
goto5 PC = OPC + H; fetch jump, and fetch the opcode there
goto6 goto Main1 wait for that fetch to arrive
The offset is relative to the opcode, but PC has already moved past it, so goto1 recovers the opcode's address into OPC. The two offset bytes arrive one at a time through the 8-bit port. The high byte is sign-extended, because offsets can be negative, and the low byte is not. goto6 does nothing: it only waits for the new opcode to reach MBR, as Main1 requires. 7 microinstructions in total.
The conditional branches reuse this code. IFEQ pops the top of stack, reads the new top into TOS, and passes the popped value through the ALU to set the Z flag. That microinstruction's JAMZ bit makes the next address depend on Z. If the branch is taken, it jumps into the middle of the GOTO routine; if not, three short microinstructions skip over the offset. IFEQ costs 11 microinstructions when taken and 8 when not.
INVOKEVIRTUAL: building a frame
The call instruction is by far the longest: 22 microinstructions after Main1. Its operand is an index into the constant pool, where the method's address is stored. The method itself starts with two 16-bit numbers, its number of parameters (counting the object reference) and its number of local variables, followed by its first opcode.
The microcode fetches all of that one byte at a time, then builds the new frame. Here is the stack in the demo below, right after INVOKEVIRTUAL calls mul(6, 7):
word 13 caller's LV = 8 ← SP
word 12 caller's PC = 9
word 11 r (local 3)
word 10 b = 7 (local 2)
word 9 a = 6 (local 1)
word 8 link pointer = 12 ← LV (was the object reference)
The caller had pushed the object reference and the two arguments. The microcode points LV at the object reference and overwrites it with a link pointer: the address, above the local variables, where it saves the caller's PC and LV. The new method's operand stack starts just above them.
IRETURN undoes this in 8 microinstructions. It follows the link pointer to restore PC and LV, and writes the return value where the object reference used to be, so the caller finds the result on top of its stack.
Running the Mic-1
The demo below is the Mic-1 written in C. The global variables are the Mic-1's registers, and every line that ends in cycle() is one microinstruction of the Mic-1 microprogram, with its label in the comment. cycle() models the memory timing: a read started in one cycle only reaches MDR or MBR at the end of the next one. The IJVM program is the mul method above, hand-assembled, called as mul(6, 7). (IJVM has no instruction to stop the machine, so the demo adds a HALT opcode, 0xFF, to its main program.)
Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.
- // A Mic-1 running IJVM: every line ending in cycle() is one microinstruction
- // of the Mic-1 microprogram, acting on the Mic-1's registers.
- int MAR, MDR, PC, MBR, SP, LV, CPP, TOS, OPC, H;
- int N, Z; // ALU flags latched by the last cycle
- int mem[32]; // word memory: constant pool and stack
- unsigned char code[] = { // method area (byte addressed)
- 0x10, 0, // 0 BIPUSH 0 OBJREF (unused)
- 0x10, 6, // 2 BIPUSH 6 a
- 0x10, 7, // 4 BIPUSH 7 b
- 0xb6, 0, 0, // 6 INVOKEVIRTUAL 0 mul(a, b)
- 0xff, // 9 HALT (not in IJVM: stops the demo)
- 0, 3, 0, 1, // 10 mul: 3 parameters (OBJREF, a, b), 1 local (r)
- 0x10, 0, // 14 BIPUSH 0
- 0x36, 3, // 16 ISTORE 3 r = 0
- 0x15, 2, // 18 ILOAD 2 loop:
- 0x99, 0, 20, // 20 IFEQ +20 if b == 0 goto done
- 0x15, 3, // 23 ILOAD 3
- 0x15, 1, // 25 ILOAD 1
- 0x60, // 27 IADD
- 0x36, 3, // 28 ISTORE 3 r = r + a
- 0x15, 2, // 30 ILOAD 2
- 0x10, 1, // 32 BIPUSH 1
- 0x64, // 34 ISUB
- 0x36, 2, // 35 ISTORE 2 b = b - 1
- 0xa7, 0xff, 0xed, // 37 GOTO -19 goto loop
- 0x15, 3, // 40 ILOAD 3 done:
- 0xac // 42 IRETURN return r
- };
- int rd_now, rd_wait, rd_addr, fetch_now, fetch_wait, fetch_addr;
- int cycles, count, shown[256];
- void rd() { rd_now = 1; }
- void wr() { mem[MAR] = MDR; }
- void fetch() { fetch_now = 1; }
- int MBRU() { return MBR; } // zero-extended
- int MBRS() { return (signed char)MBR; } // sign-extended
- // End of a clock cycle: a read started in the previous cycle arrives now,
- // so its data can be used by the microinstruction after next.
- void cycle() {
- if (rd_wait) { MDR = mem[rd_addr]; rd_wait = 0; }
- if (fetch_wait) { MBR = code[fetch_addr]; fetch_wait = 0; }
- if (rd_now) { rd_addr = MAR; rd_wait = 1; rd_now = 0; }
- if (fetch_now) { fetch_addr = PC; fetch_wait = 1; fetch_now = 0; }
- cycles++;
- }
- char *name(int op) {
- switch (op) {
- case 0x10: return "BIPUSH"; case 0x15: return "ILOAD";
- case 0x36: return "ISTORE"; case 0x60: return "IADD";
- case 0x64: return "ISUB"; case 0x99: return "IFEQ";
- case 0xa7: return "GOTO"; case 0xac: return "IRETURN";
- case 0xb6: return "INVOKEVIRTUAL";
- }
- return "?";
- }
- void alu(int v) { N = v < 0; Z = v == 0; }
- int main() {
- CPP = 0; mem[0] = 10; // constant pool: address of mul
- LV = 8; SP = 7; // empty stack for the main program
- PC = 0; MBR = code[0]; // invariant: opcode already in MBR
- while (1) {
- int op = MBR, start = cycles;
- if (op == 0xff) break;
- PC = PC + 1; fetch(); cycle(); // Main1: goto (MBR)
- switch (op) {
- case 0x10: // BIPUSH
- SP = MAR = SP + 1; cycle(); // bipush1
- PC = PC + 1; fetch(); cycle(); // bipush2
- MDR = TOS = MBRS(); wr(); cycle(); // bipush3
- break;
- case 0x15: // ILOAD
- H = LV; cycle(); // iload1
- MAR = MBRU() + H; rd(); cycle(); // iload2
- MAR = SP = SP + 1; cycle(); // iload3
- PC = PC + 1; fetch(); wr(); cycle(); // iload4
- TOS = MDR; cycle(); // iload5
- break;
- case 0x36: // ISTORE
- H = LV; cycle(); // istore1
- MAR = MBRU() + H; cycle(); // istore2
- MDR = TOS; wr(); cycle(); // istore3
- SP = MAR = SP - 1; rd(); cycle(); // istore4
- PC = PC + 1; fetch(); cycle(); // istore5
- TOS = MDR; cycle(); // istore6
- break;
- case 0x60: // IADD
- case 0x64: // ISUB
- MAR = SP = SP - 1; rd(); cycle(); // iadd1
- H = TOS; cycle(); // iadd2
- if (op == 0x60) MDR = TOS = MDR + H; // iadd3
- else MDR = TOS = MDR - H; // (isub3)
- wr(); cycle();
- break;
- case 0x99: // IFEQ
- MAR = SP = SP - 1; rd(); cycle(); // ifeq1
- OPC = TOS; cycle(); // ifeq2
- TOS = MDR; cycle(); // ifeq3
- alu(OPC); cycle(); // ifeq4: Z = OPC
- if (!Z) {
- PC = PC + 1; cycle(); // F
- PC = PC + 1; fetch(); cycle(); // F2
- cycle(); // F3
- break;
- }
- OPC = PC - 1; cycle(); // T, then goto2
- PC = PC + 1; fetch(); cycle(); // goto2
- H = MBRS() << 8; cycle(); // goto3
- H = MBRU() | H; cycle(); // goto4
- PC = OPC + H; fetch(); cycle(); // goto5
- cycle(); // goto6
- break;
- case 0xa7: // GOTO
- OPC = PC - 1; cycle(); // goto1
- PC = PC + 1; fetch(); cycle(); // goto2
- H = MBRS() << 8; cycle(); // goto3
- H = MBRU() | H; cycle(); // goto4
- PC = OPC + H; fetch(); cycle(); // goto5
- cycle(); // goto6
- break;
- case 0xb6: // INVOKEVIRTUAL
- PC = PC + 1; fetch(); cycle(); // 1
- H = MBRU() << 8; cycle(); // 2
- H = MBRU() | H; cycle(); // 3
- MAR = CPP + H; rd(); cycle(); // 4
- OPC = PC + 1; cycle(); // 5
- PC = MDR; fetch(); cycle(); // 6
- PC = PC + 1; fetch(); cycle(); // 7
- H = MBRU() << 8; cycle(); // 8
- H = MBRU() | H; cycle(); // 9
- PC = PC + 1; fetch(); cycle(); // 10
- TOS = SP - H; cycle(); // 11
- TOS = MAR = TOS + 1; cycle(); // 12
- PC = PC + 1; fetch(); cycle(); // 13
- H = MBRU() << 8; cycle(); // 14
- H = MBRU() | H; cycle(); // 15
- MDR = SP + H + 1; wr(); cycle(); // 16
- MAR = SP = MDR; cycle(); // 17
- MDR = OPC; wr(); cycle(); // 18
- MAR = SP = SP + 1; cycle(); // 19
- MDR = LV; wr(); cycle(); // 20
- PC = PC + 1; fetch(); cycle(); // 21
- LV = TOS; cycle(); // 22
- break;
- case 0xac: // IRETURN
- MAR = SP = LV; rd(); cycle(); // ireturn1
- cycle(); // ireturn2
- LV = MAR = MDR; rd(); cycle(); // ireturn3
- MAR = LV + 1; cycle(); // ireturn4
- PC = MDR; rd(); fetch(); cycle(); // ireturn5
- MAR = SP; cycle(); // ireturn6
- LV = MDR; cycle(); // ireturn7
- MDR = TOS; wr(); cycle(); // ireturn8
- break;
- }
- count++;
- if (shown[op] != cycles - start) { // print each new case once
- shown[op] = cycles - start;
- printf("%s: %d microinstructions\n", name(op), shown[op]);
- }
- }
- printf("%d IJVM instructions, %d microinstructions\n", count, cycles);
- return TOS;
- }
Press Continue. It returns 42 and prints each instruction's cost the first time it appears:
BIPUSH: 4 microinstructions
INVOKEVIRTUAL: 23 microinstructions
ISTORE: 7 microinstructions
ILOAD: 6 microinstructions
IFEQ: 8 microinstructions
IADD: 4 microinstructions
ISUB: 4 microinstructions
GOTO: 7 microinstructions
IFEQ: 11 microinstructions
IRETURN: 9 microinstructions
87 IJVM instructions, 533 microinstructions
These are exactly the lengths of the Mic-1 microprogram's routines, Main1 included. On average, one IJVM instruction took just over 6 cycles. Step through the while loop and watch MAR, MDR, PC, MBR, SP, LV, TOS and H change in the Variables panel.
The timing is real, not decoration. Delete the cycle(); after H = TOS; in the IADD case, so iadd2 and iadd3 become one cycle. The program now returns 12 instead of 42: IADD and ISUB read MDR before the stack word has arrived, and add a stale value. A real microprogrammer has to respect those delays in every routine.
The C code above doesn't pretend to be a fast interpreter. The bytecode chapter runs the same kind of JVM bytecode with an ordinary C switch loop, and measures what that costs; here the point is to see the hardware's registers and cycles.
Making it faster: Mic-2, Mic-3, Mic-4
The Mic-1 can be sped up in steps. There are three basic levers: fewer cycles per instruction (a shorter path length), a shorter cycle, or overlapping instructions.
Shorter paths. Two cheap changes come first. Main1 can be merged into the end of each routine, where a cycle is sometimes idle anyway. And a third bus lets the ALU take any register on its left input, so H = LV disappears from ILOAD, which drops from 6 to 5 cycles.
Mic-2: an instruction fetch unit. The big win is to stop using the ALU to fetch the program. The IFU has its own incrementer, fetches whole 4-byte words ahead of time into a shift register, and presents the next byte in MBR1 and the next two bytes, already joined into a 16-bit value, in MBR2. It advances PC itself as bytes are consumed. Main1 disappears entirely: each routine ends with goto (MBR1), dispatching straight to the next instruction. Counting the microinstructions of the Mic-1 and Mic-2 microprograms:
| Instruction | Mic-1 | Mic-2 |
|---|---|---|
| IADD | 4 | 3 |
| ILOAD | 6 | 3 |
| GOTO | 7 | 4 |
| IFEQ taken / not taken | 11 / 8 | 8 / 6 |
| INVOKEVIRTUAL | 23 | 11 |
mul(6, 7) in the demo | 533 | 326 |
The same run would take 326 cycles instead of 533.
Mic-3: pipelining. The Mic-2's cycle includes driving the buses, the ALU, and writing back, one after the other. The Mic-3 puts latches on the A, B and C buses, cutting the datapath into three shorter stages that work on three different microinstructions at once. A microinstruction now takes three shorter cycles, but a new one can start every cycle. When one needs a value that the previous one hasn't produced yet (a RAW dependence), it stalls. SWAP takes 11 short cycles instead of 6 long ones: about 11 versus 18 in the same time units, if each short cycle is a third of a long one. The pipelining chapter covers hazards in detail.
Mic-4: a seven-stage pipeline. The last design stops treating the microprogram as a program. A decoding unit splits the byte stream into instructions using a table of instruction lengths, and looks up where each one's micro-operations start in a ROM. A queueing unit copies those micro-operations into a queue. Four MIR registers, one per stage, carry each micro-operation down the pipeline: operands, execute, write back, memory. The stages are IFU, decoder, queue, operands, execute, write back and memory.
That's the shape of a real x86 front end: decoders turn instructions into µops, which wait in a queue to be executed. The real-cores chapter looks at Intel's Core i7 and today's versions.
The real JVM today
IJVM's opcodes are the JVM's, but no mainstream JVM runs on microcode. A few processors did execute Java bytecode in hardware, such as ARM's Jazelle extension on some 32-bit ARM chips of the 2000s, but just-in-time compilers made it pointless, and 64-bit ARM dropped it.
HotSpot, the JVM in OpenJDK, starts by interpreting bytecode with a template interpreter: a small piece of machine code generated at startup for each opcode, dispatched much like goto (MBR). When a method gets hot, it's compiled to native code, first quickly (the C1 compiler, tier 3) and then with full optimization (C2, tier 4). Running mul 20 million times on OpenJDK 26 with -XX:+PrintCompilation shows it happening within the first 20 ms:
17 8 3 Hot::mul (19 bytes)
17 9 4 Hot::mul (19 bytes)
The 19 bytes are the bytecode above. On the M2 Ultra this chapter was written on, the whole run took about 1.1 s with -Xint (interpreter only) and 0.05 s with the JIT enabled, JVM startup included. The compilers chapter explains how JIT tiers work.
Takeaways
- The Mic-1 is a complete microprogrammed CPU: ten registers, one ALU fed by H and the B bus, a word port (MAR/MDR) for data and a byte port (PC/MBR) for code, and a 512 × 36-bit control store.
- IJVM is a stack machine with the JVM's opcodes. Data is reached only through LV (locals), SP (stack) and CPP (constant pool).
Main1fetches the next byte and dispatches on the opcode in one cycle. The routines then take from 4 cycles (IADD) to 23 (INVOKEVIRTUAL), about 6 on average in our run.- The microprogram hides memory latency by doing useful work while a read is in flight; get the timing wrong and the result is wrong.
- Mic-2 adds an instruction fetch unit (533 → 326 cycles for our run), Mic-3 pipelines the datapath, and Mic-4 decodes into queued micro-operations, the organization of real CPUs.
- The real JVM uses the same bytecode, but runs it with an interpreter and JIT compilers, not microcode.