Skip to content

Level 1 · Chapter 1.4

Assembler, linker and loader

How assembly text becomes a running process: the assembler's two passes and symbol table, object files and relocations, what the linker patches, and what the loader does before the first instruction runs.

The CPU never sees mov eax, 1 or call add2. It sees bytes. Three programs turn source text into bytes running in memory:

  1. The assembler translates each source file into an object file: machine code, plus notes about everything it could not finish on its own.
  2. The linker merges the object files into one executable, choosing an address for every piece and finishing what the assembler left open.
  3. The loader — part of the operating system — maps the executable into memory, prepares the stack, and jumps to the first instruction.

A compiler usually sits in front of this chain: gcc turns C into assembly, then runs the assembler and linker for you. Everything below happens on every build, whether you wrote the assembly or not.

Labels become addresses

In assembly you write names — done, add2, message. The CPU needs numbers. Turning names into numbers is the assembler's main job, and the simulator shows the result:

Live · A label is an address
program— ▸ is the next instruction
  1. _start:
  2. mov eax, 1
  3. jmp done ; used before it is defined
  4. mov eax, 2 ; skipped
  5. done:
  6. mov ebx, eax
step 0
Loading emulator…

After the jmp, rip is 0x401003: that's the address of done. The simulator gives each instruction its own address, so done is simply _start + 3. A real assembler counts bytes instead, because x86 instructions have different lengths — but the principle is the same.

Why assemblers make two passes

Look at jmp done again. When the assembler reaches that line, it has not seen done yet, so it does not know where to jump. This is the forward reference problem, and the classic answer is to read the source twice:

  • Pass one walks through the program with a location counter that starts at 0 and grows by each instruction's length. Every time a label appears, the assembler records its name and the current counter value in the symbol table. No code is produced yet — only lengths and addresses.
  • Pass two walks through again. Now every label has a value, so each instruction can be encoded completely and written out.

For the program above, pass one produces this symbol table (in bytes, as a real assembler would count them):

SymbolValueWhy
_start0x0the first instruction
done0xcafter mov eax, 1 (5 bytes), jmp done (2 bytes) and mov eax, 2 (5 bytes): 5 + 2 + 5 = 12

In pass two, jmp done can then be encoded as eb 05, "jump 5 bytes forward" — from the end of the jmp at 0x7 to done at 0xc.

Pass one also has to know each instruction's length before it knows the addresses. On x86 that's subtle: a jump to a nearby label fits in 2 bytes (eb + an 8-bit offset), a distant one needs 5 (e9 + a 32-bit offset). Real assemblers start with the short form and re-run the layout, growing any jump that doesn't fit, until every address settles.

What an object file contains

Here is a two-file program, assembled for Linux x86-64. main.s calls a function that lives in another file:

; main.s
_start:
    mov edi, 3
    mov esi, 4
    call add2       ; defined in another file
    mov edi, eax
    jmp done
    nop
done:
    mov eax, 60     ; exit
    syscall
; add2.s
add2:
    lea eax, [rdi+rsi]
    ret

Disassembling the object file main.o gives:

0000000000000000 <_start>:
   0:  bf 03 00 00 00     mov    edi, 3
   5:  be 04 00 00 00     mov    esi, 4
   a:  e8 00 00 00 00     call   f <_start+0xf>
            b: R_X86_64_PLT32  add2-0x4
   f:  89 c7              mov    edi, eax
  11:  eb 01              jmp    14 <done>
  13:  90                 nop
0000000000000014 <done>:
  14:  b8 3c 00 00 00     mov    eax, 60
  19:  0f 05              syscall

Three things to notice:

  • Addresses start at 0. The assembler has no idea where the code will end up, so every object file pretends it starts at address 0.
  • jmp done is finished. eb 01 means "jump 1 byte forward from the next instruction". The distance between two points in the same section never changes, wherever the section ends up, so the assembler can encode it on its own.
  • call add2 is not. add2 is in another file, so the assembler writes e8 followed by a placeholder 00 00 00 00, and leaves a relocation: a note saying "at offset 0xb, put the address of add2, minus 4, relative to this spot". That's why objdump shows the call going to f, the very next instruction — it's reading the placeholder.

That second bullet is worth remembering in reverse engineering: in an unlinked .o file, every call to another file looks like a call to the next instruction.

An object file is organized into sections. The ones that matter here:

SectionHolds
.textmachine code
.data, .rodata, .bssinitialized data, read-only data, zero-initialized data
.symtabthe symbol table: symbols this file defines (like _start) and symbols it uses but doesn't define (like add2)
.rela.textthe relocations: every spot in .text the linker must patch

nm prints the symbol table: T _start (defined, global), t done (defined, local to this file), U add2 (undefined — someone else must provide it).

What the linker does

The linker takes all the object files and solves the two problems the assembler couldn't:

  1. Relocation. Every object file starts at 0, but they can't all live at 0. The linker lays them out one after another and assigns each one a real address. Suppose main.o's code (0x1b bytes) lands at 0x401000 — the usual start of code in a Linux executable — and add2.o's right after it, rounded up to a 4-byte boundary, at 0x40101c.
  2. External references. Now it can resolve U add2 against add2.o's T add2 and apply the relocation: the value is add2 − 4 − (address of the patch) = 0x40101c − 4 − 0x40100b = 0xd. The call becomes e8 0d 00 00 00: jump 0xd bytes past the end of the call, which is exactly add2.

If a symbol is still undefined after all the files have been read, you get the familiar undefined reference to 'add2'. If two files define the same global symbol, you get multiple definition. Both are linker errors, not compiler errors.

The same process works for data: an instruction that loads a global variable gets a relocation, and the linker fills in the variable's final address or offset.

When is an address decided?

Every name in a program is turned into a number at some point. When that happens is called binding time, and different names are bound at different times:

Bound at…Example
assemblyjmp done inside one file
link timecall add2 to another object file
load timeaddresses in a program that the OS places at a random base (ASLR)
first usea shared-library function resolved on its first call (lazy binding)

The later the binding, the more flexible the program — and the more work left for later.

Code that only uses relative addresses — "jump 0xd bytes ahead", "the data is 0x2f00 bytes from rip" — keeps working wherever it's placed. That's position-independent code. x86-64 makes it cheap with rip-relative addressing (lea rax, [rip+message]), and modern Linux builds executables as PIE (position-independent executables) by default, so the OS can load them at a random address.

Static and dynamic linking

Everything above is static linking: all the code the program needs is copied into the executable. The alternative is to leave library code in shared libraries (.so on Linux, .dll on Windows) and link them at run time:

  • The executable lists the libraries it needs (libc.so.6) and the symbols it imports (printf).
  • Calls to imported functions go through a small table of pointers — the GOT (global offset table) on Linux, the IAT (import address table) on Windows — that the dynamic linker fills in.
  • On Linux, each imported function gets a tiny stub in the PLT (procedure linkage table) that jumps through its GOT entry. With lazy binding, that entry first points at the resolver, which looks up the real address on the first call and overwrites the entry. Every later call goes straight to the function. Many distributions now turn lazy binding off (-z now, "full RELRO"): everything is resolved at load time and the GOT is then made read-only.

Shared libraries save memory and disk — one copy of libc serves every process — and let a library be fixed without relinking every program that uses it.

The loader

Running ./prog on Linux ends in the execve system call, and the kernel then:

  1. reads the ELF header and its program headers, which list the segments to load;
  2. maps those segments into a fresh address space — code read-only and executable, data read-write;
  3. builds the initial stack: argc, argv, the environment, and some extra information for the runtime;
  4. if the program is dynamically linked, loads the dynamic linker (ld-linux-x86-64.so.2) too, and starts that first. It maps the shared libraries, applies their relocations, and fills the GOT;
  5. jumps to the entry point — _start, whose address is stored in the ELF header.

_start is not main. It's a small piece of startup code from the C library. It hands control to the library's startup routine (__libc_start_main in glibc), which initializes the runtime, calls main, and passes its return value to exit.

Windows follows the same outline with a different file format: a PE file, whose imports are resolved into the IAT by the Windows loader before the entry point runs.

Takeaways

  • The assembler turns labels into addresses with a symbol table, usually in two passes, and writes an object file that starts at address 0.
  • What it can't finish — references to other files and absolute addresses — it leaves as relocations.
  • The linker lays out the object files, resolves undefined symbols and applies the relocations.
  • The loader maps the executable, builds the stack, loads shared libraries, and jumps to the entry point.
  • Relative addressing makes code position-independent; shared libraries are reached through the GOT/PLT (Linux) or the IAT (Windows).

In this level

  1. 1.1Registers and the register file
  2. 1.2Flags, cmp and conditional jumps
  3. 1.3Calling conventions (System V, cdecl)
  4. 1.4Assembler, linker and loader
  5. 1.5Directives, sections and macros