The CPU never sees mov eax, 1 or call add2. It sees bytes. Three programs turn source text into bytes running in memory:
- The assembler translates each source file into an object file: machine code, plus notes about everything it could not finish on its own.
- The linker merges the object files into one executable, choosing an address for every piece and finishing what the assembler left open.
- The loader — part of the operating system — maps the executable into memory, prepares the stack, and jumps to the first instruction.
A compiler usually sits in front of this chain: gcc turns C into assembly, then runs the assembler and linker for you. Everything below happens on every build, whether you wrote the assembly or not.
Labels become addresses
In assembly you write names — done, add2, message. The CPU needs numbers. Turning names into numbers is the assembler's main job, and the simulator shows the result:
- _start:
- mov eax, 1
- jmp done ; used before it is defined
- mov eax, 2 ; skipped
- done:
- mov ebx, eax
After the jmp, rip is 0x401003: that's the address of done. The simulator gives each instruction its own address, so done is simply _start + 3. A real assembler counts bytes instead, because x86 instructions have different lengths — but the principle is the same.
Why assemblers make two passes
Look at jmp done again. When the assembler reaches that line, it has not seen done yet, so it does not know where to jump. This is the forward reference problem, and the classic answer is to read the source twice:
- Pass one walks through the program with a location counter that starts at 0 and grows by each instruction's length. Every time a label appears, the assembler records its name and the current counter value in the symbol table. No code is produced yet — only lengths and addresses.
- Pass two walks through again. Now every label has a value, so each instruction can be encoded completely and written out.
For the program above, pass one produces this symbol table (in bytes, as a real assembler would count them):
| Symbol | Value | Why |
|---|---|---|
_start | 0x0 | the first instruction |
done | 0xc | after mov eax, 1 (5 bytes), jmp done (2 bytes) and mov eax, 2 (5 bytes): 5 + 2 + 5 = 12 |
In pass two, jmp done can then be encoded as eb 05, "jump 5 bytes forward" — from the end of the jmp at 0x7 to done at 0xc.
Pass one also has to know each instruction's length before it knows the addresses. On x86 that's subtle: a jump to a nearby label fits in 2 bytes (eb + an 8-bit offset), a distant one needs 5 (e9 + a 32-bit offset). Real assemblers start with the short form and re-run the layout, growing any jump that doesn't fit, until every address settles.
What an object file contains
Here is a two-file program, assembled for Linux x86-64. main.s calls a function that lives in another file:
; main.s
_start:
mov edi, 3
mov esi, 4
call add2 ; defined in another file
mov edi, eax
jmp done
nop
done:
mov eax, 60 ; exit
syscall
; add2.s
add2:
lea eax, [rdi+rsi]
ret
Disassembling the object file main.o gives:
0000000000000000 <_start>:
0: bf 03 00 00 00 mov edi, 3
5: be 04 00 00 00 mov esi, 4
a: e8 00 00 00 00 call f <_start+0xf>
b: R_X86_64_PLT32 add2-0x4
f: 89 c7 mov edi, eax
11: eb 01 jmp 14 <done>
13: 90 nop
0000000000000014 <done>:
14: b8 3c 00 00 00 mov eax, 60
19: 0f 05 syscall
Three things to notice:
- Addresses start at 0. The assembler has no idea where the code will end up, so every object file pretends it starts at address 0.
jmp doneis finished.eb 01means "jump 1 byte forward from the next instruction". The distance between two points in the same section never changes, wherever the section ends up, so the assembler can encode it on its own.call add2is not.add2is in another file, so the assembler writese8followed by a placeholder00 00 00 00, and leaves a relocation: a note saying "at offset0xb, put the address ofadd2, minus 4, relative to this spot". That's whyobjdumpshows the call going tof, the very next instruction — it's reading the placeholder.
That second bullet is worth remembering in reverse engineering: in an unlinked .o file, every call to another file looks like a call to the next instruction.
An object file is organized into sections. The ones that matter here:
| Section | Holds |
|---|---|
.text | machine code |
.data, .rodata, .bss | initialized data, read-only data, zero-initialized data |
.symtab | the symbol table: symbols this file defines (like _start) and symbols it uses but doesn't define (like add2) |
.rela.text | the relocations: every spot in .text the linker must patch |
nm prints the symbol table: T _start (defined, global), t done (defined, local to this file), U add2 (undefined — someone else must provide it).
What the linker does
The linker takes all the object files and solves the two problems the assembler couldn't:
- Relocation. Every object file starts at 0, but they can't all live at 0. The linker lays them out one after another and assigns each one a real address. Suppose
main.o's code (0x1bbytes) lands at0x401000— the usual start of code in a Linux executable — andadd2.o's right after it, rounded up to a 4-byte boundary, at0x40101c. - External references. Now it can resolve
U add2againstadd2.o'sT add2and apply the relocation: the value isadd2 − 4 − (address of the patch)=0x40101c − 4 − 0x40100b=0xd. The call becomese8 0d 00 00 00: jump0xdbytes past the end of thecall, which is exactlyadd2.
If a symbol is still undefined after all the files have been read, you get the familiar undefined reference to 'add2'. If two files define the same global symbol, you get multiple definition. Both are linker errors, not compiler errors.
The same process works for data: an instruction that loads a global variable gets a relocation, and the linker fills in the variable's final address or offset.
When is an address decided?
Every name in a program is turned into a number at some point. When that happens is called binding time, and different names are bound at different times:
| Bound at… | Example |
|---|---|
| assembly | jmp done inside one file |
| link time | call add2 to another object file |
| load time | addresses in a program that the OS places at a random base (ASLR) |
| first use | a shared-library function resolved on its first call (lazy binding) |
The later the binding, the more flexible the program — and the more work left for later.
Code that only uses relative addresses — "jump 0xd bytes ahead", "the data is 0x2f00 bytes from rip" — keeps working wherever it's placed. That's position-independent code. x86-64 makes it cheap with rip-relative addressing (lea rax, [rip+message]), and modern Linux builds executables as PIE (position-independent executables) by default, so the OS can load them at a random address.
Static and dynamic linking
Everything above is static linking: all the code the program needs is copied into the executable. The alternative is to leave library code in shared libraries (.so on Linux, .dll on Windows) and link them at run time:
- The executable lists the libraries it needs (
libc.so.6) and the symbols it imports (printf). - Calls to imported functions go through a small table of pointers — the GOT (global offset table) on Linux, the IAT (import address table) on Windows — that the dynamic linker fills in.
- On Linux, each imported function gets a tiny stub in the PLT (procedure linkage table) that jumps through its GOT entry. With lazy binding, that entry first points at the resolver, which looks up the real address on the first call and overwrites the entry. Every later call goes straight to the function. Many distributions now turn lazy binding off (
-z now, "full RELRO"): everything is resolved at load time and the GOT is then made read-only.
Shared libraries save memory and disk — one copy of libc serves every process — and let a library be fixed without relinking every program that uses it.
The loader
Running ./prog on Linux ends in the execve system call, and the kernel then:
- reads the ELF header and its program headers, which list the segments to load;
- maps those segments into a fresh address space — code read-only and executable, data read-write;
- builds the initial stack:
argc,argv, the environment, and some extra information for the runtime; - if the program is dynamically linked, loads the dynamic linker (
ld-linux-x86-64.so.2) too, and starts that first. It maps the shared libraries, applies their relocations, and fills the GOT; - jumps to the entry point —
_start, whose address is stored in the ELF header.
_start is not main. It's a small piece of startup code from the C library. It hands control to the library's startup routine (__libc_start_main in glibc), which initializes the runtime, calls main, and passes its return value to exit.
Windows follows the same outline with a different file format: a PE file, whose imports are resolved into the IAT by the Windows loader before the entry point runs.
Takeaways
- The assembler turns labels into addresses with a symbol table, usually in two passes, and writes an object file that starts at address 0.
- What it can't finish — references to other files and absolute addresses — it leaves as relocations.
- The linker lays out the object files, resolves undefined symbols and applies the relocations.
- The loader maps the executable, builds the stack, loads shared libraries, and jumps to the entry point.
- Relative addressing makes code position-independent; shared libraries are reached through the GOT/PLT (Linux) or the IAT (Windows).