Not every line of an assembly file becomes an instruction. Many lines are directives — also called pseudoinstructions — commands to the assembler itself: put this in the data section, reserve 4 bytes here, make this name visible to other files. And some lines are macros: shorthand that the assembler expands into other lines before assembling anything.
Neither produces a CPU instruction by itself. But together they decide where every byte of the program goes and what it's allowed to do.
Two syntaxes
The directives belong to the assembler, not to the CPU, so they differ from one assembler to another. This chapter uses the two you'll meet most often on Linux:
- GAS, the GNU assembler, which is what
gccandclangproduce andobjdumpreads. Its directives start with a dot:.data,.long,.globl. - NASM, popular for hand-written assembly. It uses bare keywords:
section .data,dd,global.
Microsoft's MASM has a third set again (DB, DD, PUBLIC, SEGMENT). The machine code is the same whichever you use.
Sections
A program is split into sections, each with its own purpose and its own permissions:
| Section | Holds | At run time |
|---|---|---|
.text | machine code | read + execute |
.data | global variables with an initial value | read + write |
.rodata | constants and string literals | read only |
.bss | global variables that start at zero | read + write |
You switch sections with a directive — .data, .text, or .section .rodata in GAS, section .data in NASM — and you can switch back and forth as often as you like. The assembler keeps one location counter per section and appends each line to the current one.
These permissions are what makes some bugs crash: writing to a string literal faults because .rodata is read-only, and jumping into .data faults because it isn't executable. The linker merges the sections of the same name from every object file, and the loader maps each group into memory with its permissions, as described in the previous chapter.
Data directives
Data directives place bytes in the current section. A label in front of one names the address of its first byte:
- .data
- count: .long 5 ; 4 bytes
- table: .byte 10, 20, 30 ; 3 bytes
- .align 8 ; pad to a multiple of 8
- big: .quad 0x1122334455667788
- .section .rodata
- msg: .asciz "hi" ; 'h', 'i', 0
- .bss
- buf: .zero 16
- .text
- _start:
- mov eax, DWORD PTR [rip+count]
- movzx ecx, BYTE PTR [rip+table+2]
- add eax, ecx
- mov DWORD PTR [rip+count], eax
- lea rdi, [rip+msg]
- lea rsi, [rip+buf]
Look at the data panel:
countis at0x404000and holds 5.tableis at0x404004. The panel reads its 4 bytes as0x001e140a: that's 10, 20, 30 (0a,14,1e) read little-endian, plus the padding byte00that.align 8added.bigstarts at0x404008, the next multiple of 8.msgsits in.rodata, andbufis 16 zero bytes in.bss.
The demo stopped after the second instruction: eax = 5 and ecx = 30, the third byte of table. Keep stepping and count becomes 35. The two lea instructions then load the addresses of msg and buf, not their contents.
The simulator packs all its data into one region starting at 0x404000. A real linker keeps .data, .rodata and .bss apart, in separate pages with different permissions.
The same directives in both syntaxes:
| Purpose | GAS | NASM |
|---|---|---|
| 1-byte values | .byte 1, 2 | db 1, 2 |
| 2-byte values | .word / .short | dw |
| 4-byte values | .long / .int | dd |
| 8-byte values | .quad | dq |
| string with a final 0 | .asciz "hi" / .string "hi" | db "hi", 0 |
| string without a final 0 | .ascii "hi" | db "hi" |
| n zero bytes | .zero n / .skip n | times n db 0, or resb n in .bss |
| alignment | .balign 8 / .p2align 3 | align 8 |
Two traps. On x86, GAS's .word is 2 bytes, not the CPU's 32- or 64-bit word. And .align means "n bytes" on x86 but "2ⁿ bytes" on some other targets, which is why compilers write the unambiguous .p2align (a power of two) or .balign (bytes).
What reaches the object file
Assembling the same data with a real assembler and dumping the sections gives:
Contents of section .data:
0000 05000000 0a141e00 88776655 44332211
Contents of section .rodata:
0000 686900 hi.
Every value is stored little-endian: 05000000 is 5, and 8877665544332211 is 0x1122334455667788 backwards.
.bss doesn't appear in the dump because it has no contents. The section header only records its size. Declare a 4096-byte buffer in .bss and the object file stays under a kilobyte; the loader supplies the zeroed memory at run time. That's the whole point of .bss.
In NASM, plain align pads with 0x90 bytes — nop instructions — even in a data section: the same data comes out as 0a141e90. Write align 8, db 0 to pad with zeros. If you see a stray 90 between data values in a hex dump, that's probably why.
Symbols: who can see a name
By default, a label is local to its file: two files can both have a loop label without conflict. Some directives change that:
| Purpose | GAS | NASM |
|---|---|---|
| export a name to other files | .globl main | global main |
| use a name defined elsewhere | (automatic) | extern puts |
| name a constant | .equ SIZE, 64 / .set | SIZE equ 64 |
GAS treats any name that isn't defined in the file as external, so .extern is optional there. NASM requires extern.
A constant from .equ or equ takes up no memory at all. It's a name the assembler replaces with a number while assembling, so it leaves no trace in the binary, only the number.
Reading compiler output
Most of the directives you'll see come from compilers. Here is what clang -S -O1 -masm=intel -fno-pic produces for two lines of C, int counter = 5; and const char *greet(void) { return "hi"; }, lightly trimmed. -fno-pic keeps the code position-dependent. Without it, you'd get lea rax, [rip + .L.str] instead of the mov:
.text
.globl greet
.p2align 4
.type greet,@function
greet:
mov eax, offset .L.str
ret
.Lfunc_end0:
.size greet, .Lfunc_end0-greet
.data
.globl counter
.p2align 2, 0x0
counter:
.long 5
.section .rodata.str1.1,"aMS",@progbits,1
.L.str:
.asciz "hi"
Line by line:
.globlexportsgreetandcounter— in C, anything not declaredstatic..p2align 4starts the function on a 16-byte boundary, and.p2align 2putscounteron a 4-byte boundary..typeand.sizerecord in the symbol table thatgreetis a function and how long it is. Debuggers and disassemblers use this..L.strand.Lfunc_end0start with.L, which makes them assembler-local: they're used while assembling and never reach the symbol table.nmon the object file listsgreetandcounter, but no.L.str. That's why a disassembler has to invent names likeloc_401020for jump targets: most labels never made it into the binary..rodata.str1.1with the"aMS"flags marks a section of mergeable strings, so the linker can keep a single copy of identical string literals.
Macros
A macro gives a name to a piece of text. Wherever the name appears later, the assembler replaces it with that text — the expansion — before assembling. Here is the same program in GAS and in NASM:
; GAS
.equ SYS_EXIT, 60
.macro exit code
mov edi, \code
mov eax, SYS_EXIT
syscall
.endm
.macro clamp0 reg ; if reg < 0, set it to 0
test \reg, \reg
jns 1f
xor \reg, \reg
1:
.endm
_start:
clamp0 eax
clamp0 ebx
.rept 3
nop
.endr
exit 0
; NASM
SYS_EXIT equ 60
%macro exit 1
mov edi, %1
mov eax, SYS_EXIT
syscall
%endmacro
%macro clamp0 1 ; if reg < 0, set it to 0
test %1, %1
jns %%done
xor %1, %1
%%done:
%endmacro
_start:
clamp0 eax
clamp0 ebx
times 3 nop
exit 0
A macro has a header with its name and parameters (\code in GAS, %1 in NASM), a body, and an end marker. Each call — clamp0 eax — is replaced by the body, with the parameters filled in.
Both files assemble to exactly the same 27 bytes:
0: 85 c0 test eax, eax
2: 79 02 jns 6
4: 31 c0 xor eax, eax
6: 85 db test ebx, ebx
8: 79 02 jns c
a: 31 db xor ebx, ebx
c: 90 nop
d: 90 nop
e: 90 nop
f: bf 00 00 00 00 mov edi, 0
14: b8 3c 00 00 00 mov eax, 60
19: 0f 05 syscall
There is no trace of the macros. clamp0 appears twice, fully written out, and SYS_EXIT has become 3c. Macro expansion happens entirely at assembly time, so from the binary alone you can't tell whether a macro was used. In reverse engineering, macros (and C's #define and inline functions) show up only as the same instruction pattern repeated in several places.
The label problem
clamp0 contains a label. Expand the macro twice with an ordinary label and you'd define it twice, which is an error. Each assembler has its own fix:
- GAS has numeric local labels:
1:can be defined any number of times, and1f/1bmean "the next1forward" and "the previous1backward". GAS also offers\@, a counter that is different in every expansion, to build unique names. - NASM has macro-local labels:
%%donebecomes a new name in each expansion. NASM even writes those names into the symbol table — here..@2.doneand..@3.done— so the disassembly above shows them.
Macro or function?
A macro and a function both let you write code once and use it many times, but they work very differently:
| Macro | Function (call) | |
|---|---|---|
| Handled | at assembly time | at run time |
| Copies of the code in the binary | one per use | one |
| Call and return cost | none | a call, a ret, and often a prologue |
| Recursion | only if it stops at assembly time | normal |
So macros suit short sequences used in hot code, where a call would cost more than the work itself. Functions suit anything long, since every macro use adds a full copy of the code.
Other assembly-time features
- Repetition:
.rept n….endrin GAS,times nor%repin NASM. - Conditional assembly:
.if/.else/.endifand.ifdefin GAS,%if/%ifdefin NASM. One source can then build different variants, such as 32-bit and 64-bit versions, or debug and release builds. - Inclusion:
.include "defs.s"or%include "defs.inc"pastes another file in place, usually one full of constants and macros.
Takeaways
- Directives are commands to the assembler. They produce data, switch sections, align, and control symbols, but no instructions.
- Sections separate code, initialized data, read-only data and zeroed data, and each gets its own permissions.
.bsstakes no space in the file. - Data directives store values little-endian. Alignment padding can show up as zeros — or as
90with NASM'salign. - Labels are local unless exported with
.globl/global..Llabels andequconstants never reach the binary. - Macros are text substitution at assembly time. They leave no trace in the machine code, only repeated patterns.