To the CPU, memory is just numbered bytes. It has no idea what a variable is. In C, a variable is a name the compiler gives to a range of those bytes, and its type tells the compiler two things: how many bytes the variable spans, and how to interpret them — as a signed or unsigned integer, a character, an address, a floating-point number. Once the program is compiled, the names are gone, and only the addresses and the instructions that use them remain.
This chapter follows a few variables from C down to bytes. How the hardware computes with those bytes — two's complement arithmetic, the IEEE 754 floating-point format — is the subject of the instruction-set level's chapter on data types the hardware understands.
How big is an int?
The C standard only sets minimums: char is at least 8 bits, short and int at least 16, long at least 32, and long long at least 64. The actual sizes are fixed by each platform's ABI. Measured with clang for five targets, in bytes:
| Type | Linux x86-64 | Linux ARM64 | macOS ARM64 | Windows x64 | Linux x86 (32-bit) |
|---|---|---|---|---|---|
char | 1 | 1 | 1 | 1 | 1 |
short | 2 | 2 | 2 | 2 | 2 |
int | 4 | 4 | 4 | 4 | 4 |
long | 8 | 8 | 8 | 4 | 4 |
long long | 8 | 8 | 8 | 8 | 8 |
pointer, size_t | 8 | 8 | 8 | 8 | 4 |
float / double | 4 / 8 | 4 / 8 | 4 / 8 | 4 / 8 | 4 / 8 |
long double | 16 | 16 | 8 | 8 | 12 |
Two things stand out. First, long is 8 bytes on 64-bit Unix systems but 4 bytes on 64-bit Windows. The names for these conventions are LP64 (Long and Pointer are 64 bits) and LLP64 (only Long Long and Pointer are). Code that stores a pointer in a long works on Linux and silently truncates addresses on Windows. That's why <stdint.h> exists: int32_t, uint64_t and uintptr_t have the same meaning everywhere.
Second, long double is whatever the platform decides: x86's old 80-bit format padded to 16 or 12 bytes, a true 128-bit format on ARM64 Linux, or just a double on Windows and macOS.
Signed, unsigned, and the same bits
An int and an unsigned int are the same 32 bits. Only the interpretation differs. With two's complement, the representation every modern CPU uses — and the only one C23 allows — the bit pattern 0xFFFFFFFF means −1 as an int and 4,294,967,295 as an unsigned int. Nothing in memory says which one it is; the compiler picks different instructions depending on the type, like jl (signed less-than) against jb (unsigned below) for comparisons.
When C mixes the two, the signed value is converted to unsigned. This is the source of a classic bug:
int i = -1;
unsigned int u = 1;
if (i < u) // false! i becomes 4294967295
…
gcc only warns about it with -Wextra (comparison of integer expressions of different signedness). Storing a value into a type too small for it keeps just the low bits: unsigned char c = 300; stores 44, because 300 is 0x12C and only the byte 0x2C fits.
Going the other way, arithmetic on small types is done in int: in a + b with two unsigned char values of 200 and 100, both are first promoted to int, and the result is 300, not 44.
Even plain char has a surprise: whether it's signed isn't fixed by the standard. It's signed on x86 and on Apple's ARM64, but unsigned on ARM64 Linux. The same program, char c = 200; printf("%d", c);, prints −56 on an x86 PC and 200 on a Raspberry Pi. When the difference matters, write signed char or unsigned char.
Byte order
A 4-byte int occupies four addresses. The question of which byte goes first is byte order, or endianness. x86, and ARM and RISC-V as they're used in practice, are little-endian: the least significant byte is at the lowest address. The value 0x11223344 is stored as the bytes 44 33 22 11.
C lets you see this by reading an int through a byte pointer:
- int main() {
- int x = 0x11223344;
- unsigned char *p = (unsigned char *)&x;
- return p[0];
- }
The program returns 68, which is 0x44: the lowest byte came first. On a big-endian machine it would return 0x11.
Tanenbaum's examples of big-endian machines are the SPARC and IBM mainframes. SPARC has since all but disappeared, and today big-endian lives on mainly in IBM's z mainframes, some network equipment, and in network byte order: TCP/IP headers are big-endian, which is why network code is full of htons and ntohl calls that swap bytes on little-endian machines. File formats pick one too: ELF and PE store their fields in the machine's order, PNG and Java class files in big-endian. The mixed-order problem the book describes — a record of text and integers copied between machines of opposite endianness — still exists; it just shows up in file parsers and network protocols now.
Alignment and padding
CPUs read memory most efficiently when a value sits at an address that's a multiple of its size: a 4-byte int at an address divisible by 4, an 8-byte double at a multiple of 8. This is alignment. x86 tolerates misaligned accesses at a small cost, but some architectures fault on them, and atomic operations generally require alignment. Each ABI fixes the alignment of each type, and compilers respect it.
For structs, that means the compiler inserts invisible padding bytes. Measured with clang for Linux x86-64:
struct A { char c; int x; char d; }; // sizeof = 12
struct B { int x; char c; char d; }; // sizeof = 8
In struct A, x must start at a multiple of 4, so three padding bytes follow c, and x is at offset 4. After d (offset 8), three more bytes pad the struct to 12, so that in an array of struct A every element's x stays aligned. struct B holds exactly the same fields in a different order and needs only 8 bytes. C never reorders struct fields for you: the order you write is the order in memory, which is what lets structs describe file headers and hardware registers exactly.
A struct { char c; double d; short s; } is 24 bytes on x86-64 Linux, but 16 on 32-bit x86 Linux, where the ABI only aligns double fields in structs to 4 bytes. Layout is part of the ABI, and differs between platforms.
Where each variable lives
A variable's storage duration decides where its bytes are. Compiling this file with clang and listing its symbols with nm shows the answer section by section:
int counter = 42; // D .data initialized global
int zeroed; // B .bss zero-initialized global
const int limit = 100; // R .rodata read-only
static int hidden = 7; // d .data local to this file
const char *msg = "hello"; // D .data (the pointer)
// "hello" itself goes to .rodata.str1.1
int bump(void) {
static int calls; // b .bss kept between calls
int local = counter + limit; // on the stack: no symbol at all
…
}
- Globals and
staticvariables live for the whole program, at fixed addresses, in the executable's data sections:.dataif they have a nonzero initial value,.bssif they start at zero, and.rodataif they'reconst..bsstakes no space in the file: the executable only records its size, and the loader provides zeroed memory. staticchanges two different things depending on where it's written. On a global (hidden), it makes the name private to the file:nmshows a lowercase letter, meaning a local symbol. On a variable inside a function (calls), it gives the variable a fixed address, so it keeps its value between calls. The compiler renames it to keep it unique: clang calls itbump.calls, gcccalls.0.- Local variables live in the function's stack frame and exist only during the call. They have no symbol: in the machine code,
localis just[rbp-4]or a register. - Heap memory, from
malloc, has no name at all — only a pointer to it. It's the subject of the chapter on structs, malloc and the heap.
The simulator shows all of these side by side. Step through the program below at the C level and watch the Globals panel (counter, limit, zeroed and the static calls) and the Variables panel (the struct p on the stack), then zoom in to see the same bytes in the stack and data views:
- int counter = 42;
- int zeroed;
- const int limit = 100;
- struct Point {
- char tag;
- int x;
- int y;
- };
- int bump(void) {
- static int calls;
- calls++;
- return calls;
- }
- int main() {
- struct Point p;
- p.tag = 'A';
- p.x = 0x11223344;
- p.y = -1;
- bump();
- bump();
- zeroed = counter + limit;
- return sizeof(p);
- }
The program returns 12, the size of struct Point with its padding. In the generated assembly, p.tag is written at [rbp-12], p.x at [rbp-8] and p.y at [rbp-4]: the three bytes at [rbp-11] to [rbp-9] are padding that nothing ever touches. The static calls became a global named calls.0 in .bss, and ends at 2. The simulator's compiler keeps things simple and puts limit in .data; a real compiler would put it in .rodata, where writing to it crashes the program.
Takeaways
- A variable is a named range of bytes; its type gives the size and the interpretation. The names disappear at compilation.
- Sizes come from the platform's ABI:
longis 8 bytes on Linux and macOS (LP64) but 4 on Windows (LLP64). Use<stdint.h>types when the size matters. - Signed and unsigned values share the same bits (two's complement). Mixing them converts to unsigned, so
-1 < 1uis false. Plaincharis signed on x86 but unsigned on ARM64 Linux. - x86, ARM and RISC-V run little-endian:
0x11223344is stored44 33 22 11. Networks and many file formats use big-endian. - Compilers insert padding to keep fields aligned; field order changes a struct's size (12 against 8 bytes here).
- Globals and statics live in
.data,.bssor.rodataat fixed addresses; locals live on the stack with no symbol; heap memory is reached only through pointers.