Memory is addressed byte by byte, but most numbers are bigger than a byte. A 32-bit integer takes four consecutive addresses. Which of its four bytes goes at the lowest address? There are two natural answers, and both have been used by major machines. This choice is called byte order, or endianness.
Big-endian and little-endian
Take the 32-bit value 0x11223344, stored at address 100. Its most significant byte is 0x11, its least significant byte 0x44.
- A big-endian machine stores the most significant byte first, at the lowest address: the "big end" comes first.
- A little-endian machine stores the least significant byte first.
| Address | 100 | 101 | 102 | 103 |
|---|---|---|---|---|
| big-endian | 11 | 22 | 33 | 44 |
| little-endian | 44 | 33 | 22 | 11 |
In both cases the value is the same, and the word's address is 100. Only the layout of the bytes differs. Inside a byte, there's no question to ask: software can't address individual bits, so a byte is just a number from 0 to 255 either way.
The names come from Gulliver's Travels, where two nations go to war over which end of a boiled egg to break. Danny Cohen borrowed them in a 1980 note on this very argument, On Holy Wars and a Plea for Peace. His point stands: neither order is better, but mixing them up causes real bugs.
Why each order has fans
Big-endian matches the way people write numbers: a hex dump of memory reads left to right like the number itself, and comparing two big-endian unsigned numbers byte by byte, from the lowest address, gives the same order as comparing their values.
Little-endian has an advantage for the hardware and for low-level code: a value's low byte is always at its address, whatever the value's size. Reading a 64-bit variable as a 32-, 16- or 8-bit one, at the same address, gives its low 32, 16 or 8 bits. And multi-word arithmetic, which starts from the least significant end to propagate carries, starts at the lowest address.
Neither advantage is decisive, and today the choice is mostly settled by history.
Who uses what
| Architecture | Byte order |
|---|---|
| x86, x86-64 | little-endian |
| ARM, ARM64 | bi-endian, but little-endian in practice: Android, iOS, macOS, Windows and Linux distributions all run it little-endian |
| RISC-V | little-endian |
| IBM Z mainframes | big-endian |
| POWER | bi-endian; Linux moved to little-endian (ppc64le) with POWER8 in 2014, while AIX and IBM i remain big-endian |
| SPARC, older MIPS, 68000 | big-endian |
Bi-endian processors can run in either order, selected by a control bit that the operating system sets: on ARM64, a bit in a system control register sets the order of data accesses. The classic examples of big-endian machines were the SPARC and IBM mainframes. SPARC has all but disappeared, so in practice nearly every CPU you'll meet runs little-endian. Big-endian lives on mainly in mainframes, in network protocols and in file formats.
Seeing the bytes
An x86 instruction that loads a constant carries that constant inside the instruction, in memory, so it follows the machine's byte order too. Here is the encoding of a 64-bit mov:
Try it: Type an instruction in the box (or switch to hex bytes), or click an example; click a field to expand its bits.
objdump shows48 b8 88 77 66 55 44 33 22 11 movabs rax, 0x1122334455667788
REX = 0100WRXB. W: 64-bit operands.
0100fixed1W0R0X0B- fixed=0100
- marks a REX prefix (0x40–0x4F)
- W=1
- 64-bit operand size
- R=0
- reg field not extended
- X=0
- index not extended
- B=0
- rm/base not extended
B8+r → MOVABS r64, imm64 (dst = src): the only x86 instruction with a full 64-bit immediate (10 bytes).
10111opcode000reg- opcode=10111
- B8+r
- reg=000
- rax
64-bit mode, Intel syntax. Where several encodings are valid, this picks the one clang emits; the text is what objdump -d -M intel prints. Branch targets are written relative to the instruction: jmp $+0x10.
After the prefix 48 and the opcode b8, the eight bytes of the constant appear backward: 88 77 66 55 44 33 22 11, lowest byte first. Every disassembler shows x86 immediates and addresses this way.
The same thing happens with data. The program below stores 0x11223344 in memory, then reads it back whole, as single bytes and as a 16-bit half:
Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.
- .data
- word: .long 0x11223344
- port: .short 0
- .text
- mov eax, DWORD PTR [rip+word] ; the whole word: 0x11223344
- movzx ebx, BYTE PTR [rip+word] ; the byte at the lowest address
- movzx ecx, BYTE PTR [rip+word+3] ; the byte at the highest address
- movzx esi, WORD PTR [rip+word] ; the two lowest bytes
- mov edx, 443 ; 0x01BB
- rol dx, 8 ; htons(443): swap the two bytes
- mov WORD PTR [rip+port], dx
The whole word reads back as 0x11223344. The byte at the word's address is 0x44 and the byte at address + 3 is 0x11: little-endian. The 16-bit load at the same address gets 0x3344, the low half: the little-endian advantage above. The last three instructions are the subject of the next section. The variables chapter shows the same experiment from C, with a byte pointer, and the ISA overview loads 8 bytes from a misaligned address and gets them assembled in little-endian order.
The problem: moving data between machines
Byte order is invisible as long as data stays on one machine. It matters when bytes travel: over a network, or in a file written on one machine and read on another.
Suppose a big-endian machine sends a record (the name "Ada" followed by the 32-bit integer 1815, 0x00000717) byte by byte to a little-endian machine. The bytes arrive in the same order they were sent. The string is fine, because text is a sequence of single bytes. The integer is not: the bytes 00 00 07 17, read little-endian, mean 0x17070000, 386,334,720. Swapping the bytes of every 4-byte group would fix the integer and scramble the string. There is no blind fix: whoever reads the data must know which bytes form which multi-byte field, and swap exactly those.
The practical solution is the one every protocol and file format uses: the format specifies the byte order, and each program converts every multi-byte field when reading and writing it.
Network byte order
The Internet protocols chose big-endian, called network byte order: every multi-byte field in an IP, TCP or UDP header (addresses, ports, lengths) is sent most significant byte first. The C library provides conversion functions: htons and htonl (host to network, short and long), ntohs and ntohl for the reverse.
On a big-endian machine they do nothing. On a little-endian one they swap bytes. Port 443, 0x01BB, must go on the wire as the bytes 01 BB; in a little-endian register that's the value 0xBB01, so htons(443) returns 47,873, as Python's socket.htons(443) confirms on the M2. That's what the rol dx, 8 in the demo above does: rotating a 16-bit value by 8 bits swaps its two bytes. It is exactly the instruction clang generates for htons on x86-64; for 32-bit values it uses bswap, x86's byte-reversal instruction. On ARM64, the same functions compile to rev16 and rev. Recent x86 chips also have movbe, a load or store that swaps bytes on the way.
Byte order in files
File formats make the same choice, and a hex dump shows it. PNG is big-endian: here are the first bytes of a 640 × 480 image. The width and height are the 4-byte fields at offsets 16 and 20:
00000000: 8950 4e47 0d0a 1a0a 0000 000d 4948 4452 .PNG........IHDR
00000010: 0000 0280 0000 01e0 0802 0000 00ba b34b ...............K
00 00 02 80 is 640 and 00 00 01 e0 is 480. Read little-endian by mistake, the width would be 2,147,614,720.
Executable formats record their byte order explicitly. An ELF file starts with 7f 45 4c 46 (\x7F then ELF); its sixth byte is 1 for little-endian, 2 for big-endian. The start of /bin/ls in an ARM64 Linux container:
000000 7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
000010 03 00 b7 00
02 means 64-bit and 01 means little-endian; at offset 0x12, the machine type b7 00 is 0x00B7, 183, the code for AArch64, stored low byte first.
macOS shows both orders in one file. /bin/ls is a universal binary holding an x86-64 and an ARM64 version, and its header is big-endian: it starts ca fe ba be, the magic number 0xCAFEBABE. Each version inside it is a Mach-O file in the machine's own order, starting with cf fa ed fe: the magic number 0xFEEDFACF written little-endian. The executable files chapter describes these formats.
Takeaways
- Endianness is the order of a multi-byte value's bytes in memory. Big-endian puts the most significant byte at the lowest address; little-endian puts the least significant byte there.
0x11223344is stored44 33 22 11on a little-endian machine. - x86-64, ARM64 (in practice) and RISC-V are little-endian; IBM Z is big-endian; bi-endian CPUs like ARM and POWER can run either way.
- Byte order only matters when data moves between machines or into files. Formats and protocols fix the order, and programs convert each multi-byte field; blindly swapping everything breaks text.
- Network byte order is big-endian;
htons/htonlconvert, compiling torol/bswapon x86 andrev16/revon ARM64.htons(443)is 47,873 on a little-endian machine. - PNG is big-endian; ELF and Mach-O declare their order in their header; a macOS universal binary has a big-endian header around little-endian programs.