Memory holds numbers, so text has to be stored as numbers too. A character code assigns a number to each character: 65 for A, 233 for Γ©, 128,512 for π. Everyone exchanging text has to agree on the same code, which is why character codes are standards. Three ideas need to be kept apart:
- the character, an abstract unit like "Latin small letter e with acute";
- its code point, the number the standard assigns to it;
- its encoding, the bytes used to store that number in memory or in a file.
The glyph you see on screen (the shape drawn by a font) is a fourth thing, and not part of the character code at all.
ASCII
ASCII, the American Standard Code for Information Interchange, was first published in 1963. It uses 7 bits, so it has 128 codes:
| Codes | Contents |
|---|---|
| 0x00β0x1F | 32 control characters: NUL (0x00), tab (0x09), line feed LF (0x0A), carriage return CR (0x0D), escape (0x1B), β¦ |
| 0x20 | space |
| 0x30β0x39 | the digits 0 to 9 |
| 0x41β0x5A | A to Z |
| 0x61β0x7A | a to z |
| the rest of 0x21β0x7E | punctuation and symbols: ! 0x21, @ 0x40, ~ 0x7E, β¦ |
| 0x7F | DEL |
The layout was designed with care, and programs still exploit it:
- Upper and lower case differ by exactly one bit, 0x20:
Ais 0x41 =0100 0001,ais 0x61 =0110 0001. Setting or clearing bit 5 changes the case of an ASCII letter. - The digits are in order starting at 0x30, so the value of a digit character
cisc - '0':'7' - '0'is 7. - Letters are in alphabetical order, so for plain English text, sorting by code sorts alphabetically, with all uppercase letters before all lowercase ones.
Most of the control characters were meant for teleprinters and data links and are rarely used now, but a few are everywhere. C marks the end of a string with NUL, the byte 0. Unix ends lines with LF; Windows uses the two bytes CR LF, a remnant of teleprinters that needed one code to return the carriage and another to advance the paper. Git's line-ending settings exist because of that difference.
Eight bits and code pages
Computers store a character in a byte, so ASCII leaves 128 codes, 0x80 to 0xFF, unused. Different vendors and standards filled them differently, each set of 256 characters being a code page:
- ISO 8859-1, or Latin-1, added the accented letters of Western European languages: Γ© is 0xE9, ΓΌ is 0xFC.
- Windows-1252, Microsoft's variant of Latin-1, also puts printable characters in 0x80β0x9F, where Latin-1 has control codes: β¬ is 0x80.
- Other parts of ISO 8859 covered Central European languages, Cyrillic, Greek, Arabic, Hebrew and so on, all reusing the same 128 codes.
Two standards are easy to confuse here. ISO 646 is the 7-bit international version of ASCII, whose national variants replaced a few symbols like # or [ with local letters; it added no characters. The 8-bit Latin-1 is ISO 8859-1. The original IBM PC's code page, 437, did have smiley faces, but at codes 1 and 2, among the control characters; its upper half, codes 128 to 255, held accented letters, Greek letters and box-drawing characters.
The code-page approach has a fatal flaw: the bytes don't say which code page they're in. Decode text with the wrong one and you get mojibake: garbled characters. The classic example today is UTF-8 text read as Latin-1 or Windows-1252: the two bytes of Γ©, C3 A9, come out as ΓΒ©. If you've seen "cafΓΒ©" on a web page, that's what happened.
Unicode: one number per character
Unicode, first published in 1991, gives every character in every writing system a single code point, written U+ followed by at least four hex digits: A is U+0041, Γ© is U+00E9, β¬ is U+20AC, the crab emoji π¦ is U+1F980. The first 256 code points are the same as Latin-1, so the first 128 are ASCII.
Unicode was first designed as a 16-bit code with 65,536 code points and no multibyte characters. That design has been outdated since Unicode 2.0 in 1996. Code points now run from U+0000 to U+10FFFF, a 21-bit space of 1,114,112 values, divided into 17 planes of 65,536:
| Plane | Range | Contents |
|---|---|---|
| 0, Basic Multilingual Plane | U+0000βU+FFFF | almost all modern scripts, common symbols |
| 1, Supplementary Multilingual | U+10000βU+1FFFF | historic scripts, musical and math symbols, most emoji |
| 2 and 3 | U+20000βU+3FFFF | rarer CJK (Chinese, Japanese, Korean) ideographs |
| 15 and 16 | U+F0000βU+10FFFF | private use |
Unicode 16.0, published in 2024 and the version Python 3.14's unicodedata module ships with, defines 154,998 characters; counting private-use code points and the reserved surrogates, 294,579 code points are assigned. Most of the space is still empty.
A code point is just a number; it still has to be stored as bytes. Unicode defines three encodings: UTF-8, UTF-16 and UTF-32.
UTF-8
UTF-8 stores each code point in 1 to 4 bytes. The number of leading 1 bits in the first byte says how many bytes the character takes; every following byte starts with 10. The remaining bits, marked x, hold the code point:
| Code points | Bits needed | Byte 1 | Byte 2 | Byte 3 | Byte 4 |
|---|---|---|---|---|---|
| U+0000βU+007F | 7 | 0xxxxxxx | |||
| U+0080βU+07FF | 11 | 110xxxxx | 10xxxxxx | ||
| U+0800βU+FFFF | 16 | 1110xxxx | 10xxxxxx | 10xxxxxx | |
| U+10000βU+10FFFF | 21 | 11110xxx | 10xxxxxx | 10xxxxxx | 10xxxxxx |
Encoding Γ©, U+00E9, step by step:
U+00E9 = 000 1110 1001 11 bits needed β 2-byte form
split: 00011 | 101001 5 bits, then 6 bits
fill: 110 00011 10 101001
bytes: 1100 0011 1010 1001 = C3 A9
And the crab, U+1F980, which needs 17 bits and so the 4-byte form:
U+1F980 = 0 0001 1111 1001 1000 0000 (21 bits)
split: 000 | 011111 | 100110 | 000000
bytes: 11110000 10011111 10100110 10000000 = F0 9F A6 80
The program below does the same thing in C, with shifts and masks, and prints the bytes for four characters:
Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2βL7 buttons zoom in and out one level at a time.
- int encode(int cp, unsigned char *out) {
- if (cp < 0x80) { /* 0xxxxxxx */
- out[0] = cp;
- return 1;
- }
- if (cp < 0x800) { /* 110xxxxx 10xxxxxx */
- out[0] = 0xC0 | (cp >> 6);
- out[1] = 0x80 | (cp & 0x3F);
- return 2;
- }
- if (cp < 0x10000) { /* 1110xxxx 10xxxxxx 10xxxxxx */
- out[0] = 0xE0 | (cp >> 12);
- out[1] = 0x80 | ((cp >> 6) & 0x3F);
- out[2] = 0x80 | (cp & 0x3F);
- return 3;
- }
- out[0] = 0xF0 | (cp >> 18); /* 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx */
- out[1] = 0x80 | ((cp >> 12) & 0x3F);
- out[2] = 0x80 | ((cp >> 6) & 0x3F);
- out[3] = 0x80 | (cp & 0x3F);
- return 4;
- }
- int main() {
- int text[4] = { 0x41, 0xE9, 0x20AC, 0x1F980 }; /* A Γ© β¬ crab */
- unsigned char buf[4];
- int total = 0;
- for (int i = 0; i < 4; i++) {
- int n = encode(text[i], buf);
- printf("U+%04X ->", text[i]);
- for (int j = 0; j < n; j++)
- printf(" %02x", buf[j]);
- printf("\n");
- total += n;
- }
- return total;
- }
It prints 41, c3 a9, e2 82 ac and f0 9f a6 80, and returns 10, the total number of bytes: four characters, ten bytes.
The design has several useful properties:
- ASCII text is valid UTF-8, byte for byte, and no byte of a multibyte character is below 0x80. So a program that only looks for
/,\nor NUL works unchanged on UTF-8. - It is self-synchronizing: continuation bytes always start with
10and first bytes never do, so from any position you can find the start of the next character. - Sorting UTF-8 strings byte by byte sorts them by code point.
- Each character has exactly one valid encoding. The "overlong" 2-byte form
C0 80for NUL, for instance, is invalid, and decoders must reject it. Accepting it has caused security bugs, where a filter checked for one byte but the decoder produced it from another sequence.
UTF-8 was first designed with forms of up to 6 bytes, which could encode 31-bit values, about two billion characters; since 2003 (RFC 3629), UTF-8 is limited to 4 bytes and U+10FFFF, to match UTF-16's range. UTF-8 is now the encoding of the overwhelming majority of web pages, the usual encoding of source code and of file names on Linux and macOS, and the one JSON requires for data exchanged between systems.
UTF-16 and surrogate pairs
UTF-16 stores each code point in one or two 16-bit units. The Basic Multilingual Plane takes one unit, equal to the code point. For code points above U+FFFF, 0x10000 is subtracted, and the remaining 20 bits are split into two halves of 10 bits, each added to a reserved range: the high half to 0xD800, the low half to 0xDC00. The result is a surrogate pair:
U+1F980 β 0x10000 = 0xF980 = 0000111110 0110000000 (20 bits)
high: 0xD800 + 0000111110 = 0xD83E
low: 0xDC00 + 0110000000 = 0xDD80
That's why the code points U+D800 to U+DFFF are reserved and never assigned to characters. UTF-16 was the natural successor of the original 16-bit Unicode, so the systems designed in that era use it internally: Windows, Java and JavaScript. In JavaScript, "π¦".length is 2, because length counts 16-bit units. UTF-32 stores every code point in 4 bytes: simple, but wasteful, and rarely used for storage.
One character, two spellings: normalization
Unicode can often represent the same text in more than one way. Γ© exists as a single code point, U+00E9, but it can also be written as e followed by U+0301, COMBINING ACUTE ACCENT, which attaches to the letter before it. Both display identically. In Python:
>>> import unicodedata as u
>>> nfc, nfd = u.normalize("NFC", "Γ©"), u.normalize("NFD", "Γ©")
>>> len(nfc), len(nfd), nfc == nfd
(1, 2, False)
Two strings that look the same compare as different, and a search for one misses the other. The fix is normalization: convert both to the same form before comparing. NFC composes characters wherever possible; NFD decomposes them. Most text is in NFC, but macOS file systems have historically stored names decomposed, which is a classic source of "file not found" bugs when names travel between systems.
What the user calls a character: grapheme clusters
Some things a reader sees as one character are several code points. The family emoji π¨βπ©βπ§ is three people joined by two invisible zero-width joiners (U+200D). A flag like π«π· is two "regional indicator" letters, F and R. Unicode calls what the user perceives as one character a grapheme cluster, and "how long is this string?" has four different answers:
| Text | Graphemes | Code points | UTF-16 units | UTF-8 bytes |
|---|---|---|---|---|
| Γ© (decomposed) | 1 | 2 | 2 | 3 |
| π«π· | 1 | 2 | 4 | 8 |
| π¨βπ©βπ§ | 1 | 5 | 8 | 18 |
These counts come from Python (len counts code points) and Node.js (length counts UTF-16 units; Intl.Segmenter counts graphemes). The practical consequences: truncating a string to n code points can cut an emoji in half, reversing it code point by code point breaks accents and flags, and a text field's "maximum length" means different things in different languages. When the user-visible count matters, use a grapheme segmenter.
At the CPU level, none of this exists: instructions see bytes, as the data types chapter notes.
Takeaways
- A character code maps characters to numbers (code points); an encoding turns code points into bytes.
- ASCII has 128 codes in 7 bits; case differs by bit 5 (0x20) and digits start at 0x30. 8-bit code pages like Latin-1 reused 0x80β0xFF differently, causing mojibake when mixed up.
- Unicode is no longer 16-bit: code points run to U+10FFFF (17 planes). Unicode 16.0 defines 154,998 characters.
- UTF-8 uses 1β4 bytes with the patterns
0xxxxxxx,110xxxxx 10xxxxxx, β¦ Γ© =C3 A9, π¦ =F0 9F A6 80. It is ASCII-compatible and self-synchronizing; its original 5- and 6-byte forms were removed in 2003. - UTF-16 uses surrogate pairs above U+FFFF (π¦ =
D83E DD80); it's the internal format of Windows, Java and JavaScript. - The same text can have different code points (normalization: NFC vs NFD), and one visible character can be many code points (grapheme clusters: π¨βπ©βπ§ is 5 code points and 18 bytes).