Skip to content

Foundations Β· Chapter F.5

Characters, ASCII and Unicode

How text becomes numbers: ASCII and its tricks, 8-bit code pages and mojibake, Unicode code points and planes, the UTF-8 algorithm bit by bit (with a live encoder), UTF-16 surrogate pairs, normalization, and why an emoji can be one character, five code points and 18 bytes.

Memory holds numbers, so text has to be stored as numbers too. A character code assigns a number to each character: 65 for A, 233 for Γ©, 128,512 for πŸ˜€. Everyone exchanging text has to agree on the same code, which is why character codes are standards. Three ideas need to be kept apart:

  • the character, an abstract unit like "Latin small letter e with acute";
  • its code point, the number the standard assigns to it;
  • its encoding, the bytes used to store that number in memory or in a file.

The glyph you see on screen (the shape drawn by a font) is a fourth thing, and not part of the character code at all.

ASCII

ASCII, the American Standard Code for Information Interchange, was first published in 1963. It uses 7 bits, so it has 128 codes:

CodesContents
0x00–0x1F32 control characters: NUL (0x00), tab (0x09), line feed LF (0x0A), carriage return CR (0x0D), escape (0x1B), …
0x20space
0x30–0x39the digits 0 to 9
0x41–0x5AA to Z
0x61–0x7Aa to z
the rest of 0x21–0x7Epunctuation and symbols: ! 0x21, @ 0x40, ~ 0x7E, …
0x7FDEL

The layout was designed with care, and programs still exploit it:

  • Upper and lower case differ by exactly one bit, 0x20: A is 0x41 = 0100 0001, a is 0x61 = 0110 0001. Setting or clearing bit 5 changes the case of an ASCII letter.
  • The digits are in order starting at 0x30, so the value of a digit character c is c - '0': '7' - '0' is 7.
  • Letters are in alphabetical order, so for plain English text, sorting by code sorts alphabetically, with all uppercase letters before all lowercase ones.

Most of the control characters were meant for teleprinters and data links and are rarely used now, but a few are everywhere. C marks the end of a string with NUL, the byte 0. Unix ends lines with LF; Windows uses the two bytes CR LF, a remnant of teleprinters that needed one code to return the carriage and another to advance the paper. Git's line-ending settings exist because of that difference.

Eight bits and code pages

Computers store a character in a byte, so ASCII leaves 128 codes, 0x80 to 0xFF, unused. Different vendors and standards filled them differently, each set of 256 characters being a code page:

  • ISO 8859-1, or Latin-1, added the accented letters of Western European languages: Γ© is 0xE9, ΓΌ is 0xFC.
  • Windows-1252, Microsoft's variant of Latin-1, also puts printable characters in 0x80–0x9F, where Latin-1 has control codes: € is 0x80.
  • Other parts of ISO 8859 covered Central European languages, Cyrillic, Greek, Arabic, Hebrew and so on, all reusing the same 128 codes.

Two standards are easy to confuse here. ISO 646 is the 7-bit international version of ASCII, whose national variants replaced a few symbols like # or [ with local letters; it added no characters. The 8-bit Latin-1 is ISO 8859-1. The original IBM PC's code page, 437, did have smiley faces, but at codes 1 and 2, among the control characters; its upper half, codes 128 to 255, held accented letters, Greek letters and box-drawing characters.

The code-page approach has a fatal flaw: the bytes don't say which code page they're in. Decode text with the wrong one and you get mojibake: garbled characters. The classic example today is UTF-8 text read as Latin-1 or Windows-1252: the two bytes of Γ©, C3 A9, come out as é. If you've seen "café" on a web page, that's what happened.

Unicode: one number per character

Unicode, first published in 1991, gives every character in every writing system a single code point, written U+ followed by at least four hex digits: A is U+0041, Γ© is U+00E9, € is U+20AC, the crab emoji πŸ¦€ is U+1F980. The first 256 code points are the same as Latin-1, so the first 128 are ASCII.

Unicode was first designed as a 16-bit code with 65,536 code points and no multibyte characters. That design has been outdated since Unicode 2.0 in 1996. Code points now run from U+0000 to U+10FFFF, a 21-bit space of 1,114,112 values, divided into 17 planes of 65,536:

PlaneRangeContents
0, Basic Multilingual PlaneU+0000–U+FFFFalmost all modern scripts, common symbols
1, Supplementary MultilingualU+10000–U+1FFFFhistoric scripts, musical and math symbols, most emoji
2 and 3U+20000–U+3FFFFrarer CJK (Chinese, Japanese, Korean) ideographs
15 and 16U+F0000–U+10FFFFprivate use

Unicode 16.0, published in 2024 and the version Python 3.14's unicodedata module ships with, defines 154,998 characters; counting private-use code points and the reserved surrogates, 294,579 code points are assigned. Most of the space is still empty.

A code point is just a number; it still has to be stored as bytes. Unicode defines three encodings: UTF-8, UTF-16 and UTF-32.

UTF-8

UTF-8 stores each code point in 1 to 4 bytes. The number of leading 1 bits in the first byte says how many bytes the character takes; every following byte starts with 10. The remaining bits, marked x, hold the code point:

Code pointsBits neededByte 1Byte 2Byte 3Byte 4
U+0000–U+007F70xxxxxxx
U+0080–U+07FF11110xxxxx10xxxxxx
U+0800–U+FFFF161110xxxx10xxxxxx10xxxxxx
U+10000–U+10FFFF2111110xxx10xxxxxx10xxxxxx10xxxxxx

Encoding Γ©, U+00E9, step by step:

U+00E9 = 000 1110 1001             11 bits needed β†’ 2-byte form
split:   00011 | 101001            5 bits, then 6 bits
fill:    110 00011  10 101001
bytes:   1100 0011  1010 1001   = C3 A9

And the crab, U+1F980, which needs 17 bits and so the 4-byte form:

U+1F980 = 0 0001 1111 1001 1000 0000        (21 bits)
split:    000 | 011111 | 100110 | 000000
bytes:    11110000 10011111 10100110 10000000 = F0 9F A6 80

The program below does the same thing in C, with shifts and masks, and prints the bytes for four characters:

Live Β· A UTF-8 encoder

Try it: Press Step to run one instruction, Run to animate or Continue to finish; the L2–L7 buttons zoom in and out one level at a time.

C source Β· click a line number for a breakpoint
  1. int encode(int cp, unsigned char *out) {
  2. if (cp < 0x80) { /* 0xxxxxxx */
  3. out[0] = cp;
  4. return 1;
  5. }
  6. if (cp < 0x800) { /* 110xxxxx 10xxxxxx */
  7. out[0] = 0xC0 | (cp >> 6);
  8. out[1] = 0x80 | (cp & 0x3F);
  9. return 2;
  10. }
  11. if (cp < 0x10000) { /* 1110xxxx 10xxxxxx 10xxxxxx */
  12. out[0] = 0xE0 | (cp >> 12);
  13. out[1] = 0x80 | ((cp >> 6) & 0x3F);
  14. out[2] = 0x80 | (cp & 0x3F);
  15. return 3;
  16. }
  17. out[0] = 0xF0 | (cp >> 18); /* 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx */
  18. out[1] = 0x80 | ((cp >> 12) & 0x3F);
  19. out[2] = 0x80 | ((cp >> 6) & 0x3F);
  20. out[3] = 0x80 | (cp & 0x3F);
  21. return 4;
  22. }
  23. int main() {
  24. int text[4] = { 0x41, 0xE9, 0x20AC, 0x1F980 }; /* A Γ© € crab */
  25. unsigned char buf[4];
  26. int total = 0;
  27. for (int i = 0; i < 4; i++) {
  28. int n = encode(text[i], buf);
  29. printf("U+%04X ->", text[i]);
  30. for (int j = 0; j < n; j++)
  31. printf(" %02x", buf[j]);
  32. printf("\n");
  33. total += n;
  34. }
  35. return total;
  36. }
step 0
Loading emulator…
Your program as you wrote it: the current line, its variables by name, and its output.

It prints 41, c3 a9, e2 82 ac and f0 9f a6 80, and returns 10, the total number of bytes: four characters, ten bytes.

The design has several useful properties:

  • ASCII text is valid UTF-8, byte for byte, and no byte of a multibyte character is below 0x80. So a program that only looks for /, \n or NUL works unchanged on UTF-8.
  • It is self-synchronizing: continuation bytes always start with 10 and first bytes never do, so from any position you can find the start of the next character.
  • Sorting UTF-8 strings byte by byte sorts them by code point.
  • Each character has exactly one valid encoding. The "overlong" 2-byte form C0 80 for NUL, for instance, is invalid, and decoders must reject it. Accepting it has caused security bugs, where a filter checked for one byte but the decoder produced it from another sequence.

UTF-8 was first designed with forms of up to 6 bytes, which could encode 31-bit values, about two billion characters; since 2003 (RFC 3629), UTF-8 is limited to 4 bytes and U+10FFFF, to match UTF-16's range. UTF-8 is now the encoding of the overwhelming majority of web pages, the usual encoding of source code and of file names on Linux and macOS, and the one JSON requires for data exchanged between systems.

UTF-16 and surrogate pairs

UTF-16 stores each code point in one or two 16-bit units. The Basic Multilingual Plane takes one unit, equal to the code point. For code points above U+FFFF, 0x10000 is subtracted, and the remaining 20 bits are split into two halves of 10 bits, each added to a reserved range: the high half to 0xD800, the low half to 0xDC00. The result is a surrogate pair:

U+1F980 βˆ’ 0x10000 = 0xF980 = 0000111110 0110000000   (20 bits)
high: 0xD800 + 0000111110 = 0xD83E
low:  0xDC00 + 0110000000 = 0xDD80

That's why the code points U+D800 to U+DFFF are reserved and never assigned to characters. UTF-16 was the natural successor of the original 16-bit Unicode, so the systems designed in that era use it internally: Windows, Java and JavaScript. In JavaScript, "πŸ¦€".length is 2, because length counts 16-bit units. UTF-32 stores every code point in 4 bytes: simple, but wasteful, and rarely used for storage.

One character, two spellings: normalization

Unicode can often represent the same text in more than one way. Γ© exists as a single code point, U+00E9, but it can also be written as e followed by U+0301, COMBINING ACUTE ACCENT, which attaches to the letter before it. Both display identically. In Python:

>>> import unicodedata as u
>>> nfc, nfd = u.normalize("NFC", "Γ©"), u.normalize("NFD", "Γ©")
>>> len(nfc), len(nfd), nfc == nfd
(1, 2, False)

Two strings that look the same compare as different, and a search for one misses the other. The fix is normalization: convert both to the same form before comparing. NFC composes characters wherever possible; NFD decomposes them. Most text is in NFC, but macOS file systems have historically stored names decomposed, which is a classic source of "file not found" bugs when names travel between systems.

What the user calls a character: grapheme clusters

Some things a reader sees as one character are several code points. The family emoji πŸ‘¨β€πŸ‘©β€πŸ‘§ is three people joined by two invisible zero-width joiners (U+200D). A flag like πŸ‡«πŸ‡· is two "regional indicator" letters, F and R. Unicode calls what the user perceives as one character a grapheme cluster, and "how long is this string?" has four different answers:

TextGraphemesCode pointsUTF-16 unitsUTF-8 bytes
Γ© (decomposed)1223
πŸ‡«πŸ‡·1248
πŸ‘¨β€πŸ‘©β€πŸ‘§15818

These counts come from Python (len counts code points) and Node.js (length counts UTF-16 units; Intl.Segmenter counts graphemes). The practical consequences: truncating a string to n code points can cut an emoji in half, reversing it code point by code point breaks accents and flags, and a text field's "maximum length" means different things in different languages. When the user-visible count matters, use a grapheme segmenter.

At the CPU level, none of this exists: instructions see bytes, as the data types chapter notes.

Takeaways

  • A character code maps characters to numbers (code points); an encoding turns code points into bytes.
  • ASCII has 128 codes in 7 bits; case differs by bit 5 (0x20) and digits start at 0x30. 8-bit code pages like Latin-1 reused 0x80–0xFF differently, causing mojibake when mixed up.
  • Unicode is no longer 16-bit: code points run to U+10FFFF (17 planes). Unicode 16.0 defines 154,998 characters.
  • UTF-8 uses 1–4 bytes with the patterns 0xxxxxxx, 110xxxxx 10xxxxxx, … Γ© = C3 A9, πŸ¦€ = F0 9F A6 80. It is ASCII-compatible and self-synchronizing; its original 5- and 6-byte forms were removed in 2003.
  • UTF-16 uses surrogate pairs above U+FFFF (πŸ¦€ = D83E DD80); it's the internal format of Windows, Java and JavaScript.
  • The same text can have different code points (normalization: NFC vs NFD), and one visible character can be many code points (grapheme clusters: πŸ‘¨β€πŸ‘©β€πŸ‘§ is 5 code points and 18 bytes).

Foundations

  1. F.1Levels of abstraction and a short history of computers
  2. F.2Binary and hexadecimal
  3. F.3Two’s complement and signed numbers
  4. F.4Floating point (IEEE 754)
  5. F.5Characters, ASCII and Unicode
  6. F.6Endianness
  7. F.7Parity, Hamming codes and error correction
  8. F.8Units: kilo, kibi and friends