Skip to content

Foundations · Chapter F.4

Floating point (IEEE 754)

From scientific notation to IEEE 754: how a real number becomes sign, exponent and fraction bits (5.75 and −0.15625 converted by hand), rounding to nearest-even, machine epsilon, special values, why == on floats is risky, and how Kahan summation recovers lost digits.

Integers are exact, but their range is narrow: a 64-bit integer stops at about 9.2 × 10¹⁸, and has nothing between 0 and 1. Science and engineering need numbers like Avogadro's constant, 6.022 × 10²³, and Planck's constant, 6.626 × 10⁻³⁴, in the same program. Writing both out as fixed-point numbers would take more than 50 digits, most of them meaningless zeros. Floating point solves this the way scientists do on paper: store a few significant digits, plus an exponent that says where the point goes.

This chapter builds the idea from scratch and converts a few numbers by hand. The data types chapter covers the hardware side (the formats table, fp16 and bfloat16, why 0.1 + 0.2 isn't 0.3), and the two chapters complement each other.

Scientific notation

In scientific notation, a number is written as a significand times a power of ten: 6.022 × 10²³. The number of digits in the significand sets the precision; the size of the exponent sets the range. The two are independent, which is the whole point.

Imagine a toy decimal format with a 3-digit significand and an exponent from −99 to 99. It can hold numbers from 1.00 × 10⁻⁹⁹ up to 9.99 × 10⁹⁹, but only three significant digits of each. Between 1.23 × 10⁵ and the next number, 1.24 × 10⁵, there's a gap of 1,000. Between 1.23 × 10⁻² and 1.24 × 10⁻², the gap is 0.0001. The gaps grow with the numbers, but relative to the number, they stay about the same size: a bit under 1 %. That's the defining trade-off of floating point: constant relative precision over an enormous range.

A number is normalized when its first significant digit is nonzero: 1.23 × 10⁵, not 0.123 × 10⁶ or 12.3 × 10⁴. Normalizing gives each number a single representation.

Binary floating point

Computers do the same in base 2: a number is ± significand × 2^exponent, with the significand written in binary. In binary, a normalized significand always starts with 1, the only nonzero binary digit, so it's always 1.something. Since that leading 1 is always there, it doesn't need to be stored. This hidden bit gives one bit of precision for free.

The IEEE 754 standard, published in 1985 and followed by practically every CPU since, fixes the layout. A 32-bit float (binary32) has:

FieldBitsStores
sign10 for +, 1 for −
exponent8the exponent + 127 (excess-127, see the two's complement chapter)
fraction23the bits after the point in 1.fraction

A 64-bit double (binary64) has 1, 11 and 52 bits, with a bias of 1023. The value of a normal number is:

value = (−1)^sign × 1.fraction × 2^(exponent − bias)

The exponent is biased rather than two's complement so that, apart from the sign, larger floats have larger bit patterns: positive floats compare correctly as plain integers.

Converting 5.75 by hand

  1. Convert to binary. 5 is 101. 0.75 is ½ + ¼, so .11. 5.75 = 101.11₂.
  2. Normalize. Move the point two places left: 1.0111 × 2².
  3. Sign: positive, so 0.
  4. Exponent: 2 + 127 = 129 = 1000 0001.
  5. Fraction: the bits after the leading 1, padded to 23 bits: 0111 0000 0000 0000 0000 000.

Put together:

0 10000001 01110000000000000000000
= 0100 0000 1011 1000 0000 0000 0000 0000
= 0x40B80000

The explorer below stores the number. Check the fields against the steps above, then click bits to see how each one changes the value:

Floating point · 5.75 as a float

Try it: Type a number or click an example, switch the format (top right), or click any bit to flip it.

1 sign bit8 exponent bits (bias 127)23 fraction bitsclick a bit to flip it
0x40b80000
stored as
32 bits, float (binary32)
normal
kind
5.75
reads back as
shortest decimal that round-trips
exact
typed value
stored without error
sign
0 → positive
exponent
129 − 127 = 2
significand
1.fraction = 1.4375
value
+1.4375 × 2^2 = 5.75

Rounding is round-to-nearest, ties-to-even, as in hardware. Sums like 0.1+0.2 are computed in double precision first, as a program would, then stored in the chosen format.

A negative example: −0.15625

  1. Convert to binary. 0.15625 = 5/32 = ⅛ + 1/32 = 0.00101₂.
  2. Normalize. Move the point three places right: 1.01 × 2⁻³.
  3. Sign: negative, so 1.
  4. Exponent: −3 + 127 = 124 = 0111 1100.
  5. Fraction: 0100 0000 0000 0000 0000 000.
1 01111100 01000000000000000000000 = 0xBE200000

Both numbers are exact: their binary expansions end. Type -0.15625 into the explorer above to check. Numbers like 0.1, whose binary expansion repeats forever, have to be cut off, which brings us to rounding.

Rounding

Between two neighboring floats there's a gap. Its size is called one ulp (unit in the last place): the value of the last fraction bit at that exponent. When a result falls in a gap, it must be rounded to one of the two neighbors. IEEE 754's default is round to nearest, ties to even: pick the nearer neighbor, and if the result is exactly halfway, pick the one whose last bit is 0.

Ties are rare in practice, but whole numbers just above 2²⁴ show them clearly. A float has 24 significant bits, so between 2²⁴ = 16,777,216 and 2²⁵ the gap is 2: only even integers exist. 16,777,217 is exactly halfway between 16,777,216 and 16,777,218. The tie goes to the neighbor with an even last bit, 16,777,216. But 16,777,219, halfway between 16,777,218 and 16,777,220, rounds up, because this time 16,777,220 has the even last bit:

Floating point · A tie rounded to even

Try it: Type a number or click an example, switch the format (top right), or click any bit to flip it.

1 sign bit8 exponent bits (bias 127)23 fraction bitsclick a bit to flip it
0x4b800002
stored as
32 bits, float (binary32)
normal
kind
16777220
reads back as
shortest decimal that round-trips
rounded
typed value
the nearest representable value
sign
0 → positive
exponent
151 − 127 = 24
significand
1.fraction = 1.0000002384185791015625
value
+1.0000002384185791015625 × 2^24 = 16777220

Rounding is round-to-nearest, ties-to-even, as in hardware. Sums like 0.1+0.2 are computed in double precision first, as a program would, then stored in the chosen format.

Rounding to even rather than always up avoids a systematic upward drift when many results are rounded. The standard also defines three directed rounding modes (toward zero, toward +∞ and toward −∞), used for interval arithmetic, where you compute a lower and an upper bound that are guaranteed to contain the true result.

Machine epsilon

The gap just above 1.0 is called machine epsilon. For a float, the next number after 1 is 1 + 2⁻²³ = 1.00000011920928955078125, so epsilon is 2⁻²³ ≈ 1.19 × 10⁻⁷. For a double it's 2⁻⁵² ≈ 2.22 × 10⁻¹⁶. C's <float.h> defines them as FLT_EPSILON and DBL_EPSILON; Python has sys.float_info.epsilon.

Epsilon measures relative precision: correctly rounding any result to the nearest float changes it by at most half an ulp, a relative error of at most epsilon / 2. That's where the rules of thumb "about 7 significant decimal digits for a float, about 16 for a double" come from. It also means that adding something smaller than half an ulp does nothing at all: in single precision, 1 + 2⁻²⁴ rounds back to exactly 1 (another tie, resolved to even), and in double precision, 10¹⁶ + 1 − 10¹⁶ is 0, because the gap between doubles near 10¹⁶ is 2.

Special values

IEEE 754 reserves the smallest and largest exponent fields for special cases:

Exponent fieldFractionMeaning
all zeroszero±0
all zerosnonzerosubnormal numbers: 0.fraction × 2⁻¹²⁶, no hidden 1
all oneszero±∞
all onesnonzeroNaN (not a number)

Subnormals fill the gap between the smallest normal float, 2⁻¹²⁶ ≈ 1.18 × 10⁻³⁸, and zero, down to 2⁻¹⁴⁹ ≈ 1.4 × 10⁻⁴⁵. They give up precision gradually instead of jumping to zero. The 1985 standard called them denormalized numbers; the 2008 revision renamed them subnormal.

Dividing a finite number by zero gives infinity only if the number is nonzero: in C on the machine this chapter was written on, 1.0/0.0 printed inf, -1.0/0.0 printed -inf, but 0.0/0.0 printed nan. So do ∞ − ∞, ∞ / ∞ and the square root of a negative number.

The 80-bit extended format is not one of the standard's basic formats. The 1985 standard defined single and double and left "extended" formats loosely specified; the 80-bit format is the x87 unit's implementation of one. The 2008 revision added the 16-bit and 128-bit binary formats, decimal floating point, and the fused multiply-add (a × b + c with a single rounding), which every modern CPU now has.

Why comparing floats with == is risky

Because nearly every operation rounds, two computations that are equal in mathematics often differ in their last bit:

0.1 + 0.2 == 0.3                → false
(0.1 + 0.2) + 0.3               → 0.6000000000000001
0.1 + (0.2 + 0.3)               → 0.6

Floating-point addition is not associative: the order of operations changes the result. A compiler therefore can't reorder float additions unless you allow it (with flags like -ffast-math), and a parallel sum can give slightly different answers depending on how the work was split.

The usual fix is to compare with a tolerance, relative to the size of the numbers: |a − b| ≤ rel × max(|a|, |b|), plus a small absolute tolerance for values near zero. Python's math.isclose(0.1 + 0.2, 0.3) does exactly that, and returns True. Two more traps: a NaN is unequal to everything, itself included (x != x is the classic NaN test), and −0 == +0 is true even though the bits differ.

Exact comparison is fine when the values are exact: small integers, and sums of powers of two like 5.75, are represented exactly and computed exactly.

Summing many numbers: Kahan's trick

Adding a long list of numbers loses precision at every step, and the errors accumulate. The extreme case: add 1.0 to a float 20 million times. Once the total reaches 16,777,216, adding 1 is a tie that rounds back down, and the sum stops growing. Measured with clang on the M2, the naive loop ends at 16,777,216.

Kahan summation (compensated summation), due to William Kahan, the main author of IEEE 754, keeps a second variable holding the rounding error of the last addition and feeds it back into the next one:

float sum = 0.0f, c = 0.0f;        /* c = the part that got lost */
for (int i = 0; i < n; i++) {
    float y = x[i] - c;            /* add back what was lost last time */
    float t = sum + y;             /* big + small: low bits of y are lost */
    c = (t - sum) - y;             /* (t - sum) is what was actually added */
    sum = t;
}

With the same 20 million ones, still entirely in float, it returns exactly 20,000,000. Adding 0.1f ten million times gives 1,087,937 naively and 1,000,000 with Kahan's method. Compile it without -ffast-math: that flag lets the compiler treat float arithmetic as associative, simplify (t - sum) - y to zero and reorder the loop. Built with it, the two loops over 0.1f returned the same number, 1,005,958.5625: the compensation was gone, and so was the exact answer.

Python applies the same idea for you: since Python 3.12, the built-in sum() uses a compensated algorithm, so sum([0.1] * 10) is 1.0, while a plain loop of s += 0.1 ends at 0.9999999999999999. For an exactly rounded sum, math.fsum goes further.

Takeaways

  • Floating point stores a significand and an exponent, separating precision from range, with roughly constant relative precision.
  • IEEE 754 binary32: 1 sign bit, 8 exponent bits in excess-127, 23 fraction bits after a hidden 1. binary64: 1, 11 (excess-1023), 52.
  • By hand: convert to binary, normalize to 1.f × 2^e, store e + bias and f. 5.75 = 0x40B80000, −0.15625 = 0xBE200000.
  • Results are rounded to nearest, ties to even: 16,777,217 becomes 16,777,216 but 16,777,219 becomes 16,777,220.
  • Machine epsilon is 2⁻²³ ≈ 1.19 × 10⁻⁷ for float, 2⁻⁵² ≈ 2.22 × 10⁻¹⁶ for double: about 7 and 16 decimal digits.
  • Special values: ±0, subnormals (once called "denormalized"), ±∞, and NaN, which 0/0 produces, not ∞.
  • Float arithmetic isn't associative, so compare with a tolerance, never with == on computed results; Kahan summation recovers the digits a long sum loses.

Foundations

  1. F.1Levels of abstraction and a short history of computers
  2. F.2Binary and hexadecimal
  3. F.3Two’s complement and signed numbers
  4. F.4Floating point (IEEE 754)
  5. F.5Characters, ASCII and Unicode
  6. F.6Endianness
  7. F.7Parity, Hamming codes and error correction
  8. F.8Units: kilo, kibi and friends