IEEE-754 Floating Point, Explained

After reading this you will know exactly why 0.1 + 0.2 prints 0.30000000000000004, how to read the sign, exponent and mantissa of any float, and when the rounding error will bite you.

What a float actually stores

A float does not store a number. It stores a recipe for reconstructing one number out of a fixed set of representable values. The IEEE-754 standard (the current edition is IEEE 754-2019, first published in 1985) packs that recipe into a fixed number of bits.

A double precision float (float64, the type behind every JavaScript Number and every C double) uses 64 bits split into three fields: 1 sign bit, 11 exponent bits, and 52 mantissa (fraction) bits. A single precision float (float32) uses 1 + 8 + 23 bits.

The hook: type 0.1 into the explorer. The exact decimal value stored is not 0.1. It is 0.1000000000000000055511151231257827021181583404541015625. That number is the closest float64 to 0.1, and it is off by about 5.55 \times 10^{-18}. The mismatch is not a bug in your language. It is the direct result of trying to write one tenth in binary, where it repeats forever just as one third repeats forever in decimal.

When this matters and when it does not

Floating point is the right tool for physics, graphics, statistics and anything measured from the analog world, where the input already carries more uncertainty than 2^{-52}. In those settings a relative error near 1.1 \times 10^{-16} is invisible.

It is the wrong tool for money and for exact counting past a threshold. Never store dollars as a float. Store integer cents, or use a decimal type. Financial code that adds 0.1 a thousand times drifts away from the exact answer because each addition rounds.

Do not compare floats with == after arithmetic. 0.1 + 0.2 == 0.3 is false in almost every language. Compare with a tolerance instead: check that the absolute difference is below some small epsilon you choose for the problem, not the machine epsilon.

The formula and the intuition

For a normal (non-special) float64, the stored value is:

v = (-1)^{s} \times \left(1 + \sum_{i=1}^{52} b_i \, 2^{-i}\right) \times 2^{(e - 1023)}

Here s is the sign bit (0 for positive, 1 for negative). The value e is the 11-bit exponent read as an unsigned integer from 0 to 2047, and 1023 is the fixed bias that lets the exponent represent both large and small magnitudes. Each b_i is one mantissa bit. The leading 1 + is the implicit bit: for normal numbers the standard assumes a 1 before the binary point and does not store it, buying one extra bit of precision for free.

The mantissa gives 53 bits of significance in total (52 stored plus the implicit 1). That is why float64 carries roughly 53 \times \log_{10} 2 \approx 15.95 significant decimal digits.

Two exponent values are reserved. When e = 0 the number is zero or subnormal (the implicit bit becomes 0). When e = 2047 the number is infinity (mantissa all zero) or NaN (mantissa nonzero).

Worked example: why 0.1 + 0.2 misses 0.3

Adding two numbers that do not exist exactly

Load the 0.1 + 0.2 preset. The explorer shows the exact stored value of each operand.

  1. The nearest float64 to 0.1 is slightly high: 0.1000000000000000055511151231257827…
  2. The nearest float64 to 0.2 is also slightly high: 0.2000000000000000111022302462515654…
  3. Add them exactly and you get 0.3000000000000000166533453693773481…
  4. That sum is not representable either, so it rounds to the nearest float64, which is 0.3000000000000000444089209850062616…
  5. The nearest float64 to 0.3 typed directly is 0.2999999999999999888977697537484346…, a different bit pattern.

The two results differ by exactly 1 ULP (one unit in the last place) at that magnitude, which is 2^{-54} \approx 5.55 \times 10^{-17}. Print the sum and you see 0.30000000000000004: the shortest decimal that still round-trips to that specific bit pattern. The rounding errors in the two operands did not cancel the rounding of 0.3, so the bit patterns disagree.

Every finite decimal whose only prime factors are 2 and 5 is exact in decimal, but only powers of 2 in the denominator are exact in binary. So 0.5, 0.25 and 0.75 are exact floats, while 0.1, 0.2 and 0.3 are not.

The gap between floats grows with magnitude

Floats are not evenly spaced. They are dense near zero and sparse far from it. The spacing at a given magnitude is 1 ULP, and it doubles every time the exponent increases by 1. Concretely, for a value in the range [2^{k}, 2^{k+1}), the ULP for float64 is 2^{k-52}.

Near 1.0 the exponent gives k = 0, so the gap is 2^{-52} \approx 2.22 \times 10^{-16}. That specific number, the gap just above 1.0, is machine epsilon for float64. Near 2^{53} the gap is 2^{53-52} = 2^{1} = 2, and near 2^{52} it is exactly 1.0.

The gap between adjacent float64 values in powers of 2. It crosses 1.0 (y = 0) exactly at 2^{52} and reaches 2.0 at 2^{53}.

This is why 2^{53} + 1 rounds back to 2^{53}: the next representable integer above 2^{53} is 2^{53} + 2. There is no float64 for the odd integer in between. That is the origin of JavaScript's Number.MAX_SAFE_INTEGER, which equals 2^{53} - 1 = 9007199254740991, and the reason BigInt exists.

Reading the special cases

Signed zero
There are two zeros, +0 and -0, with different sign bits. They compare equal (+0 == -0 is true), but 1/+0 is +\infty and 1/-0 is -\infty.
Subnormal
When the exponent field is all zeros and the mantissa is nonzero, the implicit bit becomes 0. These fill the gap between zero and the smallest normal float, 2^{-1022} \approx 2.23 \times 10^{-308}, down to the smallest positive float64, 2^{-1074} \approx 4.94 \times 10^{-324}. They trade precision for reach and are often 10 to 100 times slower in hardware.
Infinity
Exponent all ones, mantissa all zeros. Produced by overflow or by 1.0/0.0.
NaN
Exponent all ones, mantissa nonzero. The nonzero bits are the payload. NaN is never equal to anything, including itself: NaN == NaN is false. That self-inequality is the standard test for NaN.

Try it: flip one bit and watch the value

The most instructive move is to flip a single exponent bit and see the value jump by a factor of 2, then flip a single low mantissa bit and see it barely twitch.

Start from 1.0, whose float64 bits are sign 0, exponent 01111111111 (1023), mantissa all zeros. Flip the lowest mantissa bit and the value becomes 1.0000000000000002, a change of 2^{-52}. Flip the lowest exponent bit instead and the value becomes 2.0, doubling.

Common mistakes

Rounding for display is not the same as rounding the value. Printing 0.1 to 1 decimal shows 0.1, but the stored value is still the long tail from the first section. Formatting hides the error; it does not remove it.

Summing many floats in naive order loses precision. Adding a tiny number to a huge running total can drop the tiny number entirely once the total's ULP exceeds it. Kahan summation or pairwise summation recovers most of the lost bits.

Treating machine epsilon as a fixed tolerance is wrong. Machine epsilon (2^{-52}) is the gap near 1.0 only. Near 10^{6} the gap is about 1.16 \times 10^{-10}, far larger. Scale your tolerance to the magnitude of the numbers you compare.

Related tools

If you are studying how machines represent and manipulate data at the bit level, two neighbors help. The Karnaugh Map Minimizer shows how boolean logic behind the arithmetic units gets simplified. The Digital Logic Gate Simulator lets you build the adders and shifters that actually perform floating point operations from gates and flip-flops.

Frequently asked questions

Why does 0.1 + 0.2 equal 0.30000000000000004?

Because none of the three values is exact in binary. The stored 0.1 and 0.2 are both slightly high, their exact sum is not representable, and it rounds to a bit pattern one ULP above the float64 for 0.3. Printing that pattern with the shortest round-tripping decimal gives 0.30000000000000004.

What is machine epsilon exactly?

For float64 it is 2^{-52} \approx 2.22 \times 10^{-16}, the gap between 1.0 and the next larger float. For float32 it is 2^{-23} \approx 1.19 \times 10^{-7}. It measures relative precision, not an absolute tolerance.

Why is Number.MAX_SAFE_INTEGER equal to 2^53 minus 1?

Above 2^{53} the ULP is 2, so odd integers cannot be represented. Below and up to 2^{53} - 1 every integer has an exact float64, so that is the largest integer you can add 1 to and get a distinct value.

Are subnormal numbers safe to use?

They are correct but slow. On many CPUs an operation on a subnormal is 10 to 100 times slower than on a normal float, because it drops out of the fast hardware path. Some compilers offer a flush-to-zero mode that treats subnormals as zero to avoid the penalty.

Is float64 arithmetic deterministic across machines?

Basic operations (add, subtract, multiply, divide, square root) are correctly rounded and portable per the standard. Transcendental functions like sin and exp are not required to be correctly rounded, so their last bit can differ between libraries.