FP8: the definition
FP8 is a family of 8-bit floating-point number formats, mainly E4M3 and E5M2, used to store and multiply neural-network values with half the bits of FP16 or BF16. It is part of a broader shift to lower-precision formats, down to 4-bit FP4, that trade numeric precision for speed and memory savings.
The key points
- Every number format trades range and precision against bits; fewer bits mean less memory traffic and more math per second on hardware that supports them.
- AI has moved from FP32 to TF32, BF16 and FP16, then to FP8 on Hopper-era hardware and FP4 variants on Blackwell and newer accelerators.
- Block-scaled formats such as MXFP8, MXFP4 and NVFP4 attach a shared scale to small groups of values, which makes very short formats usable.
- Low precision works when paired with scaling, higher-precision accumulation and a few sensitive operations kept in wider formats.
Why number formats matter for AI
A floating-point number has three parts: a sign bit, exponent bits that set how large or small the value can be, and mantissa bits that set how finely values are spaced. More exponent bits give more range, more mantissa bits give more precision, and the total bit count determines how much memory and bandwidth each value costs. [3][1]
Neural networks tolerate imprecision surprisingly well, which is why hardware vendors keep adding narrower formats. The 2017 mixed-precision training paper noted that half-precision math throughput on GPUs of that era was 2 to 8 times higher than single precision, and that storing values in 16 bits nearly halves memory requirements. NVIDIA says FP8 on Hopper again halves storage and doubles throughput relative to FP16 or BF16. [2][5]
FP32, TF32, BF16 and FP16
FP32, the standard single-precision format, uses an 8-bit exponent. NVIDIA introduced TF32 with the Ampere A100 as a 19-bit tensor core format with 1 sign bit, the same 8-bit exponent as FP32 and a 10-bit mantissa. Because the range matches FP32, programs written for FP32 can use TF32 tensor cores without code changes, and NVIDIA quoted TF32 on A100 as 10 times faster than FP32 fused multiply-add on the previous V100. [3][4]
BF16, or bfloat16, also keeps an 8-bit exponent but has only 7 mantissa bits, so it covers the FP32 range with half the storage. Google's TPUs multiply in bfloat16 and accumulate in FP32 by default, and Google notes that the smaller format helps most on operations limited by memory bandwidth. [3][6]
FP16 spends more bits on the mantissa, 10 of them, but its normalised exponent range runs only from minus 14 to 15, and any value smaller than 2 to the power of minus 24 becomes zero. The mixed-precision paper found that roughly 5 percent of weight gradients in one network fell below that threshold. Its remedy, still widely used, keeps an FP32 master copy of the weights and multiplies the loss by a scale factor so small gradients stay representable. [2]
FP8: E4M3 and E5M2
The common FP8 formats were set out in a 2022 paper by authors from NVIDIA, Arm and Intel. E4M3 uses 4 exponent and 3 mantissa bits and extends its range by giving up infinities and using a single bit pattern for NaN. E5M2 uses 5 exponent and 2 mantissa bits and follows IEEE 754 conventions for special values. The authors reported FP8 training matching 16-bit quality across CNNs, RNNs and Transformers, including language models of up to 175 billion parameters. [1]
Because FP8 has so few bits, each tensor needs its own scale factor instead of the single loss scale used with FP16. NVIDIA's Transformer Engine uses E4M3 in the forward pass, where precision matters more, and E5M2 for gradients in the backward pass, where range matters more. Its delayed-scaling recipe picks each scale from the largest absolute values seen over recent iterations. [9]
Suppose a weight tensor's largest magnitude is 3.2. E4M3 can represent values up to 448, so the software multiplies the tensor by a scale of about 130 before casting, placing its largest value just under the format's limit and using the available precision, then divides the scale back out after the matrix multiply. Without that step, most small values would round to zero or lose their distinguishing bits.
Block scaling and the MX formats
Per-tensor scales struggle when a tensor mixes very large and very small values. The Open Compute Project's Microscaling (MX) specification, released as version 1.0 in September 2023 with contributions from AMD, Arm, Intel, Meta, Microsoft, NVIDIA and Qualcomm, attaches a shared scale to every block of 32 elements. The scale is stored in an 8-bit power-of-two format called E8M0. [10]
The MX specification defines MXFP8 in E5M2 and E4M3 variants, MXFP6 in two variants, MXFP4 using 4-bit E2M1 elements, and MXINT8. Vendors can support a subset and still comply. On NVIDIA Blackwell, Transformer Engine's MXFP8 recipe uses these 32-element blocks, which lets E4M3 be used for all tensors while reducing saturation compared with a single per-tensor scale. [10][9]
FP4: MXFP4 and NVFP4
A 4-bit E2M1 value has 1 sign, 2 exponent and 1 mantissa bit and can represent magnitudes only up to 6, so block scaling is essential. NVIDIA's NVFP4 format, introduced with Blackwell, uses smaller blocks of 16 values, stores each block's scale in FP8 E4M3 rather than a power of two, and adds a second FP32 scale per tensor. NVIDIA says this reduces memory by about 3.5 times versus FP16 and 1.8 times versus FP8, with accuracy on DeepSeek-R1-0528 within about 1 percent of FP8. [9][11]
FP4 has so far been used mostly for inference, but training results are appearing. A 2025 NVIDIA paper trained a 12-billion-parameter language model on 10 trillion tokens in NVFP4 and reported loss and downstream accuracy comparable to an FP8 baseline. The recipe relied on random Hadamard transforms to tame outliers, two-dimensional scaling of weights, stochastic rounding of gradients, and keeping a few sensitive layers in higher precision. [12][9]
INT8 and integer quantization
Integer formats were the first low-precision workhorse for inference. Google's first TPU was built around 8-bit integer multipliers, and a 2020 NVIDIA study showed an 8-bit integer quantization workflow could keep accuracy within 1 percent of the floating-point baseline across vision, speech and language models, including harder cases such as MobileNets and BERT-large. [13][14]
Large language models made INT8 harder. The LLM.int8 paper found that transformers develop systematic outlier features at scale, and handled them by running those few dimensions in 16-bit while keeping more than 99.9 percent of the multiplications in 8-bit, halving inference memory. The FP8 formats paper also reported that FP8 post-training quantization worked for models that resisted INT8 quantization, one reason FP8 has become popular for serving. [15][1]
The accuracy trade-offs
Low precision is never applied everywhere. Transformer Engine keeps operations such as normalisation layers and exponentials in higher precision, and its NVFP4 recipe recommends leaving the last few layers of large language models in a wider format such as MXFP8. Accumulation inside matrix units is also done in higher precision, as with the FP32 accumulation Google uses behind bfloat16 multiplies. [9][6]
Formats are also a hardware question. A format only speeds things up when the accelerator has native support for it: FP8 tensor cores arrived with Hopper, NVFP4 with Blackwell, and TPU FP8 and FP4 support varies by generation. Buying newer hardware for its low-precision throughput only pays off if your software stack and model can actually use those formats without unacceptable quality loss.
What this means when choosing compute
Headline petaflops on spec sheets are usually quoted at the lowest supported precision, so a chip's FP4 number may be four times its BF16 number. When comparing GPUs, match the precision you will really run: BF16 or FP8 for most training, FP8 or FP4 for optimised inference. On Kovara you can compare GPU models and their current cloud prices side by side, and ask Kova which options support FP8 or FP4 for your workload.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1What do E4M3 and E5M2 mean?
They describe how the 8 bits are split after the sign bit: E4M3 has 4 exponent bits and 3 mantissa bits, giving more precision, while E5M2 has 5 exponent bits and 2 mantissa bits, giving more range. NVIDIA's Transformer Engine lists maximum magnitudes of 448 for E4M3 and 57,344 for E5M2.
2Is BF16 the same as FP16?
No. Both use 16 bits, but BF16 keeps the 8-bit exponent of FP32 and has only 7 mantissa bits, so it covers the same range as FP32 with less precision. FP16 has 10 mantissa bits but a much narrower range, which is why FP16 training usually needs loss scaling.
3What is the difference between MXFP4 and NVFP4?
Both store elements as 4-bit E2M1 values with a shared scale per block. MXFP4, from the Open Compute Project MX specification, uses blocks of 32 elements with a power-of-two E8M0 scale. NVFP4, introduced with NVIDIA Blackwell, uses blocks of 16 elements with an FP8 E4M3 scale plus a second FP32 scale for the whole tensor.
4Does lower precision always reduce model quality?
Not necessarily. Published studies show FP8 training matching 16-bit results and an NVFP4-trained 12-billion-parameter model matching an FP8 baseline, but those results depend on careful scaling, keeping sensitive layers in higher precision and, for 4-bit, extra techniques such as stochastic rounding.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- arXiv · FP8 Formats for Deep Learning (Micikevicius et al., 2022) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- arXiv · Mixed Precision Training (Micikevicius et al., ICLR 2018) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- NVIDIA · NVIDIA Ampere GPU Architecture Tuning Guide ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Ampere Architecture In-Depth ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Hopper Architecture In-Depth ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- Google Cloud · The bfloat16 numerical format ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · TPU7x (Ironwood) ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud Blog · TPU 8t and TPU 8i technical deep dive ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA Transformer Engine · Using FP8 and FP4 with Transformer Engine ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Open Compute Project · OCP Microscaling Formats (MX) Specification v1.0 ↗ (opens in a new tab)Industry specification · Checked 29 September 2026
- NVIDIA Technical Blog · Introducing NVFP4 for efficient and accurate low-precision inference ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- arXiv · Pretraining Large Language Models with NVFP4 (NVIDIA, 2025) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- arXiv · In-Datacenter Performance Analysis of a Tensor Processing Unit (Jouppi et al., ISCA 2017) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- arXiv · Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation (Wu et al., 2020) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- arXiv · LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers et al., 2022) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
