Mixed precision means different parts of a model’s computation use different numeric formats — e.g., weights stored in FP16/BF16, but certain ops (like accumulation) done in FP32 to avoid numerical instability. It’s a speed/memory vs. accuracy tradeoff. This is a property of how you execute, not what the weights are.
Native precision: the format the model’s weights were originally trained in/released as (e.g., FP32, BF16, FP16). This is a property of the checkpoint.
You could run a model:
- At native precision — uniformly, in whatever format it was trained/released in
- In mixed precision — some ops downcast, others not, regardless of native format
- Quantized — deliberately compressed below native (e.g., native BF16 → quantized to INT8/FP4/NVFP4) for inference efficiency
Quantization (“quant”) generically means reducing precision further — mapping weights/activations from FP16/FP32 down to lower-bit representations (INT8, FP8, INT4, FP4) to shrink memory footprint and increase throughput on hardware that has native support for those formats. Comes at some accuracy cost, mitigated by calibration/scaling techniques.
NVFP4 is NVIDIA’s specific FP4 format for Blackwell GPUs:
- 4-bit floating point (1 sign, 2 exponent, 1 mantissa bit — E2M1)
- Uses micro-scaling: small blocks of elements (16 values) share a scale factor (stored in FP8 E4M3), rather than one scale per tensor — this preserves more dynamic range/accuracy than naive FP4
- Designed to let Blackwell hit much higher throughput (2x+ over FP8) for inference, at accuracy close to FP8 for well-calibrated models
- Competes with MXFP4 (the OCP/Microsoft-backed open microscaling FP4 standard) — NVFP4 is NVIDIA’s proprietary variant tuned for their Tensor Cores
So “run at quant” = run in some reduced-precision mode. “Run at NVFP4” = specifically use NVIDIA’s 4-bit floating-point format with block-wise scaling, requires Blackwell-generation hardware (or emulation) and a quantization/calibration pass on the model to convert it.
P.S., Floating point does NOT use a lookup table like ASCII. It uses a formula applied to the bits, split into three fields:
- [ sign | exponent | mantissa ]
- For FP32 (32 bits = 4 bytes): [1 bit sign][8 bits exponent][23 bits mantissa]
- Example — the number 6.5 in FP32:
0 10000001 10100000000000000000000
sign exponent(8) mantissa (23) - Value = 1.625 × 2^2 = 6.5 ✓
- Sign bit 0 → positive
- Exponent field 10000001 = 129, minus a bias of 127 → actual exponent = 2
- Mantissa 101… → interpreted as 1.101 in binary = 1.625
Each parameter (weight) in the model is one such 32-bit pattern, decoded via that formula into a real number like 6.5 or -0.0031. So:
- 1 parameter = 1 instance of that bit-pattern = 32 bits = 4 bytes (for FP32)
- 1 billion parameters = 1 billion of these 32-bit patterns stored back-to-back in memory = 4 billion bytes = 4 GB
Shrinking to FP16/BF16 just means: use a shorter bit pattern (16 bits) with the same sign/exponent/mantissa idea, just fewer exponent/mantissa bits allocated — so each parameter takes 2 bytes instead of 4, and the formula has less range/precision to work with.