Terminology of LLM 01

Mixed precision means different parts of a model’s computation use different numeric formats — e.g., weights stored in FP16/BF16, but certain ops (like accumulation) done in FP32 to avoid numerical instability. It’s a speed/memory vs. accuracy tradeoff. This is a property of how you execute, not what the weights are.

Native precision: the format the model’s weights were originally trained in/released as (e.g., FP32, BF16, FP16). This is a property of the checkpoint.

You could run a model:

  • At native precision — uniformly, in whatever format it was trained/released in
  • In mixed precision — some ops downcast, others not, regardless of native format
  • Quantized — deliberately compressed below native (e.g., native BF16 → quantized to INT8/FP4/NVFP4) for inference efficiency

Quantization (“quant”) generically means reducing precision further — mapping weights/activations from FP16/FP32 down to lower-bit representations (INT8, FP8, INT4, FP4) to shrink memory footprint and increase throughput on hardware that has native support for those formats. Comes at some accuracy cost, mitigated by calibration/scaling techniques.

NVFP4 is NVIDIA’s specific FP4 format for Blackwell GPUs:

  • 4-bit floating point (1 sign, 2 exponent, 1 mantissa bit — E2M1)
  • Uses micro-scaling: small blocks of elements (16 values) share a scale factor (stored in FP8 E4M3), rather than one scale per tensor — this preserves more dynamic range/accuracy than naive FP4
  • Designed to let Blackwell hit much higher throughput (2x+ over FP8) for inference, at accuracy close to FP8 for well-calibrated models
  • Competes with MXFP4 (the OCP/Microsoft-backed open microscaling FP4 standard) — NVFP4 is NVIDIA’s proprietary variant tuned for their Tensor Cores

So “run at quant” = run in some reduced-precision mode. “Run at NVFP4” = specifically use NVIDIA’s 4-bit floating-point format with block-wise scaling, requires Blackwell-generation hardware (or emulation) and a quantization/calibration pass on the model to convert it.

P.S., Floating point does NOT use a lookup table like ASCII. It uses a formula applied to the bits, split into three fields:

  • [ sign | exponent | mantissa ]
  • For FP32 (32 bits = 4 bytes): [1 bit sign][8 bits exponent][23 bits mantissa]
  • Example — the number 6.5 in FP32:
    0 10000001 10100000000000000000000
    sign exponent(8) mantissa (23)
  • Value = 1.625 × 2^2 = 6.5 ✓
  • Sign bit 0 → positive
  • Exponent field 10000001 = 129, minus a bias of 127 → actual exponent = 2
  • Mantissa 101… → interpreted as 1.101 in binary = 1.625

Each parameter (weight) in the model is one such 32-bit pattern, decoded via that formula into a real number like 6.5 or -0.0031. So:

  • 1 parameter = 1 instance of that bit-pattern = 32 bits = 4 bytes (for FP32)
  • 1 billion parameters = 1 billion of these 32-bit patterns stored back-to-back in memory = 4 billion bytes = 4 GB

Shrinking to FP16/BF16 just means: use a shorter bit pattern (16 bits) with the same sign/exponent/mantissa idea, just fewer exponent/mantissa bits allocated — so each parameter takes 2 bytes instead of 4, and the formula has less range/precision to work with.

Leave a Reply