BF16
A 16-bit floating-point format with an 8-bit exponent that preserves the dynamic range of FP32 while reducing memory and compute cost.
Related Concepts
- Quantization
- FP16
- INT8
- Mixed Precision
A 16-bit floating-point format with an 8-bit exponent that preserves the dynamic range of FP32 while reducing memory and compute cost.