Quantization
The representation of model weights, activations, or intermediate values with reduced numerical precision to decrease memory use and often improve inference speed.
Generalisation and Measurement Boundary
Quantization reduces numerical precision and resource cost but is distinct from pruning, distillation, or architectural compression.
Related Machine-Learning Concepts
- BF16
- INT8
- Inference
- Product Quantization
Related technical article: Artificial Intelligence: Philosophy, Theory and Practice.
Direct source: The primary paper or official specification for Quantization is linked here for verification.