Model quantization is a technique that lowers the numeric precision of a model's weights and activations to reduce memory footprint and accelerate inference on CPUs, GPUs, and edge hardware. Every production inference stack worth mentioning, NVIDIA TensorRT, Google LiteRT, and PyTorch ExecuTorch among them, ships first-class support for model quantization because the math is unforgiving: a 32-bit floating-point weight tensor that occupies four bytes per parameter becomes a single byte at INT8 and half a byte at INT4, before any kernel optimization kicks in. The result is a smaller binary, a lower memory footprint, and faster compute on integer-friendly silicon, with an accuracy cost that is usually small and always measurable. The trade-offs split cleanly along two axes: when the conversion happens (post-training quantization or quantization-aware training, often abbreviated PTQ and QAT) and what the target precision is (INT8, INT4, half-precision FP16, FP8, FP4, or a mixed scheme). The rest of this explainer walks through the mechanics, the format landscape, and how the major frameworks differ in what they support.
What Model Quantization Does to a Neural Network

Model quantization reduces the bit-width of the numeric types used to store a model's weights and activations, replacing 32-bit floating-point (FP32) values with 8-bit integers (INT8) or lower, which cuts the memory a deployed model occupies and the compute each inference call requires. Google's LiteRT documentation defines the operation directly: quantization works by reducing the precision of the numbers used to represent a model's parameters, which by default are 32-bit floating point numbers, and the result is a smaller model size and faster computation (Google LiteRT, Model optimization). The compression ratio is fixed by the data type: dropping FP32 weights to INT8 cuts storage by four times before any other optimization. LiteRT lists quantization alongside pruning and clustering as the three optimization techniques it supports for shrinking edge models. The same primary documentation notes that activations can stay in floating point even when weights are quantized, which keeps numerical headroom on operations sensitive to dynamic range. The hardware payoff lands on integer execution units, which most modern AI accelerators from NVIDIA, Arm, and x86 CPU vendors expose, and the practical effect is shorter latency and lower power per inference. The ONNX runtime, vLLM, and the Hugging Face Optimum toolkit all consume quantized models produced by these flows. For context on how different silicon targets handle the resulting integer math, see how NPUs, GPUs, and TPUs handle AI inference workloads.
- FP32: the default 32-bit floating-point precision used during training, four bytes per parameter.
- INT8: 8-bit signed integer representation, one byte per parameter and the standard production target.
- Weights: the learned parameters that quantization most often compresses; the largest share of model size.
- Activations: the per-layer intermediate values; quantizing them too unlocks faster integer kernels.
- Scale and zero-point: the calibration constants that map a floating-point range onto a quantized integer range.
Post-Training Quantization: Fastest Path to a Smaller Model
Post-training quantization (PTQ) applies model quantization after training is complete, requiring no retraining, making it the default first step for most production deployments. Google's LiteRT documentation describes PTQ as a conversion technique that can reduce model size while also improving CPU and hardware accelerator latency, with little degradation in model accuracy (Google LiteRT, Post-training quantization). Within PTQ, several flavors exist. The simplest is dynamic range quantization, which quantizes weights to INT8 statically while activations are quantized on the fly during inference, with activations stored in floating point between layers. According to the same LiteRT documentation, dynamic range quantization achieves a 4x reduction in the model size, and the same page documents how LiteRT mixes floating-point and quantized kernels for different parts of the graph when a fully quantized kernel is not available. Full integer quantization goes further, converting both weights and activations to INT8 using a calibration dataset, and is the path that yields the largest latency gains on integer-only accelerators. Float16 quantization, by contrast, halves precision rather than going to integers, producing a 2x reduction in model size while keeping the GPU-friendly floating-point math intact (Google LiteRT, Float16 post-training quantization). On the GPU side, NVIDIA's TensorRT documentation states that TensorRT enables high-performance inference by supporting quantization, a technique that reduces model size and accelerates computation by representing floating-point values with lower-precision data types (NVIDIA TensorRT, Working with quantized types). For teams running PTQ as part of a managed pipeline, the same workflow plugs into how AWS SageMaker, Google Vertex AI, and Azure ML handle model deployment.
- Dynamic range quantization: weights quantized to INT8 statically, activations quantized on the fly, no calibration data required.
- Full integer quantization: both weights and activations quantized to INT8 using a representative calibration dataset.
- Float16 quantization: weights converted from FP32 to half-precision FP16, halving model size while keeping floating-point math.
- Weight-only quantization: only the weights are compressed (often to INT4), with activations kept at higher precision for accuracy.
Quantization-Aware Training: Preserving Accuracy at Lower Precision

Model quantization is not one format but a family of precision tiers, each presenting a different balance of memory savings, throughput gain, and accuracy risk. NVIDIA's TensorRT documentation lists its supported quantized data types as INT8 (signed 8-bit integer), INT4 (signed 4-bit integer, weight-only quantization), FP8E4M3 (FP8, 8-bit floating point with 4 exponent and 3 mantissa bits), and FP4E2M1 (FP4, 4-bit floating point with 2 exponent and 1 mantissa bit) (NVIDIA TensorRT, Working with quantized types). INT8 is the production workhorse: well-supported on every major accelerator, low accuracy risk for most networks, and a clean 4x reduction in weight storage from FP32. INT4 doubles the savings again but is typically weight-only, since activations at four bits rarely survive without QAT. Half-precision FP16 keeps a floating-point dynamic range, halves storage versus FP32, and is the default on GPUs for transformer inference when integer math is not acceptable. FP8 and FP4 are the newer floating-point integer hybrids that NVIDIA's recent Hopper and Blackwell generations target specifically for large language model inference. The GGUF format, popularized by the llama.cpp project, packages quantized large language model weights into a single file for CPU-first edge inference and supports a range of mixed-precision schemes (Q4_K_M, Q5_K_M, Q8_0, and others) that combine block-wise INT4 weights with selective higher-precision blocks. The bitsandbytes library, widely used in the Hugging Face ecosystem, provides INT8 and 4-bit inference paths for PyTorch models, including the NF4 (NormalFloat) data type used by QLoRA fine-tuning. Related kernels and runtimes such as ExLlamaV2, Marlin, and SmoothQuant target the same precision tiers, while ONNX serves as a portable interchange format and vLLM consumes the resulting weights for high-throughput serving. The picture for the rest of the serving stack is covered in how AI models are served at scale.
| Format | Bits per weight | Memory vs FP32 | Typical use | Primary tooling |
|---|---|---|---|---|
| half-precision FP16 | sixteen | 2x smaller | GPU transformer inference | TensorRT, LiteRT |
| INT8 | eight | 4x smaller | CPU, GPU, NPU production default | TensorRT, LiteRT, ExecuTorch, SmoothQuant |
| FP8 (E4M3) | eight | 4x smaller | LLM inference on Hopper, Blackwell | TensorRT |
| INT4 | four | 8x smaller | Weight-only LLM compression | TensorRT, GPTQ, AWQ, bitsandbytes, ExLlamaV2, Marlin |
| FP4 (E2M1) | four | 8x smaller | Aggressive LLM inference, QAT-aided | TensorRT |
| GGUF (mixed) | two to eight | 2x to 16x smaller | CPU edge LLM inference | llama.cpp |
Framework Support: TensorRT, LiteRT, ExecuTorch, and Hugging Face

Model quantization is supported across all major inference stacks, but the available data types, calibration workflows, and hardware targets differ significantly between frameworks. NVIDIA TensorRT targets NVIDIA GPUs and uses a symmetric quantization scheme, where both activations and weights are mapped to quantized values centered around zero, with three scale granularities documented as per-tensor quantization, per-channel quantization, and block quantization (NVIDIA TensorRT, Working with quantized types). Per-channel quantization is the standard for convolutional networks because it preserves the dynamic range of each output channel independently. Google LiteRT covers mobile and edge CPUs, GPUs, and NPUs, and ships dynamic range quantization, full integer quantization, and Float16 quantization as conversion-time options, with the dynamic-range flow documented as keeping activations stored in floating point. PyTorch ExecuTorch is the on-device runtime for PyTorch and frames its quantization page as advanced techniques for model compression and performance optimization (PyTorch ExecuTorch, Quantization and optimization). The ExecuTorch Arm VGF backend supports 8-bit symmetric weights with 8-bit asymmetric activations via the PT2E quantization flow, both static and dynamic activations, and per-channel and per-tensor schemes, with the caveat that weight-only quantization is not supported on the VGF backend per the same ExecuTorch documentation. The Hugging Face ecosystem layers on top of PyTorch: bitsandbytes provides INT8 and 4-bit kernels, the GPTQ algorithm produces post-training INT4 quantized weights with a reconstruction-loss objective, AWQ (Activation-aware Weight Quantization) selects salient weights to preserve at higher precision, and the Optimum library wraps these flows for export to ONNX and other runtimes. QLoRA combines NF4 quantization with low-rank adapters for memory-efficient fine-tuning on consumer GPUs. The hardware targets span NVIDIA Hopper and Blackwell GPUs, Qualcomm Snapdragon, Intel, and Apple silicon, while adjacent toolchains include TensorFlow, Keras, and OpenVINO, and calibration commonly draws on datasets such as WikiText and C4. On the kernel and runtime side, projects such as Triton, DeepSpeed, and Apple MLX accelerate quantized inference on architectures from Ampere onward, while Brevitas and Neural Magic target quantization-aware training and sparsity, and post-quantization quality is tracked against benchmarks like MMLU and HellaSwag. The training-side context for all of these workflows is covered in how machine learning training pipelines work.
- NVIDIA TensorRT: NVIDIA GPU target; INT8, INT4 (weight-only), FP8E4M3, FP4E2M1; PTQ and QAT; symmetric scheme; per-tensor, per-channel, block.
- Google LiteRT: mobile, CPU, GPU, NPU; dynamic range, full integer, Float16; PTQ-first; optimization alongside pruning and clustering.
- PyTorch ExecuTorch: on-device PyTorch runtime; PT2E and QAT; per-channel and per-tensor; static and dynamic activations; partial quantization on Arm VGF.
- Hugging Face stack: bitsandbytes for INT8 and 4-bit; GPTQ, AWQ, and SmoothQuant for post-training INT4 large language model weights; Optimum for ONNX export; QLoRA for NF4 fine-tuning.
- llama.cpp and GGUF: CPU-first inference, packaged single-file GGUF format, mixed-precision blocks (Q4_K_M and similar) for memory-constrained edge deployments; ExLlamaV2 and Marlin cover GPU-side INT4 kernels.
When to Use PTQ, QAT, or Weight-Only Quantization
Choosing the right model quantization path depends on your accuracy budget, whether retraining is feasible, and the hardware target. For a classification network or a moderately sized encoder shipping to a mobile CPU, dynamic range PTQ to INT8 is the default starting point: no calibration data needed, a 4x size reduction per the LiteRT documentation cited earlier, and usually negligible accuracy loss. When the deployment target is an integer-only NPU or DSP, full integer PTQ unlocks the largest latency gains because every kernel runs on integer hardware. For large language models on consumer GPUs, weight-only INT4 quantization via GPTQ or AWQ is the path that fits a 7B or 13B Llama, Mistral, or Qwen model into 8 to 16 GB of VRAM, with activations kept in half-precision FP16 to protect accuracy on attention layers; HBM-equipped data-center GPUs from NVIDIA tolerate the same scheme but get the largest gains from FP8 on Hopper and Blackwell. When the accuracy drop from PTQ exceeds the application's threshold (typically more than one to two percentage points on the relevant benchmark), QAT is the answer: it trades training-time compute for the smallest possible accuracy delta at INT8, INT4, FP8, or FP4. For CPU-only edge devices, the llama.cpp toolchain with a GGUF Q4_K_M or Q5_K_M file remains the dominant pattern, optimized for the integer SIMD paths on modern Arm and x86 cores; ExLlamaV2 and Marlin offer comparable INT4 throughput on NVIDIA GPUs. QLoRA-style NF4 quantization sits in a different niche: it is a fine-tuning pattern that quantizes the base model once and trains small low-rank adapters on top, keeping memory footprint low without committing to a single deployment precision. The decision is rarely irreversible. Most teams ship PTQ first, monitor accuracy in production, and graduate to QAT or a finer-grained scheme only when measurements demand it.
- Size-first, mobile CPU: PTQ with dynamic range quantization to INT8; no calibration data; ship the smaller binary.
- Latency-first, NPU or DSP: PTQ with full integer quantization to INT8; collect a calibration set of representative inputs.
- Accuracy gap too large: switch to QAT; fine-tune with fake-quantization nodes; target INT8, INT4, FP8, or FP4 as needed.
- Large language models on consumer GPUs: weight-only INT4 via GPTQ, AWQ, or SmoothQuant; keep activations at half-precision FP16 for attention numerics.
- CPU-only edge inference: convert to GGUF with llama.cpp; pick a mixed-precision tier (Q4_K_M, Q5_K_M) matched to the RAM ceiling.
- Memory-efficient fine-tuning: QLoRA with NF4 base weights and low-rank adapters trained on top.
Accuracy and Size Trade-offs: What Primary-Source Benchmarks Show
Model quantization benchmarks from vendor documentation give concrete anchors for the memory and accuracy trade-offs practitioners should expect. Google's LiteRT documentation states that dynamic range quantization achieves a 4x reduction in the model size, and that PTQ can reduce model size while also improving CPU and hardware accelerator latency, with little degradation in model accuracy (Google LiteRT, Post-training quantization). The same vendor source documents Float16 quantization as delivering a 2x reduction in model size (Google LiteRT, Float16 post-training quantization), a useful middle ground that preserves floating-point numerics for layers that need it. The accuracy story is more nuanced than the size story. LiteRT's primary documentation phrases the accuracy outcome carefully (little degradation), not as a guarantee, and the actual delta depends on the model architecture, the dataset, and which tensors are quantized. INT8 with per-channel quantization typically stays within a fraction of a percentage point on image classification networks. INT4 weight-only quantization on large language models, as practiced by GPTQ and AWQ, often costs one to three points on standardized perplexity benchmarks, depending on the model family and calibration data. FP8 and FP4 results vary more widely: NVIDIA documents the data types and the framework support, but does not blanket-recommend them for all models, and the editorial discipline is to validate against the target task. For broader help selecting tooling around these trade-offs, see how to evaluate and choose a machine learning platform for your workloads.
- FP32 to INT8 dynamic range: 4x reduction in model size per LiteRT, typically minimal accuracy degradation on classification and detection.
- FP32 to half-precision FP16: 2x reduction in model size per LiteRT, near-lossless on most architectures, GPU-friendly.
- INT4 weight-only on large language models: roughly 8x smaller weights versus FP32, modest perplexity cost when calibrated with GPTQ, AWQ, or SmoothQuant.
- QAT versus PTQ at the same precision: QAT typically narrows the accuracy gap, especially at INT4 and below.
- What to validate: measured task accuracy on a held-out set, end-to-end latency on the target hardware, and memory footprint at runtime, not just file size.
References
- Google LiteRT, Model optimization (quantization, pruning, clustering)
- Google LiteRT, Post-training quantization
- Google LiteRT, Float16 post-training quantization
- NVIDIA TensorRT, Working with quantized types
- PyTorch ExecuTorch, Quantization and optimization overview
- PyTorch ExecuTorch, Arm VGF backend quantization
Further reading
Frequently Asked Questions
What is model quantization in plain terms?
Model quantization converts a neural network's stored numbers from high-precision floats (FP32) to smaller integer or lower-precision types such as INT8 or INT4. The result is a model that occupies less RAM, loads faster, and runs inference more cheaply on CPUs, GPUs, and edge chips, with a typically small and measurable accuracy trade-off.
What is the difference between PTQ and QAT?
Post-training quantization (PTQ) compresses an already-trained model without any retraining, making it the fastest path to a smaller binary. Quantization-aware training (QAT) inserts fake-quantization nodes during the training loop so the model learns to tolerate lower precision before compression, preserving more accuracy at the cost of additional training time.
How much does INT8 quantization reduce model size?
Google's LiteRT documentation states that dynamic range quantization achieves a 4x reduction in model size when converting from FP32 to INT8 representations. The actual reduction for a given model depends on what fraction of its parameters are eligible for quantization and whether activations are also quantized.
Is model quantization safe to use in production?
Model quantization is production-standard across NVIDIA TensorRT, Google LiteRT, and PyTorch ExecuTorch. Google's LiteRT documentation notes that PTQ can reduce model size and improve latency with little degradation in accuracy. The practical risk depends on how aggressively you quantize: INT8 PTQ is low-risk for most classification and generation tasks; INT4 or FP4 requires validation against your accuracy requirements before shipping.









