Training
GPTQ
Reduce model sizeQuantizationReferenceHessian-based layerwise error compensation; the classic 4-bit weight quant.
- Objective
- Reduce model size · Lossy
- Relieves
- CapacityBandwidth
- Targets
- Weights
- Format
- INT4 / INT3 (W4A16)
- Granularity
- Per-group
- Lifecycle
- PTQ · no gradients
- Calibration
- Small calibration set
- Compression
- ~4×
- Quality
- Near-FP16 at 4-bit
- Hardware
- INT4 dequant kernels; broad support
- Runtimes
- vLLMTensorRT-LLMSGLangTransformers
In the wild
AngelSlim INT4 pipeline (Tencent)
GPTQ and GPTAQ (their accuracy-recovering variant) at W4A16 across Qwen3 / DeepSeek / Hunyuan, benchmarked on CEVAL/MMLU/GSM8K with minimal loss.
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.