howtospark
Training

GPTQ

Reduce model sizeQuantizationReference

Hessian-based layerwise error compensation; the classic 4-bit weight quant.

Objective
Reduce model size · Lossy
Relieves
CapacityBandwidth
Targets
Weights
Format
INT4 / INT3 (W4A16)
Granularity
Per-group
Lifecycle
PTQ · no gradients
Calibration
Small calibration set
Compression
~4×
Quality
Near-FP16 at 4-bit
Hardware
INT4 dequant kernels; broad support
Runtimes
vLLMTensorRT-LLMSGLangTransformers
In the wild
AngelSlim INT4 pipeline (Tencent)

GPTQ and GPTAQ (their accuracy-recovering variant) at W4A16 across Qwen3 / DeepSeek / Hunyuan, benchmarked on CEVAL/MMLU/GSM8K with minimal loss.

A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.