Training
SmoothQuant
Reduce model sizeQuantizationReferenceMigrates activation outliers into the weights to unlock W8A8.
- Objective
- Reduce model size · Lossy
- Relieves
- ComputeBandwidth
- Targets
- Weights · Activations
- Format
- W8A8 (INT8)
- Granularity
- Per-channel scaling
- Lifecycle
- PTQ · no gradients
- Calibration
- Small calibration set
- Compression
- ~2×
- Quality
- Near-FP16 at W8A8
- Hardware
- INT8 tensor cores
- Runtimes
- vLLMTensorRT-LLM
In the wild
AngelSlim FP8/INT8 quantization (Tencent)
Static/dynamic FP8 and dynamic INT8 W8A8 recipes with published accuracy tables — the low-drama 2× shrink tier.
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.