Training
SqueezeLLM
Reduce model sizeQuantizationReferenceNon-uniform quantization that places quant levels by second-order sensitivity, with rare outliers kept sparse in FP16 — more bits, effectively, where they matter.
- Objective
- Reduce model size · Lossy
- Relieves
- CapacityBandwidth
- Targets
- Weights
- Format
- 3-bit non-uniform LUT + dense-and-sparse split
- Granularity
- Sensitivity-weighted, per-channel codebooks
- Lifecycle
- PTQ · no gradients
- Calibration
- Small calibration set
- Compression
- ~5× (3-bit)
- Quality
- Strong at 3-bit; sensitivity weighting beats uniform 3-bit
- Hardware
- LUT dequant + sparse-matrix kernels
- Runtimes
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.