Training
SpQR
Reduce model sizeQuantizationReferenceVariable bitrate for weights: isolate the sensitive outliers in high precision and push everything else to ~3 bits.
On the Spark
The clearest 'VBR' quant — spends bits where the model is sensitive, matching how gguf K-quants mix per-tensor bit-widths.
- Objective
- Reduce model size · Lossy
- Relieves
- CapacityBandwidth
- Targets
- Weights
- Format
- ~3-bit + high-precision outliers (sparse)
- Granularity
- Per-group, sensitivity-driven mixed precision
- Lifecycle
- PTQ · no gradients
- Calibration
- Small calibration set
- Compression
- ~5× (near-4-bit average)
- Quality
- Near-lossless vs FP16 at ~3.4 bits/weight
- Hardware
- Sparse-outlier + low-bit dequant kernels
- Runtimes
- Transformers
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.