howtospark
Training

SpQR

Reduce model sizeQuantizationReference

Variable bitrate for weights: isolate the sensitive outliers in high precision and push everything else to ~3 bits.

On the Spark

The clearest 'VBR' quant — spends bits where the model is sensitive, matching how gguf K-quants mix per-tensor bit-widths.

Objective
Reduce model size · Lossy
Relieves
CapacityBandwidth
Targets
Weights
Format
~3-bit + high-precision outliers (sparse)
Granularity
Per-group, sensitivity-driven mixed precision
Lifecycle
PTQ · no gradients
Calibration
Small calibration set
Compression
~5× (near-4-bit average)
Quality
Near-lossless vs FP16 at ~3.4 bits/weight
Hardware
Sparse-outlier + low-bit dequant kernels
Runtimes
Transformers

A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.