howtospark
Training

AQLM / QuIP# (2-bit codebook)

Reduce model sizeQuantizationReference

Codebook / vector quantization pushing weights toward ~2 bits.

Objective
Reduce model size · Lossy
Relieves
Capacity
Targets
Weights
Format
2-bit vector / codebook
Granularity
Vector quantization
Lifecycle
PTQ · no gradients
Calibration
Full corpus
Compression
~8× (2-bit)
Quality
Usable at 2-bit; below 4-bit methods
Hardware
Custom codebook decode kernels
Runtimes
vLLMTransformers
In the wild
vllm-Moet 2-bit expert codebook

Routed experts of GLM-5.2 (753B) / Kimi-K2.7 (1T) at 2-bit via a sign-symmetric {-4,-1,1,4} codebook and hand-written Blackwell SASS kernels; dense stack kept at FP8/NVFP4.

AngelSlim sub-2-bit research quant (Tencent)

Tequila (ternary) and Sherry / STQ1_0 (1.25-bit) — the frontier below 2 bits, from the same toolkit that ships the boring INT4/FP8 paths.

A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.