Training
AQLM / QuIP# (2-bit codebook)
Reduce model sizeQuantizationReferenceCodebook / vector quantization pushing weights toward ~2 bits.
- Objective
- Reduce model size · Lossy
- Relieves
- Capacity
- Targets
- Weights
- Format
- 2-bit vector / codebook
- Granularity
- Vector quantization
- Lifecycle
- PTQ · no gradients
- Calibration
- Full corpus
- Compression
- ~8× (2-bit)
- Quality
- Usable at 2-bit; below 4-bit methods
- Hardware
- Custom codebook decode kernels
- Runtimes
- vLLMTransformers
In the wild
vllm-Moet 2-bit expert codebook
Routed experts of GLM-5.2 (753B) / Kimi-K2.7 (1T) at 2-bit via a sign-symmetric {-4,-1,1,4} codebook and hand-written Blackwell SASS kernels; dense stack kept at FP8/NVFP4.
AngelSlim sub-2-bit research quant (Tencent)
Tequila (ternary) and Sherry / STQ1_0 (1.25-bit) — the frontier below 2 bits, from the same toolkit that ships the boring INT4/FP8 paths.
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.