Training
KV-cache quantization
Reduce KV cache sizeKV compressionReferenceQuantize the KV cache to fit longer context or a larger batch.
- Objective
- Reduce KV cache size
- Relieves
- KV cache
- Targets
- KV cache
- Format
- KV8 / KV4 (INT8 / INT4)
- Granularity
- Per-token / per-head
- Lifecycle
- PTQ · no gradients
- Calibration
- No calibration
- Compression
- 2–4× KV
- Quality
- KV8 near-lossless; KV4 small loss
- Hardware
- Runtime KV-quant support
- Runtimes
- vLLMTensorRT-LLMSGLang
In the wild
vllm-Moet NVFP4 KV cache
Packed 352 B/token vs a 656 B baseline (+38% KV pool capacity) at 128K–512K windows.
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.