howtospark
Training

KV-cache quantization

Reduce KV cache sizeKV compressionReference

Quantize the KV cache to fit longer context or a larger batch.

Objective
Reduce KV cache size
Relieves
KV cache
Targets
KV cache
Format
KV8 / KV4 (INT8 / INT4)
Granularity
Per-token / per-head
Lifecycle
PTQ · no gradients
Calibration
No calibration
Compression
2–4× KV
Quality
KV8 near-lossless; KV4 small loss
Hardware
Runtime KV-quant support
Runtimes
vLLMTensorRT-LLMSGLang
In the wild
vllm-Moet NVFP4 KV cache

Packed 352 B/token vs a 656 B baseline (+38% KV pool capacity) at 128K–512K windows.

A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.