Q8_0llama.cppthinking off
Stub recipe for the Q8_0 build of Laguna-XS 2.1 (poolside) on DGX Spark. Benchmark runs link here; serving budget, steps, and flags are still to be filled in.
Bench card
No measured runs yet — be the first: install the spark-benchmark skill.
Same quantized model
- Laguna-XS 2.1 Q4_K_M (llama.cpp)vanilla
- Laguna-XS 2.1 (poolside) — BF16 (vLLM)
- Laguna-XS 2.1 (poolside) — BF16 (llama.cpp)
Overview
_Stub recipe_ — auto-generated so its benchmark runs have a home to attach to. TODO: per-node memory budget, serving steps, key vLLM/llama.cpp flags, and troubleshooting.