Tag
#quantization
2 open-source projects filed under this tag.
01 Python
kvarn
KVarN is a vLLM KV-cache quantization backend from Huawei that delivers 3-5x more context capacity with FP16-level accuracy — one flag, no calibration.
#llm#quantization#vllm#inference
★ 366 18
Read
→
02 Python
turbovec
Turbovec is a Rust vector index with Python bindings built on Google's TurboQuant algorithm — faster than FAISS with zero training phase and 16x compression.
#vector-search#rust#python#rag
★ 8.7k 805
Read
→