Tag
#vllm
2 open-source projects filed under this tag.
01 Python
kvarn
KVarN is a vLLM KV-cache quantization backend from Huawei that delivers 3-5x more context capacity with FP16-level accuracy — one flag, no calibration.
#llm#quantization#vllm#inference
★ 366 18
Read
→
02 Python
lmcache
LMCache is a KV cache management layer that slashes LLM inference latency by reusing computed attention states across requests, sessions, and GPU clusters.
#llm#inference#kv-cache#vllm
★ 8.7k 1.3k
Read
→