OpenRadar

Tag

#vllm

2 open-source projects filed under this tag.

01 Python

kvarn

KVarN is a vLLM KV-cache quantization backend from Huawei that delivers 3-5x more context capacity with FP16-level accuracy — one flag, no calibration.

#llm#quantization#vllm#inference
366 18 Read
02 Python

lmcache

LMCache is a KV cache management layer that slashes LLM inference latency by reusing computed attention states across requests, sessions, and GPU clusters.

#llm#inference#kv-cache#vllm
8.7k 1.3k Read