OpenRadar

Tag

#inference

3 open-source projects filed under this tag.

01 Python

kvarn

KVarN is a vLLM KV-cache quantization backend from Huawei that delivers 3-5x more context capacity with FP16-level accuracy — one flag, no calibration.

#llm#quantization#vllm#inference
366 18 Read
02 Python

lmcache

LMCache is a KV cache management layer that slashes LLM inference latency by reusing computed attention states across requests, sessions, and GPU clusters.

#llm#inference#kv-cache#vllm
8.7k 1.3k Read
03 Python

tokenspeed

TokenSpeed is an open-source LLM inference engine built for agentic workloads — 580 tokens/sec on Qwen3.5-397B with TensorRT-LLM performance and vLLM-level ease of use.

#llm#inference#ai-agents#gpu
1.4k 142 Read