← Explore
TOPIC

#vllm

Open source repositories tagged with #vllm, ranked by health score.

nodetool-ai
nodetool-ai/nodetool
TypeScript
89
health

The open-source, agent-first creative workspace.

490
syv-ai
syv-ai/qwen38-27b-rtx3090
Python
89
health

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

462
xorbitsai
xorbitsai/inference
Python
89
health

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

9.5k
mostlygeek
mostlygeek/llama-swap
Go
89
health

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

5.4k
vllm-project
vllm-project/semantic-router
Go
88
health

A programmable Mixture-of-Models router for heterogeneous LLM inference

5.2k
kvcache-ai
kvcache-ai/Mooncake
C++
88
health

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

6.3k
mudler
mudler/vllm.cpp
C++
77
health

a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)

321