← Explore
TOPIC

#llm-inference

Open source repositories tagged with #llm-inference, ranked by health score.

felladrin
felladrin/MiniSearch
TypeScript
89
health

Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space

581
syv-ai
syv-ai/qwen38-27b-rtx3090
Python
89
health

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

462
avifenesh
avifenesh/memra
Rust
89
health

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

321
openvinotoolkit
openvinotoolkit/openvino
C++
89
health

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

10.7k
drumih
drumih/turbo-fieldfare
Swift
88
health

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

6.3k
flashinfer-ai
flashinfer-ai/flashinfer
Python
87
health

FlashInfer: Kernel Library for LLM Serving

6.2k
lemonade-sdk
lemonade-sdk/lemonade
C++
87
health

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

5.4k
spiceai
spiceai/spiceai
Rust
83
health

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

3.1k
Avarok-Cybersecurity
Avarok-Cybersecurity/atlas
Rust
82
health

Pure Rust Inference Engine

658
anthony-chaudhary
anthony-chaudhary/fak
Go
73
health

Create your Agentic AIs.

31