← Explore
TOPIC

#moe

Open source repositories tagged with #moe, ranked by health score.

gittensor-ai-lab
gittensor-ai-lab/sparkinfer
C++
94
health

Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.

★ 89
SharpAI
SharpAI/SwiftLM
Swift
89
health

⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

★ 778
carloslfu
carloslfu/slotstream
Swift
88
health

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.

★ 420
NVIDIA
NVIDIA/cudnn-frontend
Python
87
health

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

★ 959
townsendmerino
townsendmerino/goinfer
Go
81
health

Pure-Go, no-cgo local LLM inference — run Gemma, Qwen, Llama and friends from safetensors or GGUF in a single static binary.

★ 18