← Explore
TOPIC

#qwen

Open source repositories tagged with #qwen, ranked by health score.

lightseekorg
lightseekorg/tokenspeed
Python
89
health

TokenSpeed is a speed-of-light LLM inference engine.

2.0k
syv-ai
syv-ai/qwen38-27b-rtx3090
Python
89
health

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

462
xorbitsai
xorbitsai/inference
Python
89
health

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

9.5k
avifenesh
avifenesh/memra
Rust
89
health

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

321
zhongkaifu
zhongkaifu/TensorSharp
C#
89
health

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

376
happier-dev
happier-dev/happier
TypeScript
88
health

Web, Desktop & Mobile client for Codex, Claude Code, OpenCode, Kimi, Augment Code, Qwen, fully end-to-end encrypted

1.5k
QwenLM
QwenLM/qwen-code
TypeScript
88
health

An open-source AI coding agent that lives in your terminal.

27.3k
invergent-ai
invergent-ai/surogate
C++
88
health

Training/Fine-tuning at the speed of light

812
PowerBeef
PowerBeef/Vocello
Swift
88
health

Vocello: a local, private voice studio for Apple Silicon. Write a script, pick or describe a voice, and generate speech on-device, faster than realtime on an 8 GB M2 Mac mini. Native Swift + MLX, no Python. Mac app out now, iPhone beta on TestFlight. (Formerly QwenVoice.)

354
Indras-Mirror
Indras-Mirror/llama.cpp-turboq-mtp
C++
87
health

Fused TBQ4 Flash Attention + MTP + Shared Tensors + Qwen35 SWA Hybrid for llama.cpp — 82+ tok/s, lossless 4.25 bpv KV cache, SWA-bounded deep-context decode (w/ long-range recall) on RTX 4090

90
lemonade-sdk
lemonade-sdk/lemonade
C++
87
health

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

5.4k