← Explore
TOPIC

#cuda

Open source repositories tagged with #cuda, ranked by health score.

zhongkaifu
zhongkaifu/TensorSharp
C#
89
health

A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability

★ 557
timoncool
timoncool/YuE2-Studio
TypeScript
88
health

Local AI song generator with an editable score — YuE2 on your GPU: full songs with vocals, sheet music, covers, exact replay. Native Windows app, no Python, installer with auto-update.

★ 484
cupy
cupy/cupy
Python
88
health

NumPy & SciPy for GPU

★ 12.4k
LMCache
LMCache/LMCache
Python
88
health

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

★ 12.0k
NVIDIA
NVIDIA/cuda-rust
Rust
88
health

NVIDIA's CUDA platform for Rust. Host runtime crates plus Tile (cutile-rs) and SIMT (cuda-oxide) kernel programming models in idiomatic Rust.

★ 3.7k
invergent-ai
invergent-ai/surogate
C++
88
health

Train and serve LLMs at extreme speed and massive throughput.

★ 852
Code-Amadeus
Code-Amadeus/Amadeus
Python
87
health

Real-time multimodal desktop agent evolving toward a persistent AI OS interface (0.15 α).

★ 301
NVIDIA
NVIDIA/cudnn-frontend
Python
87
health

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

★ 961
pytorch
pytorch/TensorRT
Python
87
health

PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT

★ 3.0k
LuxCoreRender
LuxCoreRender/LuxCore
C++
87
health

LuxCore source repository

★ 1.3k
gpustack
gpustack/gpustack
Python
86
health

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

★ 5.8k
SemiAnalysisAI
SemiAnalysisAI/InferenceX
Python
85
health

Open Source Inference Research Platform Standard / 开源推理研究平台

★ 1.8k
headpiece747
headpiece747/ninfer-5090-windows
C++
84
health

Native Windows port of NInfer engine for RTX 5090. Features Qwen3.8-27B with QUASAR and NInfer models, MTP/DFlash2 with vision and 262,144 context

★ 55
ariannamethod
ariannamethod/notorch
C
82
health

neural networks in pure C

★ 27
kibae
kibae/onnxruntime-server
C++
69
health

ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.

★ 199
Wangmerlyn
Wangmerlyn/KeepGPU
Python
49
health

KeepGPU is a simple CLI app that keeps your GPUs running.

★ 40