← Explore
TOPIC

#inference

Open source repositories tagged with #inference, ranked by health score.

uxlfoundation
uxlfoundation/oneDNN
C++
91
health

Open-source library of optimized deep learning operations (matmul, convolution, attention) for CPUs (x64, AArch64, RISC-V) and Intel GPUs. Used by PyTorch, TensorFlow, OpenVINO, and ONNX Runtime.

★ 4.1k
openvinotoolkit
openvinotoolkit/openvino
C++
90
health

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

★ 11.0k
FlashML-org
FlashML-org/FreeVideo
Python
89
health

Make videos on the computer you already own. FreeVideo runs MiniMax H3 in as little as 8 GB of VRAM and 16 GB of RAM, and adapts its acceleration path to your hardware.

★ 923
timtoole02
timtoole02/Camelid
Rust
89
health

Camelid: a Rust-native local inference backend with evidence-gated model compatibility.

★ 215
SharpAI
SharpAI/SwiftLM
Swift
89
health

⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

★ 778
llm-d
llm-d/llm-d-router
Go
87
health

llm-d Router: The intelligent entry point for inference requests

★ 373
vllm-project
vllm-project/semantic-router
Go
87
health

A programmable Mixture-of-Models router for heterogeneous LLM inference

★ 6.0k
vllm-project
vllm-project/vllm-ascend
Python
85
health

Community maintained hardware plugin for vLLM on Huawei Ascend

★ 2.9k
rwilliamspbg-ops
rwilliamspbg-ops/Ghostlink
Rust
85
health

Distributed LLM inference fabric for heterogeneous local clusters. Features zero-copy SPSC ring buffers, automated tuning, and an OpenAI-compatible API.

★ 49
vllm-project
vllm-project/vllm-omni
Python
84
health

A framework for efficient model inference with omni-modality models

★ 7.1k
zoompilot
zoompilot/jetlink
Swift
80
health

Runs openpilot's large driving model on Jetson Orin Nano, macOS, iOS, Android, CUDA laptops with only your Comma, and a USB3 cable

★ 27
NVIDIA
NVIDIA/TensorRT-Model-Connect
Python
78
health

From PyTorch model to end-to-end TensorRT inference experience in two commands—AI-native, cross-platform, and built for the best possible user experience.

★ 271