Open source repositories tagged with #inference-server, ranked by health score.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Runs openpilot's large driving model on Jetson Orin Nano, macOS, iOS, Android, CUDA laptops with only your Comma, and a USB3 cable
ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.