Model Deployment65 / 390 projects
Model deployment leaderboard featuring the hottest AI model serving and deployment projects on GitHub for production inference.
A high-throughput and memory-efficient inference and serving engine for LLMs
An open-source, self-hostable PaaS alternative to Vercel, Heroku & Netlify that lets you easily deploy static sites, databases, full-stack applications and 280+ one-click services on your own servers.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Open Source Alternative to Vercel, Netlify and Heroku.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Chef Infra, a powerful automation platform that transforms infrastructure into code automating how infrastructure is configured, deployed and managed across any environment, at any scale
A framework for efficient model inference with omni-modality models
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models
In this repository, I will share some useful notes and references about deploying deep learning-based models in production.
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
Build data pipelines, the easy way 🛠️
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
OpenMMLab Model Deployment Framework
Open-source inference server and production cluster for all the models your agent needs.
Community maintained hardware plugin for vLLM on Ascend
Turn any computer or edge device into a command center for your computer vision projects.
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
AICI: Prompts as (Wasm) Programs
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
Machine Learning Platform and Recommendation Engine built on Kubernetes
An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.
Deploy a ML inference service on a budget in less than 10 lines of code.
Declarative Intent Driven Platform Orchestrator for Internal Developer Platform (IDP).
Hopsworks - Data-Intensive AI platform with a Feature Store
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
The simplest way to serve AI/ML models in production
A throughput-oriented high-performance serving framework for LLMs
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
A highly optimized LLM inference acceleration engine for Llama and its variants.
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
Deployment platform for code-first internal tools. Deploy web apps declaratively, on a single-node or on Kubernetes, with OIDC/SAML auth and RBAC.
Train and Deploy an ML REST API to predict crypto prices, in 10 steps
An open-source computer vision framework to build and deploy apps in minutes
Model Deployment at Scale on Kubernetes 🦄️
A self-hosted, open-source Platform as a Service that enables easy swarm deployments, load balancing, automatic SSL, metrics, analytics and more.
🐶 A tool to package, serve, and deploy any ML model on any platform. Archived to be resurrected one day🤞
KubeStellar - a flexible solution for multi-cluster configuration management for edge, multi-cloud, and hybrid cloud
Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
Manage federated learning workload using cloud native technologies.
FedScale is a scalable and extensible open-source federated learning (FL) platform.
Deploy and scale serverless machine learning app - in 4 steps.
BentoDiffusion: A collection of diffusion models served with BentoML
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.
A multi-functional library for full-stack Deep Learning. Simplifies Model Building, API development, and Model Deployment.
A scalable, high-performance serving system for federated learning models
MONAI Deploy App SDK offers a framework and associated tools to design, develop and verify AI-driven applications in the healthcare imaging domain.
Deploy DL/ ML inference pipelines with minimal extra code.
Serving PyTorch models with TorchServe :fire:
:scroll: Deploy Home-Assistant easily with Fabric
Route inference across providers.
Pluto provides a unified programming interface that allows you to seamlessly tap into cloud capabilities and develop your cloud and AI applications.
A wannabe Ollama equivalent for Apple MlX models
Hypergol is a Data Science/Machine Learning productivity toolkit to accelerate any projects into production with autogenerated code, standardised structure for data and ML and parallel processing out-of-the-box.
🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffusion.cpp
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.
Source: GitHub API · Curated · Realtime