GPU & Inference Optimization226 / 442 projects
Inference optimization leaderboard featuring the hottest GPU acceleration and model optimization projects on GitHub.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
A high-throughput and memory-efficient inference and serving engine for LLMs
C/C++ implementation of LLM inference. Runs on CPU, GPU, and edge devices. Supports GGUF quantized models.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Port of OpenAI's Whisper model in C/C++
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Making large AI models cheaper, faster and more accessible
SGLang is a high-performance serving framework for large language models and multimodal models.
The fastai deep learning library
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
Nano vLLM
A vector index built on TurboQuant, written in Rust with Python bindings
Open3D: A Modern Library for 3D Data Processing
Open Machine Learning Compiler Framework
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
Large Language Model Text Generation Inference
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Production ready toolkit to run AI locally
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
cuDF - GPU DataFrame Library
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT and NVIDIA Jetson.
Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrete GPU such as Arc, Flex and Max); seamlessly integrate with llama.cpp, Ollama, HuggingFace, LangChain, LlamaIndex, vLLM, DeepSpeed, Axolotl, etc.
Accessible large language models via k-bit quantization for PyTorch.
A Datacenter Scale Distributed Inference Serving Framework
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.
A Python framework for GPU-accelerated simulation, robotics, and machine learning.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
FlashInfer: Kernel Library for LLM Serving
A framework for efficient model inference with omni-modality models
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
High-performance TensorFlow library for quantitative finance.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Superduper: End-to-end framework for building custom AI applications and agents.
cuML - RAPIDS Machine Learning Library
A programmable Mixture-of-Models router for heterogeneous LLM inference
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持国产cpu/gpu/npu 昇腾生态,支持RDMA,支持pytorch/tf/mxnet/deepspeed/paddle/colossalai/horovod/ray/volcano等分布式
An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.
Time series forecasting with PyTorch
Optimized primitives for collective multi-GPU communication
An easy to use PyTorch to TensorRT converter
On-device AI across mobile, embedded and edge for PyTorch
MegEngine 是一个快速、可拓展、易于使用且支持自动求导的深度学习框架
TNN: developed by Tencent Youtu Lab and Guangying Lab, a uniform deep learning inference framework for mobile、desktop and server. TNN is distinguished by several outstanding features, including its cross-platform capability, high performance, model compression and code pruning. Based on ncnn and Rapidnet, TNN further strengthens the support and performance optimization for mobile devices, and also draws on the advantages of good extensibility and high performance from existed open source efforts. TNN has been deployed in multiple Apps from Tencent, such as Mobile QQ, Weishi, Pitu, etc. Contributions are welcome to work in collaborative with us and make TNN a better framework.
Fast inference engine for Transformer models
Lightning fast C++/CUDA neural network framework
Serve, optimize and scale PyTorch models in production
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
Deep Learning GPU Training System
[ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops 🦖
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Fast and flexible AutoML with learning guarantees.
LightSeq: A High Performance Library for Sequence Processing and Generation
Sparsity-aware deep learning inference runtime for CPUs
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
Build, personalize and control your own LLMs. From data pre-processing to fine-tuning, xTuring provides an easy way to personalize open-source LLMs. Join our discord community: https://discord.gg/TgHXuSJEk6
Open-source inference server and production cluster for all the models your agent needs.
Pytorch domain library for recommendation systems
Community maintained hardware plugin for vLLM on Ascend
Deep Learning Server and CLI for Torch and TensorRT
Fast ML inference & training for ONNX models in Rust
Turn any computer or edge device into a command center for your computer vision projects.
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
micronet, a model compression and deploy lib. compression: 1、quantization: quantization-aware-training(QAT), High-Bit(>2b)(DoReFa/Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference)、Low-Bit(≤2b)/Ternary and Binary(TWN/BNN/XNOR-Net); post-training-quantization(PTQ), 8-bit(tensorrt); 2、 pruning: normal、regular and group convolutional channel pruning; 3、 group convolution structure; 4、batch-normalization fuse for quantization. deploy: tensorrt, fp32/fp16/int8(ptq-calibration)、op-adapt(upsample)、dynamic_shape
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
AICI: Prompts as (Wasm) Programs
Ultrafast serverless GPU inference, sandboxes, and background jobs
🔥 Real-time NVIDIA GPU dashboard
A toolkit to optimize ML models for deployment for Keras and TensorFlow, including quantization and pruning.
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
Deploy a ML inference service on a budget in less than 10 lines of code.
Efficient computing methods developed by Huawei Noah's Ark Lab
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
GraphVite: A General and High-performance Graph Embedding System
Examples of programs built using Modal
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Neural Network Compression Framework for enhanced OpenVINO™ inference
NVTabular is a feature engineering and preprocessing library for tabular data designed to quickly and easily manipulate terabyte scale datasets used to train deep learning based recommender systems.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form building blocks for more easily writing high performance applications.
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
A tool for converting ONNX files to LiteRT/TFLite/TensorFlow, PyTorch native code (nn.Module), TorchScript (.pt), state_dict (.pt), Exported Program (.pt2), and Dynamo ONNX. It also supports direct conversion from LiteRT to PyTorch.
A throughput-oriented high-performance serving framework for LLMs
TinyChatEngine: On-Device LLM Inference Library
Bolt is a deep learning library with high performance and heterogeneous flexibility.
Official implementation of Half-Quadratic Quantization (HQQ)
[NeurIPS 2020] MCUNet: Tiny Deep Learning on IoT Devices; [NeurIPS 2021] MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning; [NeurIPS 2022] MCUNetV3: On-Device Training Under 256KB Memory
A uniform interface to run deep learning models from multiple frameworks
A flexible, high-performance carrier for machine learning models(『飞桨』服务化部署框架)
[ICLR2024 spotlight] OmniQuant is a simple and powerful quantization technique for LLMs.
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
Performance-optimized wheels for TensorFlow (SSE, AVX, FMA, XLA, MPI)
An open-source computer vision framework to build and deploy apps in minutes
cuVS - a library for vector search and clustering on the GPU
End-to-end documentation to set up your own local & fully private LLM server on Debian. Equipped with chat, web search, RAG, model management, MCP servers, image generation, and TTS.
High-performance Vision library in Python. Scale your research, not boilerplate.
Run TensorFlow models in C++ without installation and without Bazel
GPU worker client for the Talos network. Pairs with your Talos account, serves open-model inference jobs over a WebSocket, and reports uptime for payouts.
OpenCL 1.2 implementation for Tensorflow
Estimate whether a Hugging Face model fits and fine-tunes on your local GPU.
Multimodal RL training framework for diffusion & omni models
The open source, end-to-end computer vision platform. Label, build, train, tune, deploy and automate in a unified platform that runs on any cloud and on-premises.
[ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization
TensorFlow models accelerated with NVIDIA TensorRT
Graph Data Science: an abstraction layer in Python for building knowledge graphs, integrated with popular graph libraries – atop Pandas, NetworkX, RAPIDS, RDFlib, pySHACL, PyVis, morph-kgc, pslpython, pyarrow, etc.
Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.
Library for faster pinned CPU <-> GPU transfer in Pytorch
GPU environment and cluster management with LLM support
A Pytorch Knowledge Distillation library for benchmarking and extending works in the domains of Knowledge Distillation, Pruning, and Quantization.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
convert mmdetection model to tensorrt, support fp16, int8, batch input, dynamic shape etc.
Minimum-distortion embedding with PyTorch
⚡ boost inference speed of T5 models by 5x & reduce the model size by 3x.
Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)
Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. Join our Discord communty: https://discord.com/invite/TgHXuSJEk6
Static suckless single batch CUDA-only qwen3-0.6B mini inference engine
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
Examples for Cerebrium Serverless GPUs
A high-performance inference system for large language models, designed for production environments.
RDNA-native LLM inference engine in Rust.
RDNA-native LLM inference engine in Rust.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
⚠️ Legacy repository for Geti v2.x. For Geti v3.0+, visit https://github.com/open-edge-platform/geti
Efficient AI Inference & Serving
Quantization library for PyTorch. Support low-precision and mixed-precision quantization, with hardware implementation through TVM.
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
RAG (Retrieval-augmented generation) ChatBot that provides answers based on contextual information extracted from a collection of Markdown files.
[NeurIPS 2024] KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models(NeurIPS 2024 Spotlight)
T-GATE: Temporally Gating Attention to Accelerate Diffusion Model for Free!
An open-source tool for sequence learning in NLP built on TensorFlow.
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.
Super performant RAG pipelines for AI apps. Summarization, Retrieve/Rerank and Code Interpreters in one simple API.
[ICCV 2023] Q-Diffusion: Quantizing Diffusion Models.
Vector search engine inside Milvus, integrating FAISS, HNSW, DiskANN.
An innovative library for efficient LLM inference via low-bit quantization
A unified, high-performance framework for training LLMs, VLMs, diffusion, and embodied models on NVIDIA GPUs and Kunlun XPUs.
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
On-device LLM Inference Powered by X-Bit Quantization
ClearML Agent - MLOps/LLMOps made easy. MLOps/LLMOps scheduler & orchestration solution
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.
☁️ Terraform plugin for machine learning workloads: spot instance recovery & auto-termination | AWS, GCP, Azure, Kubernetes
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
bw24 — from-scratch LLM inference for RTX 5090 (sm_120a) and H100 (sm_90a)
Autoscale LLM (vLLM, SGLang, LMDeploy) inferences on Kubernetes (and others)
[COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
HDLTex: Hierarchical Deep Learning for Text Classification
[ICML'21 Oral] I-BERT: Integer-only BERT Quantization
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.
FasterAI: Prune and Distill your models with FastAI and PyTorch
针对pytorch模型的自动化模型结构分析和修改工具集,包含自动分析模型结构的模型压缩算法库
OramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more utilities.
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。
An official lightweight library for the RaBitQ algorithm and its applications in vector search.
One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
Enhancing LLMs with LoRA
This project collects GPU benchmarks from various cloud providers and compares them to fixed per token costs. Use our tool for efficient LLM GPU selections and cost-effective AI models. LLM provider price comparison, gpu benchmarks to price per token calculation, gpu benchmark table
HierarchicalKV is a part of NVIDIA Merlin and provides hierarchical key-value storage to meet RecSys requirements. The key capability of HierarchicalKV is to store key-value feature-embeddings on high-bandwidth memory (HBM) of GPUs and in host memory. It also can be used as a generic key-value storage.
RapidFire AI: Rapid AI Customization from RAG to Fine-Tuning
An Open Source Deep Learning Inference Engine Based on FPGA
A scalable, high-performance serving system for federated learning models
A Deep learning library for neutrino telescopes
Model compression for ONNX
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
Lightweight turnkey solution for AI
AI inference, packed simply. A blazing-fast, zero-dependency WebGPU runtime to run GGUF models directly in the browser. Features a symmetric API for seamless local execution and cloud provider routing. Built with Rust & C++.
A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
Developer Hub for NVIDIA Alpamayo, containing ready-to-use recipes for fine-tuning, reinforcement-learning post-training, quantization, and deployment.
Nano vLLM with vLLM v1's request scheduling strategy and chunked prefill
✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers
Source: GitHub API · Curated · Realtime