ML Libraries & Tools891 / 620 projects
Machine learning & deep learning libraries, NLP libraries, RL libraries, AutoML tools, and general ML frameworks
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
The fastest path to AI-powered full stack observability, even for lean teams.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
scikit-learn: machine learning in Python
Deep Learning for humans
Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.
Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, including supervised learning, market dynamics modeling, and RL, and is now equipped with https://github.com/microsoft/RD-Agent to automate R&D process.
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Caffe: a fast open framework for deep learning.
Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the cloud https://cloud.qdrant.io/
SGLang is a high-performance serving framework for large language models and multimodal models.
Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C and more. Runs on single machine, Hadoop, Spark, Dask, Flink and DataFlow
The fastai deep learning library
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
The Modular Platform (includes MAX & Mojo)
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
☁️ Build multimodal AI applications with cloud-native stack
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Open standard for machine learning interoperability
A WebGL accelerated JavaScript library for training and deploying ML models.
A fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
Private AI platform for agents, assistants and enterprise search. Built-in Agent Builder, Deep research, Document analysis, Multi-model support, and API connectivity for agents.
Datasets, Transforms and Models specific to Computer Vision
Microsoft Cognitive Toolkit (CNTK), an open source deep-learning toolkit
Library of deep learning models and datasets designed to make deep learning more accessible and accelerate ML research.
Topic Modelling for Humans
Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.
Machine Learning Toolkit for Kubernetes
🦉 Data Versioning and ML Experiments
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
Fast and flexible image augmentation library. Paper about the library: https://www.mdpi.com/2078-2489/11/2/125
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Tensor library for machine learning
Nano vLLM
Image augmentation for machine learning experiments.
Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
NLTK Source
Bullet Physics SDK: real-time collision detection and multi-physics simulation for VR, games, visual effects, robotics, machine learning etc.
ChatterBot is a machine learning, conversational dialog engine for creating chat bots
A toolkit for making real world machine learning and data analysis applications in C++
A very simple framework for state-of-the-art Natural Language Processing (NLP)
An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.
Open3D: A Modern Library for 3D Data Processing
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Open Machine Learning Compiler Framework
💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
Turi Create simplifies the development of custom machine learning models.
Refine high-quality datasets and visual AI models
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
Large Language Model Text Generation Inference
Open source annotation tool for machine learning practitioners.
Fast and Accurate ML in 3 Lines of Code
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
Content aware image resize library
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Build, Manage and Deploy AI/ML Systems
A Python Automated Machine Learning tool that optimizes machine learning pipelines using genetic programming.
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.
TensorFlow-based neural network library
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane.
AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding
Deep learning library featuring a higher-level API for TensorFlow.
Simple, Pythonic, text processing--Sentiment analysis, part-of-speech tagging, noun phrase extraction, translation, and more.
A python library for user-friendly forecasting and anomaly detection on time series.
Containers for machine learning
Running large language models on a single GPU for throughput-oriented scenarios.
ML.NET is an open source and cross-platform machine learning framework for .NET.
AutoML library for deep learning
ModelScope: bring the notion of Model-as-a-Service to life.
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT and NVIDIA Jetson.
Web mining module for Python, with tools for scraping, natural language processing, machine learning, network analysis and visualization.
🧙 Build, run, and manage data pipelines for integrating and transforming data.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
The backtesting engine that gives you an unfair advantage. Run thousands of trading ideas before others finish one.
Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per second 🚀
Accessible large language models via k-bit quantization for PyTorch.
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. DoWhy is based on a unified language for causal inference, combining causal graphical models and potential outcomes frameworks.
Production infrastructure for machine learning at scale
An open source python library for automated feature engineering
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.
Deep Learning and Reinforcement Learning Library for Scientists and Engineers
Visualizations for machine learning datasets
The Open Source Feature Store for AI/ML
🔮 Libraries & tools for enabling Machine Learning driven user-experiences on the web
A Python Package to Tackle the Curse of Imbalanced Datasets in Machine Learning
Flower: A Friendly Federated AI Framework
The AI search platform
Officially maintained, supported by PaddlePaddle, including CV, NLP, Speech, Rec, TS, big models and so on.
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..
Postgres with GPUs for ML/AI apps.
ClearML - Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution
A Python scikit for building and analyzing recommender systems
A Flexible and Powerful Parameter Server for large-scale machine learning
Friendly machine learning for the web! 🤖
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
A flexible, high-performance serving system for machine learning models
The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.
A Neural Net Training Interface on TensorFlow, with focus on speed + flexibility
Puffing up reinforcement learning
Aim 💫 — An easy-to-use & supercharged open-source experiment tracker.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Swift for TensorFlow
Time series Timeseries Deep Learning Machine Learning Python Pytorch fastai | State-of-the-art Deep Learning library for Time Series and Sequences in Pytorch / fastai
The Swift machine learning library.
A system for quickly generating training data with weak supervision
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
Extract Keywords from sentence or Replace keywords in sentences.
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
mlpack: a fast, header-only C++ machine learning library
Deep Reinforcement Learning for Keras.
A C++ standalone library for machine learning
OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.
Core ML tools contain supporting tools for Core ML model conversion, editing, and validation.
cuML - RAPIDS Machine Learning Library
Simple and Distributed Machine Learning
Probabilistic time series modeling in Python
Brings SQL and AI together.
A Next-Generation Training Engine Built for Ultra-Large MoE Models
A library of extension and helper modules for Python's data analysis and machine learning libraries.
Image augmentation library in Python for machine learning.
A Python implementation of LightFM, a hybrid recommendation algorithm.
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.
A machine learning software for extracting information from scholarly documents
Time series forecasting with PyTorch
Run Keras models in the browser, with GPU support using WebGL
MlFinLab helps portfolio managers and traders who want to leverage the power of machine learning by providing reproducible, interpretable, and easy to use tools.
On-device AI across mobile, embedded and edge for PyTorch
A C library for parsing/normalizing street addresses around the world. Powered by statistical NLP and open geo data.
Lightning ⚡️ fast forecasting with statistical and econometric models.
A probabilistic programming language in TensorFlow. Deep generative models, variational inference.
An Engine-Agnostic Deep Learning Framework in Java
Machine Learning Containers for NVIDIA Jetson and JetPack-L4T
MegEngine 是一个快速、可拓展、易于使用且支持自动求导的深度学习框架
High-level library to help with training and evaluating neural networks in PyTorch flexibly and transparently.
An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models
ALICE (Automated Learning and Intelligence for Causation and Economics) is a Microsoft Research project aimed at applying Artificial Intelligence concepts to economic decision making. One of its goals is to build a toolkit that combines state-of-the-art machine learning techniques with econometrics in order to bring automation to complex causal inference problems. To date, the ALICE Python SDK (econml) implements orthogonal machine learning algorithms such as the double machine learning work of Chernozhukov et al. This toolkit is designed to measure the causal effect of some treatment variable(s) t on an outcome variable y, controlling for a set of features x.
Relax! Flux is the ML library that doesn't make you tensor
A Rust machine learning framework.
Data augmentation for NLP
🤗 AutoTrain Advanced
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...
Tengine is a lite, high performance, modular inference engine for embedded device
Merlion: A Machine Learning Framework for Time Series Intelligence
RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
搜索所有中文NLP数据集,附常用英文NLP数据集
[EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction
RAG (Retrieval Augmented Generation) Framework for building modular, open source applications for production by TrueFoundry
A fast library for AutoML and tuning. Join our Discord: https://discord.gg/Cppx2vSPVP.
A JavaScript deep learning and reinforcement learning library.
Massively Parallel Deep Reinforcement Learning. 🔥
Serve, optimize and scale PyTorch models in production
NeuralProphet: A simple forecasting package
modular quant framework.
MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。
Scalable and user friendly neural :brain: forecasting algorithms.
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
Machine Learning Pipelines for Kubeflow
Deep Learning GPU Training System
State of the Art Natural Language Processing
Everything you need to know to build your own RAG application
Build data pipelines, the easy way 🛠️
⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / Ultralytics / MMEngine / Keras etc.
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
Language model tokenization at GB/s
PyTorch implementation of Advantage Actor Critic (A2C), Proximal Policy Optimization (PPO), Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation (ACKTR) and Generative Adversarial Imitation Learning (GAIL).
A retargetable MLIR-based machine learning compiler and runtime toolkit.
TensorFlowOnSpark brings TensorFlow programs to Apache Spark clusters.
Fast Python Collaborative Filtering for Implicit Feedback Datasets
Superfast AI decision making and intelligent processing of multi-modal data.
DeepVariant is an analysis pipeline that uses a deep neural network to call genetic variants from next-generation DNA sequencing data.
Java dataframe and visualization library
A high performance and generic framework for distributed DNN training
AI Infra / AI Orchestration / AI Control Plane
Module for automatic summarization of text documents and HTML pages.
High-Performance Symbolic Regression in Python and Julia
The fastest ⚡️ way to build data pipelines. Develop iteratively, deploy anywhere. ☁️
Alink is the Machine Learning algorithm platform based on Flink, developed by the PAI team of Alibaba computing platform.
Models, data loaders and abstractions for language processing, powered by PyTorch
Fast, flexible and easy to use probabilistic modelling in Python.
Source: GitHub API · Curated · Realtime