Models3201 projects
A curated leaderboard of the hottest AI model projects on GitHub, including LLMs, vision models, multimodal models, and more, ranked by Stars.
Robust Speech Recognition via Large-Scale Weak Supervision
DeepSeek-V3 (⭐103K+), is a leading open-source project in Models, built with Python, a benchmark project in its domain, ideal for open_source AI development workflows.
Leading open-weight LLM series by DeepSeek. Mixture-of-Experts architecture, competitive with top proprietary models.
DeepSeek-R1 (⭐91.9K), is a leading open-source project in Models, built with AI, a benchmark project in its domain, ideal for open_source AI development workflows.
opencv (⭐89.8K), is a leading open-source project in Models, built with C++, focusing on c plus plus, with computer vision capabilities, a benchmark project in its domain, ideal for open_source AI development workflows.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Tesseract Open Source OCR Engine (main repository)
stable-diffusion (⭐73.1K), is a leading open-source project in Models, built with Jupyter Notebook, a benchmark project in its domain, ideal for open_source AI development workflows.
Powerful node-based Stable Diffusion workflow editor. Modular pipeline for advanced image generation.
AI that turns screenshots into code. Drop in an image, get HTML/CSS/JS out. Supports Tailwind CSS.
Stability AI's foundational text-to-image model. Open weights, high quality, huge ecosystem of tools.
Meta's open-source LLM family. State-of-the-art performance, available in 8B to 405B parameters.
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
The world's simplest facial recognition api for Python and the command line
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Deep learning text-to-speech. Supports 1100+ languages, voice cloning, and fine-tuning.
ChatGLM-6B: An Open Bilingual Dialogue Language Model | 开源双语对话语言模型
TensorFlow code and pre-trained models for BERT
A generative speech model for daily dialogue.
🔊 Text-Prompted Generative Audio Model
Pure Javascript OCR for more than 100 Languages 📖🎉🖥
GFPGAN aims at developing Practical Algorithms for Real-world Face Restoration.
Instant voice cloning by MIT and MyShell. Audio foundation model.
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
Cross-platform, customizable ML solutions for live and streaming media.
Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
Let us control diffusion models!
OpenMMLab Detection Toolbox and Benchmark
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Real-time face swap for PC streaming or video calls
Code and documentation to train Stanford's Alpaca models, and generate the data.
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
State-of-the-art 2D and 3D Face Analysis Project
Deezer source separation library including pretrained models.
SoftVC VITS Singing Voice Conversion
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest AI-driven technologies. The solution offers an industry leading WebUI, and serves as the foundation for multiple commercial products.
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
Generative Models by Stability AI
DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
SoTA open-source TTS
Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow
Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
Image-to-Image Translation in PyTorch
Code for the paper "Language Models are Unsupervised Multitask Learners"
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
DeepSeek Coder: Let the Code Write Itself
Contexts Optical Compression
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Meta's generative audio AI. MusicGen for music, AudioGen for sound effects, and EnCodec for compression.
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
FinGPT: Open-Source Financial Large Language Models! Revolutionize 🔥 We release the trained model on HuggingFace.
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
A TTS model capable of generating ultra-realistic dialogue in one pass.
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
Bring portraits to life!
Torch implementation of neural style algorithm
[NeurIPS 2022] Towards Robust Blind Face Restoration with Codebook Lookup Transformer
Open-source multimodal LLM with vision, speech, and text. Strong performance on mobile and edge devices.
JavaScript API for face detection and face recognition in the browser and nodejs with tensorflow.js
Stable Diffusion with Core ML on Apple Silicon
Janus-Series: Unified Multimodal Understanding and Generation Models
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
Instant neural graphics primitives: lightning fast NeRF and more
Wan: Open and Advanced Large-Scale Video Generative Models
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
Wan: Open and Advanced Large-Scale Video Generative Models
pix2tex: Using a ViT to convert images of equations into LaTeX code.
This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
GPT-3: Language Models are Few-Shot Learners
Bringing Old Photo Back to Life (CVPR 2020 oral)
StableLM: Stability AI Language Models
kaldi-asr/kaldi is the official location of the Kaldi project.
Face recognition with deep neural networks.
End-to-End Object Detection with Transformers
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
Face recognition using Tensorflow
Object Detection toolkit based on PaddlePaddle. It supports object detection, instance segmentation, multiple object tracking and real-time multi-person keypoint detection.
Implementation of paper - YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
High-Resolution Image Synthesis with Latent Diffusion Models
An open source implementation of CLIP.
🎬 火宝短剧 - 基于AI的一站式短剧生成平台 《一句话生成完整短剧,从剧本到成片全自动化》 Huobao Drama - An AI-Powered End-to-End Short Drama Generator "One Sentence to Complete Drama: Fully Automated from Script to Final Video"
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
首家工业级全流程 AI 影视生产平台。Industry-first professional AI Agent platform for controllable film & video production. From shorts to live-action with Hollywood-standard workflows.
PaddleFormers is an easy-to-use library of pre-trained large language model zoo based on PaddlePaddle.
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
Enjoy the magic of Diffusion models!
Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
HunyuanVideo: A Systematic Framework For Large Video Generation Model
COLMAP - Structure-from-Motion and Multi-View Stereo
超轻量级中文ocr,支持竖排文字识别, 支持ncnn、mnn、tnn推理 ( dbnet(1.8M) + crnn(2.5M) + anglenet(378KB)) 总模型仅4.7M
OCR & Document Extraction using vision models
100+ Chinese Word Vectors 上百种预训练中文词向量
An open-source tool-augmented conversational language model from Fudan University
Multi-layer Recurrent Neural Networks (LSTM, GRU, RNN) for character-level language models in Torch
Meta's AI that brings children's drawings to life. Automatically animates humanoid figures.
An open-source NLP research library, built on PyTorch.
A collaboration friendly studio for NeRFs
Super Resolution for images using deep learning.
A PyTorch-based Speech Toolkit
Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
PyTorch implementation of the U-Net for image semantic segmentation with high quality images
A paper list of object detection using deep learning.
YOLOv10: Real-Time End-to-End Object Detection [NeurIPS 2024]
🐍 Geometric Computer Vision Library for Spatial AI
A fast, local neural text to speech system
LAVIS - A One-stop Library for Language-Vision Intelligence
Kimi K2 is the large language model series developed by Moonshot AI team
一款入门级的人脸、视频、文字检测以及识别的项目.
Custom photo generation with Stable Diffusion. Generate personalized images from a few photos.
TensorFlow CNN for fast style transfer ⚡🖥🎨🖼
[CVPR 2024] Official repository for "MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model"
Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
Image-to-image translation with conditional adversarial nets
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
YOLOX is a high-performance anchor-free YOLO, exceeding yolov3~v5 with MegEngine, ONNX, TensorRT, ncnn, and OpenVINO supported. Documentation: https://yolox.readthedocs.io/
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
An easy 1-click way to create beautiful artwork on your PC using AI, with no tech knowledge. Provides a browser UI for generating images from text prompts and images. Just enter your text prompt, and see the generated image.
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
Pre-Training with Whole Word Masking for Chinese BERT(中文BERT-wwm系列模型)
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.
tiny vision language model
End-to-End Speech Processing Toolkit
OpenMMLab Semantic Segmentation Toolbox and Benchmark.
Best Practices, code samples, and documentation for Computer Vision.
The code for our newly accepted paper in Pattern Recognition 2020: "U^2-Net: Going Deeper with Nested U-Structure for Salient Object Detection."
A PyTorch implementation of the Transformer model in "Attention is All You Need".
ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.
Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!
A modern approach for Computer Vision on the web
Easy-to-use image segmentation library with awesome pre-trained model zoo, supporting wide-range of practical tasks in Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, 3D Segmentation, etc.
Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.
Keras implementations of Generative Adversarial Networks.
Multilingual Document Layout Parsing in a Single Vision-Language Model
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Multilingual Document Layout Parsing in a Single Vision-Language Model
Speech recognition module for Python, supporting several engines and APIs, online and offline.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
A python library built to empower developers to build applications and systems with self-contained Computer Vision capabilities
Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
Text-audio foundation model from Boson AI
BELLE: Be Everyone's Large Language model Engine(开源中文对话大模型)
An implementation of model parallel GPT-2 and GPT-3-style models using the mesh-tensorflow library.
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
Pythonic AI generation of images and videos
Leading free and open-source face recognition system
PaddlePaddle GAN library, including lots of interesting applications like First-Order motion transfer, Wav2Lip, picture repair, image editing, photo2cartoon, image style transfer, GPEN, and so on.
Code for the paper "Jukebox: A Generative Model for Music"
PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models
all kinds of text classification models and more with deep learning
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
Run Stable Diffusion on Mac natively
fast-stable-diffusion + DreamBooth
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
Stable Diffusion web UI
Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLM
Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages
Leveraging BERT and c-TF-IDF to create easily interpretable topics.
⚡ TabPFN: Foundation Model for Tabular Data ⚡
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Source: GitHub API · Curated · Realtime