Vision Models527 / 1896 projects
Vision models category featuring top computer vision AI projects on GitHub, including text-to-image, text-to-video, image editing, and visual understanding.
opencv (⭐89.8K), is a leading open-source project in Models, built with C++, focusing on c plus plus, with computer vision capabilities, a benchmark project in its domain, ideal for open_source AI development workflows.
Powerful node-based Stable Diffusion workflow editor. Modular pipeline for advanced image generation.
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
Cross-platform, customizable ML solutions for live and streaming media.
OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest AI-driven technologies. The solution offers an industry leading WebUI, and serves as the foundation for multiple commercial products.
Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
Image-to-Image Translation in PyTorch
Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
Instant neural graphics primitives: lightning fast NeRF and more
pix2tex: Using a ViT to convert images of equations into LaTeX code.
Face recognition using Tensorflow
An open source implementation of CLIP.
Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.
COLMAP - Structure-from-Motion and Multi-View Stereo
A collaboration friendly studio for NeRFs
Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
🐍 Geometric Computer Vision Library for Spatial AI
Custom photo generation with Stable Diffusion. Generate personalized images from a few photos.
Image-to-image translation with conditional adversarial nets
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022
Best Practices, code samples, and documentation for Computer Vision.
The code for our newly accepted paper in Pattern Recognition 2020: "U^2-Net: Going Deeper with Nested U-Structure for Salient Object Detection."
Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
Leading free and open-source face recognition system
PaddlePaddle GAN library, including lots of interesting applications like First-Order motion transfer, Wav2Lip, picture repair, image editing, photo2cartoon, image style transfer, GPEN, and so on.
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
a cross-platform image super-resolution tool
Real-Time High-Resolution Background Matting
Synthesizing and manipulating 2048x1024 images with conditional GANs
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
open Multiple View Geometry library. Basis for 3D computer vision and Structure from Motion.
(CGCSTCD'2017) An easy, flexible, and accurate plate recognition project for Chinese licenses in unconstrained situations. CGCSTCD = China Graduate Contest on Smart-city Technology and Creative Design
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
A Unified Toolkit for Deep Learning Based Document Image Analysis
ECCV2022 - Real-Time Intermediate Flow Estimation for Video Frame Interpolation
computer vision and sports
A PyTorch Implementation of Single Shot MultiBox Detector
Tiled Diffusion and VAE optimize, licensed under CC BY-NC-SA 4.0
Unofficial Implementation of DragGAN - "Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold" (DragGAN 全功能实现,在线Demo,本地部署试用,代码、模型已全部开源,支持Windows, macOS, Linux)
StableSwarmUI, A Modular Stable Diffusion Web-User-Interface, with an emphasis on making powertools easily accessible, high performance, and extensibility.
Torchreid: Deep learning person re-identification in PyTorch.
🔎 Super-scale your images and run experiments with Residual Dense and Adversarial Networks.
Fast face detection, pupil/eyes localization and facial landmark points detection library in pure Go.
Official implementation of "Neuralangelo: High-Fidelity Neural Surface Reconstruction" (CVPR 2023)
SenseTime Research platform for single object tracking, implementing algorithms like SiamRPN and SiamMask.
[ECCV 2022] This is the official implementation of BEVFormer, a camera-only framework for autonomous driving perception, e.g., 3D object detection and semantic map segmentation.
:man: Code for "Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression"
[ICCV 2019] Monocular depth estimation from a single image
An OBS plugin for removing background in portrait images (video), making it easy to replace the background when recording or streaming.
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
Official implementations for paper: Anydoor: zero-shot object-level image customization
An open-source framework for training large multimodal models.
3D ResNets for Action Recognition (CVPR 2018)
人像卡通化探索项目 (photo-to-cartoon translation project)
Interactive Image Generation via Generative Adversarial Networks
SOTA Re-identification Methods and Toolbox
Unofficial implementation of Image Super-Resolution via Iterative Refinement by Pytorch
[CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
Scenic: A Jax Library for Computer Vision Research and Beyond
A python library for self-supervised learning on images.
The PyTorch improved version of TPAMI 2017 paper: Face Alignment in Full Pose Range: A 3D Total Solution.
🔥🔥High-Performance Face Recognition Library on PaddlePaddle & PyTorch🔥🔥
[CVPR19/TPAMI23] SiamMask: A Framework for Fast Online Object Tracking and Segmentation
Visual tracking library based on PyTorch.
[CVPR2020] Adversarial Latent Autoencoders
3D Computer Vision Framework
Automatic colorization using deep neural networks. "Colorful Image Colorization." In ECCV, 2016.
FCOS: Fully Convolutional One-Stage Object Detection (ICCV'19)
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) by way of Textual Inversion (https://arxiv.org/abs/2208.01618) for Stable Diffusion (https://arxiv.org/abs/2112.10752). Tweaks focused on training faces, objects, and styles.
The official PyTorch implementation of Towards Fast, Accurate and Stable 3D Dense Face Alignment, ECCV 2020.
[CVPR 2025] MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors
Foundation Architecture for (M)LLMs
Largest multi-label image database; ResNet-101 model; 80.73% top-1 acc on ImageNet
OpenVSLAM: A Versatile Visual SLAM Framework
[ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection"
Kandinsky 2 — multilingual text2image latent diffusion model
Optical character recognition for Japanese text, with the main focus being Japanese manga
Deep learning software for colorizing black and white images with a few clicks.
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
Image Deblurring using Generative Adversarial Networks
Automatically remove the mosaics in images and videos, or add mosaics to them.
A Python package for fast and robust Image Stitching
This is the repo for our new project Highly Accurate Dichotomous Image Segmentation
Contrastive unpaired image-to-image translation, faster and lighter training than cyclegan (ECCV 2020, in PyTorch)
One-step image-to-image with Stable Diffusion turbo: sketch2image, day2night, and more
YOLO ROS: Real-Time Object Detection for ROS
ICCV2019 - Learning to Paint With Model-based Deep Reinforcement Learning
A sketch extractor for anime/illustration.
A simple interface for editing natural photos with generative neural networks.
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
Custom Diffusion: Multi-Concept Customization of Text-to-Image Diffusion (CVPR 2023)
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
:unlock: Lip Reading - Cross Audio-Visual Recognition using 3D Architectures
A Keras port of Single Shot MultiBox Detector
[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
Discovering Interpretable GAN Controls [NeurIPS 2020]
MobileNetV2-YoloV3-Nano: 0.5BFlops 3MB HUAWEI P40: 6ms/img, YoloFace-500k:0.1Bflops 420KB:fire::fire::fire:
Meta-Transformer for Unified Multimodal Learning
Generate images from texts. In Russian
Face Mask Detection system based on computer vision and deep learning using OpenCV and Tensorflow/Keras
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
High quality, fast, modular reference implementation of SSD in PyTorch
ChainerCV: a Library for Deep Learning in Computer Vision
Implementation EfficientDet: Scalable and Efficient Object Detection in PyTorch
Visual intelligence for your home.
This repository is the official implementation of Disentangling Writer and Character Styles for Handwriting Generation (CVPR 2023)
Real-time and accurate open-vocabulary end-to-end object detection
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
Generative Adversarial Transformers
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.
A clean and readable Pytorch implementation of CycleGAN
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
Build computer vision models in a fraction of the time and with less data.
[ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
Paint by Example: Exemplar-based Image Editing with Diffusion Models
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Fast computer vision library for SFM, calibration, fiducials, tracking, image processing, and more.
PyTorch implementation for SDEdit: Image Synthesis and Editing with Stochastic Differential Equations
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
Official PyTorch implementation of "VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization" (CVPR 2021)
Vehicle detection using machine learning and computer vision techniques for Udacity's Self-Driving Car Engineer Nanodegree.
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Visit PixelLib's official documentation https://pixellib.readthedocs.io/en/latest/
Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
EntitySeg Toolbox: Towards Open-World and High-Quality Image Segmentation
Lightweight models for real-time semantic segmentationon PyTorch (include SQNet, LinkNet, SegNet, UNet, ENet, ERFNet, EDANet, ESPNet, ESPNetv2, LEDNet, ESNet, FSSNet, CGNet, DABNet, Fast-SCNN, ContextNet, FPENet, etc.)
Metric learning and retrieval pipelines, models and zoo.
The official implementation of Segment Any 3D GAussians (AAAI-25)
Deep Image Matting
[CVPR 2023 Highlight] Neural Kernel Surface Reconstruction
[CVPR2022] Geometric Transformer for Fast and Robust Point Cloud Registration
This repository implements a demo of the networks described in "How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)" paper.
Pytorch implementation of MixNMatch
[ICLR 2026] A Training-free Iterative Framework for Long Story Visualization
A real-time method that estimates the 3D human pose directly in the popular Bio Vision Hierarchy (BVH) format, given estimations of the 2D body joints originating from monocular color images. Our contributions include: (a) A novel and compact 2D pose NSRM representation. (b) A human body orientation classifier and an ensemble of orientation-tuned
Pix2Seq codebase: multi-tasks with generative modeling (autoregressive and diffusion)
[ECCV 2026] Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
Gated-Shape CNN for Semantic Segmentation (ICCV 2019)
[ICRA 2025 Best Paper] MAC-VO: Metrics-aware Covariance for Learning-based Stereo Visual Odometry
:oncoming_automobile: "MORE THAN VEHICLE COUNTING!" This project provides prediction for speed, color and size of the vehicles with TensorFlow Object Counting API.
This project extends the idea of the innovative architecture of Kolmogorov-Arnold Networks (KAN) to the Convolutional Layers, changing the classic linear transformation of the convolution to learnable non linear activations in each pixel.
[CVPR 2024 Oral] Rethinking Inductive Biases for Surface Normal Estimation
Official PyTorch implementation for the paper High-Resolution Virtual Try-On with Misalignment and Occlusion-Handled Conditions (ECCV 2022).
A Kitti Road Segmentation model implemented in tensorflow.
[CVPR 2022] FaceFormer: Speech-Driven 3D Facial Animation with Transformers
[CVPR 2016] Unsupervised Feature Learning by Image Inpainting using GANs
3DMatch - a 3D ConvNet-based local geometric descriptor for aligning 3D meshes and point clouds.
침착한 생성모델 학습기
[NeurIPS 2021] You Only Look at One Sequence
[CVPR 2020] CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement
Learning to Regress 3D Face Shape and Expression from an Image without 3D Supervision
A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。
CVPR2023-Occupancy-Prediction-Challenge
Pytorch implementation for "Large-Scale Long-Tailed Recognition in an Open World" (CVPR 2019 ORAL)
Open-source SOTA multi-image editing model
Fine-tune SAM (Segment Anything Model) for computer vision tasks such as semantic segmentation, matting, detection ... in specific scenarios
PointFlow : 3D Point Cloud Generation with Continuous Normalizing Flows
Official PyTorch implementation of "Camera Distance-aware Top-down Approach for 3D Multi-person Pose Estimation from a Single RGB Image", ICCV 2019
TensorFlow Implementation of Unsupervised Cross-Domain Image Generation
Learning to Adapt Structured Output Space for Semantic Segmentation, CVPR 2018 (spotlight)
:zap: A newly designed ultra lightweight anchor free target detection algorithm, weight only 250K parameters, reduces the time consumption by 10% compared with yolo-fastest, and the post-processing is simpler
[ICCV 2021 Oral] PoinTr: Diverse Point Cloud Completion with Geometry-Aware Transformers
Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions (ICCV 2023)
Official repository accompanying a CVPR 2022 paper EMOCA: Emotion Driven Monocular Face Capture And Animation. EMOCA takes a single image of a face as input and produces a 3D reconstruction. EMOCA sets the new standard on reconstructing highly emotional images in-the-wild
Vision-and-Language Navigation in Continuous Environments using Habitat
[ICLR2025] Kolmogorov-Arnold Transformer
A Python toolkit for fine-tuning Geospatial Foundation Models (GFMs).
[CVPR 2024] GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models
Implementations of NeRF variants based on Taichi + PyTorch
[ICRA 2022] An opensource framework for cooperative detection. Official implementation for OPV2V.
A very simple CNN project to recognize gestures made in American Sign Language
Example code for the FLAME 3D head model. The code demonstrates how to sample 3D heads from the model, fit the model to 3D keypoints and 3D scans.
[CVPR 2022] "MonoScene: Monocular 3D Semantic Scene Completion": 3D Semantic Occupancy Prediction from a single image
[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that can magically restore any image to perfect-4K!
PyTorch implementation of DeepLabV3, trained on the Cityscapes dataset.
Code for APDrawingGAN: Generating Artistic Portrait Drawings from Face Photos with Hierarchical GANs (CVPR 2019 Oral)
This is a implementation of the 3D FLAME model in PyTorch
Real-time 3D face tracking and reconstruction from 2D video
[ECCV 2024 - Oral] ACE0 is a learning-based structure-from-motion approach that estimates camera parameters of sets of images by learning a multi-view consistent, implicit scene representation.
IJCAI2023 - Collaborative Neural Rendering using Anime Character Sheets
YoloDotNet - A C# .NET 8.0 project for Classification, Object Detection, OBB Detection, Segmentation and Pose Estimation in both images and live video streams.
Rich-Text-to-Image Generation
Learning infinite-resolution image processing with GAN and RL from unpaired image datasets, using a differentiable photo editing model.
Official PyTorch implementation of Revisiting Image Pyramid Structure for High Resolution Salient Object Detection (ACCV 2022)
Precise RoI Pooling with coordinate gradient support, proposed in the paper "Acquisition of Localization Confidence for Accurate Object Detection" (https://arxiv.org/abs/1807.11590).
[CVPR 2021] Anycost GANs for Interactive Image Synthesis and Editing
🤖🖌️ Generate photo-realistic textures based on source images or (soon) PBR materials. Remix, remake, mashup! Useful if you want to create variations on a theme or elaborate on an existing texture.
[ECCV 2024] InstructIR: High-Quality Image Restoration Following Human Instructions https://huggingface.co/spaces/marcosv/InstructIR
Source: GitHub API · Curated · Realtime