Image Understanding381 / 915 projects
Visual understanding model leaderboard featuring the hottest AI image recognition projects on GitHub, including OCR, object detection, and segmentation.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Tesseract Open Source OCR Engine (main repository)
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
Pure Javascript OCR for more than 100 Languages 📖🎉🖥
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
OpenMMLab Detection Toolbox and Benchmark
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow
Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
pix2tex: Using a ViT to convert images of equations into LaTeX code.
This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
Object Detection toolkit based on PaddlePaddle. It supports object detection, instance segmentation, multiple object tracking and real-time multi-person keypoint detection.
超轻量级中文ocr,支持竖排文字识别, 支持ncnn、mnn、tnn推理 ( dbnet(1.8M) + crnn(2.5M) + anglenet(378KB)) 总模型仅4.7M
OCR & Document Extraction using vision models
Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
A paper list of object detection using deep learning.
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
YOLOX is a high-performance anchor-free YOLO, exceeding yolov3~v5 with MegEngine, ONNX, TensorRT, ncnn, and OpenVINO supported. Documentation: https://yolox.readthedocs.io/
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
Best Practices, code samples, and documentation for Computer Vision.
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
A python library built to empower developers to build applications and systems with self-contained Computer Vision capabilities
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
OpenMMLab's next-generation platform for general 3D object detection.
NanoDet-Plus⚡Super fast and lightweight anchor-free object detection model. 🔥Only 980 KB(int8) / 1.8MB (fp16) and run 97FPS on cellphone🔥
Transforms PDF, Documents and Images into Enriched Structured Data
YOLOv6: a single-stage object detection framework dedicated to industrial applications.
A Unified Toolkit for Deep Learning Based Document Image Analysis
OpenPCDet Toolbox for LiDAR-based 3D Object Detection.
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
computer vision and sports
The pytorch re-implement of the official efficientdet with SOTA performance in real time and pretrained weights.
A PyTorch Implementation of Single Shot MultiBox Detector
Segmentation models with pretrained backbones. Keras and TensorFlow Keras.
PointNet and PointNet++ implemented by pytorch (pure python) and on ModelNet, ShapeNet and S3DIS.
OpenMMLab Text Detection, Recognition and Understanding Toolbox
[ECCV 2022] This is the official implementation of BEVFormer, a camera-only framework for autonomous driving perception, e.g., 3D object detection and semantic map segmentation.
Single Shot MultiBox Detector in TensorFlow
A Python package for segmenting geospatial data with the Segment Anything Model (SAM)
A simplified implemention of Faster R-CNN that replicate performance from origin paper
OpenMMLab Pre-training Toolbox and Benchmark
Tensorflow Faster RCNN for Object Detection
🔥 TensorFlow Code for technical report: "YOLOv3: An Incremental Improvement"
AdelaiDet is an open source toolbox for multiple instance-level detection and recognition tasks.
OpenMMLab YOLO series toolbox and benchmark. Implemented RTMDet, RTMDet-Rotated,YOLOv5, YOLOv6, YOLOv7, YOLOv8,YOLOX, PPYOLOE, etc.
text detection mainly based on ctpn model in tensorflow, id card detect, connectionist text proposal network
FCOS: Fully Convolutional One-Stage Object Detection (ICCV'19)
A PyTorch implementation of the YOLO v3 object detection algorithm
D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement [ICLR 2025 Spotlight]
[ICRA'23] BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation
[验证码识别-训练] This project is based on CNN/ResNet/DenseNet+GRU/LSTM+CTC/CrossEntropy to realize verification code identification. This project is only for training the model.
A Repo For Document AI
DAMO-YOLO: a fast and accurate object detection method with some new techs, including NAS backbones, efficient RepGFPN, ZeroHead, AlignedOTA, and distillation enhancement.
🔥🔥🔥🔥 (Earlier YOLOv7 not official one) YOLO with Transformers and Instance Segmentation, with TensorRT acceleration! 🔥🔥🔥
A Simple and Versatile Framework for Object Detection and Instance Recognition
Mask RCNN in TensorFlow
A tensorflow implementation of EAST text detector
[python3.6] 运用tf实现自然场景文字检测,keras/pytorch实现ctpn+crnn+ctc实现不定长场景文字OCR识别
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
[CVPR 2023 Highlight] InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
[ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection"
Optical character recognition for Japanese text, with the main focus being Japanese manga
[ECCV2024] API code for T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
Lightweight Python library for adding real-time multi-object tracking to any detector.
YoloV3 Implemented in Tensorflow 2.0
YOLO ROS: Real-Time Object Detection for ROS
Lightweight nudity detection
An extension of Open3D to address 3D Machine Learning tasks
detrex is a research platform for DETR-based object detection, segmentation, pose estimation and other visual recognition tasks.
YOLOv4, YOLOv4-tiny, YOLOv3, YOLOv3-tiny Implemented in Tensorflow 2.0, Android. Convert YOLO v4 .weights tensorflow, tensorrt and tflite
You Only Look Once for Panopitic Driving Perception.(MIR2022)
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
Use of Attention Gates in a Convolutional Neural Network / Medical Image Classification and Segmentation
Scaled-YOLOv4: Scaling Cross Stage Partial Network
[DEIMv2] Real Time Object Detection Meets DINOv3
Codes for our paper "CenterNet: Keypoint Triplets for Object Detection" .
A Keras port of Single Shot MultiBox Detector
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
Set of methods to ensemble boxes from different object detection models, including implementation of "Weighted boxes fusion (WBF)" method.
SECOND for KITTI/NuScenes object detection
Deep Hough Voting for 3D Object Detection in Point Clouds
SOLO and SOLOv2 for instance segmentation, ECCV 2020 & NeurIPS 2020.
MobileNetV2-YoloV3-Nano: 0.5BFlops 3MB HUAWEI P40: 6ms/img, YoloFace-500k:0.1Bflops 420KB:fire::fire::fire:
Implementation of "YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception".
This is a pytorch repository of YOLOv4, attentive YOLOv4 and mobilenet YOLOv4 with PASCAL VOC and COCO
Frustum PointNets for 3D Object Detection from RGB-D Data
A PyTorch impl of EfficientDet faithful to the original Google impl w/ ported weights
Training and Detecting Objects with YOLO3
[CVPR 2025] DEIM: DETR with Improved Matching for Fast Convergence
High quality, fast, modular reference implementation of SSD in PyTorch
动态语义SLAM 目标检测+VSLAM+光流/多视角几何动态物体检测+octomap地图+目标数据库
World's first general purpose 3D object detection codebse.
Complete YOLO v3 TensorFlow implementation. Support training on your own dataset.
[CVPR 2023] Official implementation of the paper "Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation"
[CVPR2026] Detect Anything via Next Point Prediction
"Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement" (ICCV 2023 Top-10 Cited 🏆) & (NTIRE 2024 Runner-Up 🏆) & (NTIRE 2025 Winner 🏆) & (NTIRE 2026 Winner 🏆)
[ICLR 2023 Spotlight] Vision Transformer Adapter for Dense Predictions
Implementation EfficientDet: Scalable and Efficient Object Detection in PyTorch
Single-Shot Refinement Neural Network for Object Detection, CVPR, 2018
Real-time and accurate open-vocabulary end-to-end object detection
Using Diffusion Models to Segment/Reconstruct Organs from Medical Images [AAAI Most influential Paper]
PANet for Instance Segmentation and Object Detection
The PyTorch Implementation based on YOLOv4 of the paper: "Complex-YOLO: Real-time 3D Object Detection on Point Clouds"
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
xmnlp:提供中文分词, 词性标注, 命名体识别,情感分析,文本纠错,文本转拼音,文本摘要,偏旁部首,句子表征及文本相似度计算等功能
Build computer vision models in a fraction of the time and with less data.
StarDist - Object Detection with Star-convex Shapes
Grounding DINO 1.5: IDEA Research's Most Capable Open-World Object Detection Model Series
(CVPR 2021 Oral) Open World Object Detection
Visit PixelLib's official documentation https://pixellib.readthedocs.io/en/latest/
[ECCV2022] PETR: Position Embedding Transformation for Multi-View 3D Object Detection & [ICCV2023] PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images
Caffe implementation of multiple popular object detection frameworks
EntitySeg Toolbox: Towards Open-World and High-Quality Image Segmentation
Lightweight models for real-time semantic segmentationon PyTorch (include SQNet, LinkNet, SegNet, UNet, ENet, ERFNet, EDANet, ESPNet, ESPNetv2, LEDNet, ESNet, FSSNet, CGNet, DABNet, Fast-SCNN, ContextNet, FPENet, etc.)
High level network definitions with pre-trained weights in TensorFlow
MetaSeg: Packaged version of the Segment Anything repository
:art: Pytorch YOLO v5 训练自己的数据集超详细教程!!! :art: (提供PDF训练教程下载)
A general list of resources to image text localization and recognition 场景文本位置感知与识别的论文资源与实现合集 シーンテキストの位置認識と識別のための論文リソースの要約
Code for 3D object detection for autonomous driving
Pix2Seq codebase: multi-tasks with generative modeling (autoregressive and diffusion)
Repository of Vision Transformer with Deformable Attention (CVPR2022) and DAT++: Spatially Dynamic Vision Transformerwith Deformable Attention
Medical SAM 2: Segment 3D Medical Images Via Segment Anything Model 2
:oncoming_automobile: "MORE THAN VEHICLE COUNTING!" This project provides prediction for speed, color and size of the vehicles with TensorFlow Object Counting API.
Semi-Supervised Learning, Object Detection, ICCV2021
[ICLR 2024] Official PyTorch implementation of FasterViT: Fast Vision Transformers with Hierarchical Attention
This repository is an unoffical PyTorch implementation of Medical segmentation in 2D and 3D.
A Kitti Road Segmentation model implemented in tensorflow.
[ICLR2022] official implementation of UniFormer
[NeurIPS 2021] You Only Look at One Sequence
[CVPR 2020] CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement
A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。
Implementation of YOLO v3 object detector in Tensorflow (TF-Slim)
Official Pytorch Code for "Medical Transformer: Gated Axial-Attention for Medical Image Segmentation" - MICCAI 2021
:zap: A newly designed ultra lightweight anchor free target detection algorithm, weight only 250K parameters, reduces the time consumption by 10% compared with yolo-fastest, and the post-processing is simpler
Darknet/YOLO object detection framework
This is the third party implementation of the paper Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.
OCR software for recognition of handwritten text
A powerful baseline for image classification, face recognition and image retrieval with Pytorch
YoloDotNet - A C# .NET 8.0 project for Classification, Object Detection, OBB Detection, Segmentation and Pose Estimation in both images and live video streams.
This repo is the codebase for our team to participate in DOTA related competitions, including rotation and horizontal detection.
CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包
[CVPR 2020] CenterMask : Real-Time Anchor-Free Instance Segmentation
Precise RoI Pooling with coordinate gradient support, proposed in the paper "Acquisition of Localization Confidence for Accurate Object Detection" (https://arxiv.org/abs/1807.11590).
[CVPR 2020] CenterMask : Real-time Anchor-Free Instance Segmentation
A mobilenet SSD based face detector, powered by tensorflow object detection api, trained by WIDERFACE dataset.
MXNet port of SSD: Single Shot MultiBox Object Detector. Reimplementation of https://github.com/weiliu89/caffe/tree/ssd
experiments on Paper <Bag of Tricks for Image Classification with Convolutional Neural Networks> and other useful tricks to improve CNN acc
Random Erasing Data Augmentation. Experiments on CIFAR10, CIFAR100 and Fashion-MNIST
Bounding Box Regression with Uncertainty for Accurate Object Detection (CVPR'19)
DSOD: Learning Deeply Supervised Object Detectors from Scratch. In ICCV 2017.
Hierarchical perception library in Python for pose estimation, object detection, instance segmentation, keypoint estimation, face recognition, etc.
Multi-object trackers in Python
More readable and flexible yolov5 with more backbone(gcn, resnet, shufflenet, moblienet, efficientnet, hrnet, swin-transformer, etc) and (cbam,dcn and so on), and tensorrt
Real-time object detection on Android using the YOLO network with TensorFlow
Implementation of Bottleneck Transformer in Pytorch
Train a state-of-the-art yolov3 object detector from scratch!
an experiment for yolo-v1, including training and testing.
Gaussian YOLOv3: An Accurate and Fast Object Detector Using Localization Uncertainty for Autonomous Driving (ICCV, 2019)
基于深度学习的驾驶员分心驾驶行为(疲劳+危险行为)预警系统使用YOLOv5+Deepsort实现驾驶员的危险驾驶行为的预警监测
FreeAnchor: Learning to Match Anchors for Visual Object Detection (NeurIPS 2019)
🚀🚀🚀 YOLO series of PaddlePaddle implementation, PP-YOLOE+, RT-DETR, YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLOv10, YOLO11, YOLOX, YOLOv5u, YOLOv7u, YOLOv6Lite, RTMDet and so on. 🚀🚀🚀
based on yolo-high-level project (detect\pose\classify\segment\):include yolov5\yolov7\yolov8\ core ,improvement research ,SwintransformV2 and Attention Series. training skills, business customization, engineering deployment C
This repository provides train&test code, dataset, det.&rec. annotation, evaluation script, annotation tool, and ranking.
YOLOv7 Object Tracking Using PyTorch, OpenCV and Sort Tracking
Full implementation of YOLOv3 in PyTorch
Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
(Pretrained weights provided) EfficientDet: Scalable and Efficient Object Detection implementation by Signatrix GmbH
[CVPR 2022] SparseInst: Sparse Instance Activation for Real-Time Instance Segmentation
Next-generation Video instance recognition framework on top of Detectron2 which supports InstMove (CVPR 2023), SeqFormer(ECCV Oral), and IDOL(ECCV Oral))
[CVPR 2025 Highlight] Official code and models for Encoder-only Mask Transformer (EoMT).
Easy & Modular Computer Vision Detectors, Trackers & SAM - Run YOLOv9,v8,v7,v6,v5,R,X in under 10 lines of code.
OBBDetection is an oriented object detection library, which is based on MMdetection.
A Wide Range of Custom Functions for YOLOv4, YOLOv4-tiny, YOLOv3, and YOLOv3-tiny Implemented in TensorFlow, TFLite, and TensorRT.
[CVPR 2024] Aligning and Prompting Everything All at Once for Universal Visual Perception
[CVPR 2022 Oral] Official implementation of DN-DETR
Keras implementation of yolo v3 object detection.
This is the implementation of GGHL (A General Gaussian Heatmap Label Assignment for Arbitrary-Oriented Object Detection)
Official code for Conformer: Local Features Coupling Global Representations for Visual Recognition
Camouflaged Object Detection, CVPR 2020 (Oral)
Unofficial implementation for [ECCV'22] "Exploring Plain Vision Transformer Backbones for Object Detection"
[CVPR 2023] Official code release of our paper "BiFormer: Vision Transformer with Bi-Level Routing Attention"
Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers [CVPR 2021]
🚀 A high performance real-time object detection solution using YOLO11 ⚡️ powered by ONNX-Runtime
I transfer the backend of yolov3 into Mobilenetv1,VGG16,ResNet101 and ResNeXt101
Detectorch - detectron for PyTorch
This repo is implemented based on detectron2 and centernet
Computer vision based vehicle detection and tracking using Tensorflow Object Detection API and Kalman-filtering
Keras package for region-based convolutional neural networks (RCNNs)
Cool experiments at the intersection of Computer Vision and Sports ⚽🏃
An accurate GUI element detection approach based on old-fashioned CV algorithms [Upgraded on 5/July/2021]
Pytorch implementation of ResUnet and ResUnet ++
Code for AAAI 2021 paper: R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating Object
An easy to understand and better performance version of CenterNet
In-Browser Object Detection using Tiny YOLO on Tensorflow.js
Out-of-the-box code and models for CMU's object detection and tracking system for multi-camera surveillance videos. Speed optimized Faster-RCNN model. Tensorflow based. Also supports EfficientDet. WACVW'20
A brand logo detection system using tensorflow object detection API.
Source: GitHub API · Curated · Realtime