Vision-Language112 / 161 projects
Vision-Language model leaderboard featuring the hottest VLM projects on GitHub that understand both images and natural language.
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
LAVIS - A One-stop Library for Language-Vision Intelligence
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
Solve Visual Understanding with Reinforced VLMs
PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/spaces/TencentARC/Caption-Anything https://huggingface.co/spaces/VIPLab/Caption-Anything
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Bottom-up attention model for image captioning and VQA, based on Faster R-CNN and Visual Genome
The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
Implementation of 🦩 Flamingo, state-of-the-art few-shot visual question answering attention net out of Deepmind, in Pytorch
JoyCaption is an image captioning Visual Language Model (VLM) being built from the ground up as a free, open, and uncensored model for the community to use in training Diffusion models.
Oscar and VinVL
Unofficial pytorch implementation for Self-critical Sequence Training for Image Captioning. and others.
X-modaler is a versatile and high-performance codebase for cross-modal analytics(e.g., image captioning, video captioning, vision-language pre-training, visual question answering, visual commonsense reasoning, and cross-modal retrieval).
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
TensorFlow Implementation of "Show, Attend and Tell"
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
NEO Series: Native Vision-Language Models from First Principles
Open-source SOTA multi-image editing model
[CVPR 2026🔥] 🧑🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator that produces Lottie JSONs.
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
Evaluating text-to-image/video/3D models with VQAScore
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
PyTorch source code for "Stacked Cross Attention for Image-Text Matching" (ECCV 2018)
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
[CVPR 2021] VirTex: Learning Visual Representations from Textual Annotations
Bilinear attention networks for visual question answering
Meshed-Memory Transformer for Image Captioning. CVPR 2020
Official Pytorch implementation of "OmniNet: A unified architecture for multi-modal multi-task learning" | Authors: Subhojeet Pramanik, Priyanka Agrawal, Aman Hussain
[ ICLR 2024 ] Official Codebase for "InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists"
An open-source implementation for training LLaVA-NeXT.
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
Video to Text: Natural language description generator for some given video. [Video Captioning]
Code for paper "Attention on Attention for Image Captioning". ICCV 2019
Implementation of "Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning"
Image Captioning using InceptionV3 and beam search
Transformer-based image captioning extension for pytorch/fairseq
A neural network to generate captions for an image using CNN and RNN with BEAM Search.
Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions. CVPR 2019
Phi-4 for Mac: Locally-run Vision and Language Models for Apple Silicon
Phi-3.5 for Mac: Locally-run Vision and Language Models for Apple Silicon
Implementation of 'X-Linear Attention Networks for Image Captioning' [CVPR 2020]
Image Captioning Using Transformer
A modular library built on top of Keras and TensorFlow to generate a caption in natural language for any input image.
Language Models Can See: Plugging Visual Controls in Text Generation
Automatic image captioning model based on Caffe, using features from bottom-up attention.
PyTorch code for "Fine-grained Image Captioning with CLIP Reward" (Findings of NAACL 2022)
Foundation models based medical image analysis
[ICLR 2026] An official implementation of "CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning"
A pytorch implementation of On the Automatic Generation of Medical Imaging Reports.
Image Captions Generation with Spatial and Channel-wise Attention
Official pytorch implementation of paper "Dual-Level Collaborative Transformer for Image Captioning" (AAAI 2021).
GRIT: Faster and Better Image-captioning Transformer (ECCV 2022)
Code for "Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner" in ICCV 2017
[DEPRECATED] A Neural Network based generative model for captioning images using Tensorflow
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
Official Code for 'RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words' (CVPR 2021)
Computer vision tools for fairseq, containing PyTorch implementation of text recognition and object detection
Generating Captions for images using Deep Learning
CLIPxGPT Captioner is Image Captioning Model based on OpenAI's CLIP and GPT-2.
Show and Tell : A Neural Image Caption Generator
TensorFlow (TensorLayer) Implementation of Image Captioning
ZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.
Pytorch Implementation of Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
Image Captioning based on Bottom-Up and Top-Down Attention model
Using pretrained encoder and language models to generate captions from multimedia inputs.
[NeurIPS'25 | CVPR'26] The official repo of OralGPT & MMOral Bench.
[ICLR2025 Oral] ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
Sleek, mobile-friendly web UI for NVIDIA LocateAnything-3B — open-vocabulary object detection & grounding on your own GPU, via one docker compose up.
A minimal implementation of LLaVA-style VLM with interleaved image & text & video processing ability.
VL-JEPA (Vision-Language Joint Embedding Predictive Architecture) in MLX
mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections. (EMNLP 2022)
Implementation code of the work "Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning"
[CVPR 2020] Transform and Tell: Entity-Aware News Image Captioning
This repository contains the implementation for the paper "Revisiting Few Shot Object Detection with Vision-Language Models"
AI-powered visual reasoning tools for broadcast & ProAV. PTZ camera tracking, object detection, scene analysis using Moondream VLM. By StreamGeeks & PTZOptics.
Improving Chest X-Ray Report Generation by Leveraging Warm-Starting
[Paper][ISWC 2021] Zero-shot Visual Question Answering using Knowledge Graph
[ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"
[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
Multimodal Instruction Tuning for Llama 3
Combining ViT and GPT-2 for image captioning. Trained on MS-COCO. The model was implemented mostly from scratch.
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Toy-scale unified multimodal model experiments — encoder-free understanding & generation with Mixture-of-Transformers on MLX/Apple Silicon
An end-to-end image captioning project using a CNN encoder (ResNet-50) and LSTM decoder in PyTorch. Includes vocabulary building, preprocessing, training with BLEU evaluation, and inference. Generates natural language captions for images with saved metrics, model checkpoints, and visualization outputs.
Code repository for "Post-pre-training for Modality Alignment in Vision-Language Foundation Models" (CVPR2025)
This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"
Inverse DALL-E for Optical Character Recognition
Tool to automatically generate text descriptions for images using Ollama vision models (LLaVA, Qwen3-VL, Llama Vision)
[AAAI 2025] Official Implementation of I-HallA v1.0
Source: GitHub API · Curated · Realtime