Multimodal Models195 / 648 projects
Multimodal models category featuring top cross-modal AI projects on GitHub, including vision-language models, audio-visual understanding, and unified multimodal models.
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Open-source multimodal LLM with vision, speech, and text. Strong performance on mobile and edge devices.
Janus-Series: Unified Multimodal Understanding and Generation Models
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Solve Visual Understanding with Reinforced VLMs
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
OpenMMLab Pre-training Toolbox and Benchmark
🪩 Create Disco Diffusion artworks in one line
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
Foundation Architecture for (M)LLMs
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Easily compute clip embeddings and build a clip retrieval system with them
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
On-device, real-time multimodal AI with features similar to GPT-Live
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.
Meta-Transformer for Unified Multimodal Learning
Fast Multimodal LLM on Mobile Devices
Multimodal-GPT
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Visual intelligence for your home.
This repository is the official implementation of Disentangling Writer and Character Styles for Handwriting Generation (CVPR 2023)
[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
Implementation of CoCa, Contrastive Captioners are Image-Text Foundation Models, in Pytorch
Unifying 3D Mesh Generation with Language Models
MOVA: Towards Scalable and Synchronized Video–Audio Generation
Codebase for Aria - an Open Multimodal Native MoE
🩺 首个会看胸部X光片的中文多模态医学大模型 | The first Chinese Medical Multimodal Model that Chest Radiographs Summarization.
[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
[ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
A Framework of Small-scale Large Multimodal Models
Pix2Seq codebase: multi-tasks with generative modeling (autoregressive and diffusion)
NEO Series: Native Vision-Language Models from First Principles
This is the third party implementation of the paper Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
[ECCV 2024] InstructIR: High-Quality Image Restoration Following Human Instructions https://huggingface.co/spaces/marcosv/InstructIR
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
[CVPR 2019]: Pluralistic Image Completion
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
Official implementation of SEED-LLaMA (ICLR 2024).
Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"
Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
Explore the Multimodal “Aha Moment” on 2B Model
Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets
[ECCV 2024] Tokenize Anything via Prompting
Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
PyTorch source code for "Stacked Cross Attention for Image-Text Matching" (ECCV 2018)
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
Official implementation for "Break-A-Scene: Extracting Multiple Concepts from a Single Image" [SIGGRAPH Asia 2023]
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
Multi-Modal learning toolkit based on PaddlePaddle and PyTorch, supporting multiple applications such as multi-modal classification, cross-modal retrieval and image caption.
[VLDB' 25] ChatTS: LLM for Time Series Understanding and Reasoning
A multimodal Chinese LLM based on LLaMA and Alpaca (VisualCLA).
A CLI tool/python module for generating images from text using guided diffusion and CLIP from OpenAI.
An open-source implementation for training LLaVA-NeXT.
Open-AI's DALL-E for large scale training in mesh-tensorflow.
Towards Generalist Biomedical AI
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Learning.
Visual Med-Alpaca is an open-source, multi-modal foundation model designed specifically for the biomedical domain, built on the LLaMa-7B.
Virtual Sparse Convolution for Multimodal 3D Object Detection
Creates an index of images, queries a local LLM and adds tags to the image metadata
Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.
LLaVA-Interactive-Demo
PyTorch Implementation for Paper "Emotionally Enhanced Talking Face Generation" (ICCVW'23 and ACM-MMW'23)
Joint speech-language model - respond directly to audio!
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
多模态情感分析——基于BERT+ResNet的多种融合方法
An open source implementation of "Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning", an all-new multi modal AI that uses just a decoder to generate both text and images
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
Implementation of "PaLM-E: An Embodied Multimodal Language Model"
Fusing Histology and Genomics via Deep Learning - IEEE TMI
Phi-4 for Mac: Locally-run Vision and Language Models for Apple Silicon
Phi-3.5 for Mac: Locally-run Vision and Language Models for Apple Silicon
Language Models Can See: Plugging Visual Controls in Text Generation
UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
[CVPR 2023] Referring Image Matting
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
End-to-end Training for Multimodal Recommendation Systems
An out-of-the-box local Web UI for DeepSeek-OCR. Built with FastAPI + Vue.js, it supports PDF/Image uploads, progress tracking, and result visualization with bounding boxes. Easily experience the power of a top-tier OCR model.
Official Code for 'RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words' (CVPR 2021)
Implementation of Zorro, Masked Multimodal Transformer, in Pytorch
[MICCAI 2024, top 11%] Official Pytorch implementation of Mammo-CLIP: A Vision Language Foundation Model to Enhance Data Efficiency and Robustness in Mammography
mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections. (EMNLP 2022)
[ICLR 2025] 🏄 OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer
[NAACL 2024] MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning
Robust multimodal image registration via keypoints
Democratization of "PaLI: A Jointly-Scaled Multilingual Language-Image Model"
Official implementation of the paper "FLIP: Cross-domain Face Anti-spoofing with Language Guidance". (ICCV 2023)
Neural Machine Translation with universal Visual Representation (ICLR 2020)
[ICLR 2024] Official repository for "Vision-by-Language for Training-Free Compositional Image Retrieval"
Robust multimodal integration method implemented in PyTorch and TensorFlow
Unofficial implementation and experiments related to Set-of-Mark (SoM) 👁️
[NeurIPS 2024] GenRL: Multimodal-foundation world models enable grounding language and video prompts into embodied domains, by turning them into sequences of latent world model states. Latent state sequences can be decoded using the decoder of the model, allowing visualization of the expected behavior, before training the agent to execute it.
The official implementation of 'Align and Attend: Multimodal Summarization with Dual Contrastive Losses' (CVPR 2023)
CLIP: Connecting Text and Image (Learning Transferable Visual Models From Natural Language Supervision)
[ICML 2024] The offical Implementation of "DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning"
A Step-by-Step Implementation of Google Veo 3 Architecture from Scratch
ICML 2025 - Impossible Videos
Improving Chest X-Ray Report Generation by Leveraging Warm-Starting
[ICRA'23] Dataset of Moving Object Detection; Official Implementation of "RGB-Event Fusion for Moving Object Detection in Autonomous Driving"
Quick Long Video Understanding [TMLR2025]
My implementation of Kosmos2.5 from the paper: "KOSMOS-2.5: A Multimodal Literate Model"
[NeurIPS'25 Spotlight] Official implementation of "JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation"
GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection (AAAI 2024)
Image Classification Testing with LLMs
Baselines for CCKS 2022 Task "Link Prediction for Multimodal Product Knowledge Graph"
An open-source multimodal large language model based on baichuan-7b.
[ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
[CVPR 2024] Unified Multi-Sensor Tracker With One Parameter Set
Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"
[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding
[IJCAI 2024] Continual Multimodal Knowledge Graph Construction
[ICLR 2025] Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
Multimodal Graph Learning: how to encode multiple multimodal neighbors with their relations into LLMs
Modality-Transferable-MER, multimodal emotion recognition model with zero-shot and few-shot abilities.
[SOICT 2024] LLM-Powered Video Search: A Comprehensive Multimedia Retrieval System
[CVPR 2021] SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.
Graph Distillation for Action Detection
[ICLRW 2024] Efficient Remote Sensing with Harmonized Transfer Learning and Modality Alignment
Pytorch Implementation of "Adaptive Co-attention Network for Named Entity Recognition in Tweets" (AAAI 2018)
[ICCV 2025] Official Implementation of "ProLearn: Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation"
[ICLR 2026] Official code repository for "⚡️VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration"
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
[ICLR 2025] MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Efficient Multimodal Foundation Model Adaptation for Recommendation
Official repository of the paper "MIRAGE: A multimodal foundation model and benchmark for comprehensive retinal OCT image analysis", published in npj Digital Medicine (2025).
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
PyTorch implementation of LIMoE
Multimodal Instruction Tuning for Llama 3
An implementation of Deep Generalized Canonical Correlation Analysis (DGCCA or Deep GCCA) with pytorch.
[IEEE TMI 2024] MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images
A Unified Framework for Benchmarking Generative Electrocardiogram-Language Models (ELMs)
Combining ViT and GPT-2 for image captioning. Trained on MS-COCO. The model was implemented mostly from scratch.
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
[NeurIPS 2024] Official PyTorch implementation of "Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives"
🐝 | From Data to Prognosis: Embedding Multimodal Oncology Data for Precision Medicine
[𝐧𝐚𝐭𝐮𝐫𝐞 𝐦𝐚𝐜𝐡𝐢𝐧𝐞 𝐢𝐧𝐭𝐞𝐥𝐥𝐢𝐠𝐞𝐧𝐜𝐞] ImmunoStruct enables multimodal deep learning for immunogenicity prediction
Multimodal Hashtag Prediction with instagram data & pytorch (2nd Place on OpenResource Hackathon 2019)
A multimodal Bert for the Chinese domain.
Toy-scale unified multimodal model experiments — encoder-free understanding & generation with Mixture-of-Transformers on MLX/Apple Silicon
State-of-the-art embedding models fine-tuned for the ecommerce domain. +67% increase in evaluation metrics vs ViT-B-16-SigLIP.
DiffBlender: Scalable and Composable Multimodal Text-to-Image Diffusion Models
[ICML'25] MedTok: Multimodal Medical Code Tokenizer
Implementation of GATO style Generalist Multimodal model capable of image, text, RL and Robotics tasks
CCKS2022 Task9 Subtask2: Product Item Alignment.
Baseline model for multimodal classification based on images and text. Text representation obtained from pretrained BERT base model and image representation obtained from VGG16 pretrained model.
Visual Instruction Tuning for Qwen2 Base Model
An end-to-end image captioning project using a CNN encoder (ResNet-50) and LSTM decoder in PyTorch. Includes vocabulary building, preprocessing, training with BLEU evaluation, and inference. Generates natural language captions for images with saved metrics, model checkpoints, and visualization outputs.
This repository contains the supporting code for multimodal sentiment analysis experiments.
[NeurIPS-2024] The offical Implementation of "Instruction-Guided Visual Masking"
Code repository for "Post-pre-training for Modality Alignment in Vision-Language Foundation Models" (CVPR2025)
[WACV 2025] Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
[NeurIPS 2025] The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning"
Easy wrapper for inserting LoRA layers in CLIP.
An efficient multi-modal instruction-following data synthesis tool and the official implementation of Oasis https://arxiv.org/abs/2503.08741.
Implementation of the "Learn No to Say Yes Better" paper.
ACM MM'23: A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal Recommendation
Finetune HunyuanImage 3.0, a 80B unified understanding and generation model
This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"
[MICCAI 2023] Fundus-Enhanced Disease-Aware Distillation Model for Retinal Disease Classification from OCT Images
Inverse DALL-E for Optical Character Recognition
A Multimodal Transformer: Fusing Clinical Notes With Structured EHR Data for Interpretable In-Hospital Mortality Prediction
[NeurIPS 2024] Official Repository of Multi-Object Hallucination in Vision-Language Models
The open source implementation of the cross attention mechanism from the paper: "JOINTLY TRAINING LARGE AUTOREGRESSIVE MULTIMODAL MODELS"
Tool to automatically generate text descriptions for images using Ollama vision models (LLaVA, Qwen3-VL, Llama Vision)
[CVPR '23] Unite and Conquer: Plug & Play Multi-Modal Synthesis using Diffusion Models
Code and models for the paper "2D3MF: Deepfake Detection using Multi Modal Middle Fusion"
[NeurIPS'23] Parts of Speech–Grounded Subspaces in Vision-Language Models
[WACV 2025] I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
Source: GitHub API · Curated · Realtime