Text-to-Video151 / 192 projects
Text-to-video model leaderboard featuring the hottest AI video generation projects on GitHub, including video synthesis, frame interpolation, and editing.
Bring portraits to life!
Wan: Open and Advanced Large-Scale Video Generative Models
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
首家工业级全流程 AI 影视生产平台。Industry-first professional AI Agent platform for controllable film & video production. From shorts to live-action with Hollywood-standard workflows.
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
HunyuanVideo: A Systematic Framework For Large Video Generation Model
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
[CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
HunyuanVideo-1.5: A leading lightweight video generation model
Advancing Open-source World Models
[ECCV 2024] Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
[ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
A unified inference and post-training framework for accelerated video generation.
[ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
AI Agent 驱动的开源可自部署视频工作台:将小说与剧本转为角色、场景、道具资产、分镜、视频和剪映草稿,支持跨镜头一致性、多供应商与费用追踪 | Self-hosted AI video workspace for stories, storyboards and short-form video production
MAGI-1: Autoregressive Video Generation at Scale
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
[CVPR 2026] PersonaLive! : Expressive Portrait Image Animation for Live Streaming
A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models
[ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
短剧平台 AI Short Film Motion Comic Generation Platform Industrial AI Motion Comic & Video Workbench
MuseV: Infinite-length and High Fidelity Virtual Human Video Generation with Visual Conditioned Parallel Denoising
FLUX, Stable Diffusion, SDXL, SD3, LoRA, Fine Tuning, DreamBooth, Training, Automatic1111, Forge WebUI, SwarmUI, DeepFake, TTS, Animation, Text To Video, Tutorials, Guides, Lectures, Courses, ComfyUI, Google Colab, RunPod, Kaggle, NoteBooks, ControlNet, TTS, Voice Cloning, AI, AI News, ML, ML News, News, Tech, Tech News, Kohya, Midjourney, RunPod
High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
Lightweight Image Video Action Generation Inference Framework
Long Video Gen Infrastructure
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Helios: Real Real-Time Long Video Generation Model
[ICML 2026] Video generation via code
[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
Official Pytorch Implementation for "TokenFlow: Consistent Diffusion Features for Consistent Video Editing" presenting "TokenFlow" (ICLR 2024)
Official implementation of "MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling"
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pytorch
[AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text-to-Video Generation using Pose-Free Videos"
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
[TPAMI 2025🔥] MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Code for Motion Representations for Articulated Animation paper
We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-processing, conditioned on a reference image and audio.
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
Official code of Motus: A Unified Latent Action World Model
Code for SCIS-2025 Paper "UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation".
🎬 seedance2接入 开源本地 AI 短剧 & 漫剧生成工具 —— 从故事到成片一站式完成,数据不出本机,短剧工作流管理平台,高灵活度,AI真人剧,AI漫剧本地搞定。 Open-source local AI short drama maker: story → storyboard → video, fully offline, your data stays yours. 纳米流水线
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
✨ Hotshot-XL: State-of-the-art AI text-to-GIF model trained to work alongside Stable Diffusion XL
[ECCV 2024 Oral] MotionDirector: Motion Customization of Text-to-Video Diffusion Models.
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations (CVPR 2026 Findings)
[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
Official JAX implementation of MAGVIT: Masked Generative Video Transformer
Fine-Grained Open Domain Image Animation with Motion Guidance
[ICLR 2024] SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
Official repo for VideoComposer: Compositional Video Synthesis with Motion Controllability
Video generation from text&image, 1st-gen
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
[AAAI 2025] Follow-Your-Click: This repo is the official implementation of "Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts"
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
[NeurIPS 2024] A Generalizable World Model for Autonomous Driving
[ICLR 2024] Official pytorch implementation of "ControlVideo: Training-free Controllable Text-to-Video Generation"
[ACM MM 2025] Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
[ICLR2026] SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
[SIGGRAPH 2025] Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
Kandinsky 5.0: A family of diffusion models for Video & Image generation
[CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
Generate video from text using AI
CLIP + FFT/DWT/RGB = text to image/video
Implementation of Phenaki Video, which uses Mask GIT to produce text guided videos of up to 2 minutes in length, in Pytorch
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
[ICML 2024] MagicPose(also known as MagicDance): Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt.
[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
[ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control”
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
Generative World Renderer: an AI-native Renderer for Games and Virtual Worlds.
[ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
Finetune ModelScope's Text To Video model using Diffusers 🧨
[ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
A Minimalist, Batteries-included Repository for Advancing World Model Science.
[ICLR'25] SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
Implementation of MagViT2 Tokenizer in Pytorch
[SIGGRAPH ASIA 2024 TCS] AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
[ICLR 2025] Autoregressive Video Generation without Vector Quantization
Code and data for "AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks" [TMLR 2024]
A unified framework for easy reinforcement learning in Flow-Matching models
[ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
Official codes of VEnhancer: Generative Space-Time Enhancement for Video Generation
Official code, models, and data for Vista4D: Video Reshooting with 4D Point Clouds (CVPR 2026 Highlight)
Let's finetune video generation models!
Implementation of NÜWA, state of the art attention network for text to video synthesis, in Pytorch
[ECCV 2024] FreeInit: Bridging Initialization Gap in Video Diffusion Models
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
[ICCV 2025] Light-A-Video: Training-free Video Relighting via Progressive Light Fusion
FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers
Official Code for DiffMorpher: Unleashing the Capability of Diffusion Models for Image Morphing (CVPR 2024)
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions (NeurIPS 2025)
[ICCV 2025] Official Pytorch Implementation of FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait.
[AAAI-2026]FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
The pytorch implementation of our CVPR 2023 paper "Conditional Image-to-Video Generation with Latent Flow Diffusion Models"
[ECCV 2024 Oral] EDTalk - Official PyTorch Implementation
[ICML 2026] ByteDance's All-in-One Video Generation Model for Human-Object Interaction Video Generation
Codes for ID-Specific Video Customized Diffusion
[CVPR'23] MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation
AI short drama & micro-drama video generator — turns any idea into a complete short-form drama using multi-agent AI pipeline (screenwriter → storyboard → frames → video). Seedance 2 VIP, Kling 3.0 Pro, Veo 3.1, Sora 2.
Python wrapper for ByteDance's Seedance 2.5 API — Text-to-Video, Image-to-Video, realistic human faces, native 4K, consistent character generation.
[ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"
[ICCV 2025 ⭐highlight⭐] Implementation of VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
[ECCV 2026] Official PyTorch implementation of RefAlign: Representation Alignment for Reference-to-Video Generation
[ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
An open source code repository of driving world models, with training, inferencing, evaluation tools, and pretrained checkpoints.
ICCV 2025 | TesserAct: Learning 4D Embodied World Models
Official PyTorch Implementation of Unified Video Action Model (RSS 2025)
[CVPR 2022] StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2
[ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
Official Code for Epona: Autoregressive Diffusion World Model for Autonomous Driving (ICCV 2025)
[TPAMI 2025, NeurIPS 2024] Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization
LPM 1.0: Video-based Character Performance Model
[CVPR 2025 Highlight] VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation (ECCV 2024)
[NeurIPS D&B Track 2024] Official implementation of HumanVid
[CVPR-2025] The official code of HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
A real-time streaming conversational video system that transforms text interactions into continuous, high-fidelity video responses using autoregressive diffusion.
Python wrapper for ByteDance's Seedance 2.0 , Seedance 2.5 and Seedance 2 Mini API — Text-to-Video, Image-to-Video, realistic human faces, 1080p, consistent character generation.
GitHub repository for "Improving Video Generation for Multi-functional Applications"
[ICCV 2025] LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
[CVPR 2024] | LAMP: Learn a Motion Pattern for Few-Shot Based Video Generation
OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
Official repo for: Epipolar Geometry Improves Video Generation Models
Python wrapper for Black Forest Labs' FLUX 3 video API — Text-to-Video and Image-to-Video with native synchronized audio.
[ECCV 2026] Official implementation of paper LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
A Step-by-Step Implementation of Google Veo 3 Architecture from Scratch
ICML 2025 - Impossible Videos
EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
ComfyUI custom nodes for Veo 3.1 video generation — text-to-video, image-to-video, reference-to-video, extend, and 4K upscale via MuAPI
Documentation for "GeminiGen AI" API - Transform your ideas into stunning AI-generated images, videos, speech.
AI video and image creation platform. Text-to-video, video-to-video, inpainting, and real-time editing.
Source: GitHub API · Curated · Realtime