Text-to-Image421 / 395 projects
Text-to-image model leaderboard featuring the hottest AI image generation projects on GitHub, including Stable Diffusion and open-source implementations.
Powerful node-based Stable Diffusion workflow editor. Modular pipeline for advanced image generation.
Stability AI's foundational text-to-image model. Open weights, high quality, huge ecosystem of tools.
Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest AI-driven technologies. The solution offers an industry leading WebUI, and serves as the foundation for multiple commercial products.
Image-to-Image Translation in PyTorch
Software that can generate photos from paintings, turn horses into zebras, perform style transfer, and more.
Custom photo generation with Stable Diffusion. Generate personalized images from a few photos.
Image-to-image translation with conditional adversarial nets
An easy 1-click way to create beautiful artwork on your PC using AI, with no tech knowledge. Provides a browser UI for generating images from text prompts and images. Just enter your text prompt, and see the generated image.
Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
PaddlePaddle GAN library, including lots of interesting applications like First-Order motion transfer, Wav2Lip, picture repair, image editing, photo2cartoon, image style transfer, GPEN, and so on.
Run Stable Diffusion on Mac natively
fast-stable-diffusion + DreamBooth
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
SD.Next: All-in-one WebUI for AI generative image and video creation, captioning and processing
SUPIR aims at developing Practical Algorithms for Photo-Realistic Image Restoration In the Wild. Our new online demo is also released at suppixel.ai.
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
Tiled Diffusion and VAE optimize, licensed under CC BY-NC-SA 4.0
Unofficial Implementation of DragGAN - "Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold" (DragGAN 全功能实现,在线Demo,本地部署试用,代码、模型已全部开源,支持Windows, macOS, Linux)
StableSwarmUI, A Modular Stable Diffusion Web-User-Interface, with an emphasis on making powertools easily accessible, high performance, and extensibility.
Create 🔥 videos with Stable Diffusion by exploring the latent space and morphing between text prompts
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
Simple command line tool for text to image generation using OpenAI's CLIP and Siren (Implicit neural representation network). Technique was originally created by https://twitter.com/advadnoun
Official implementations for paper: Anydoor: zero-shot object-level image customization
Unofficial implementation of Image Super-Resolution via Iterative Refinement by Pytorch
Stable diffusion for real-time music generation
Outpainting with Stable Diffusion on an infinite canvas
🪩 Create Disco Diffusion artworks in one line
Segment Anything for Stable Diffusion WebUI
min(DALL·E) is a fast, minimal port of DALL·E Mini to PyTorch
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) by way of Textual Inversion (https://arxiv.org/abs/2208.01618) for Stable Diffusion (https://arxiv.org/abs/2112.10752). Tweaks focused on training faces, objects, and styles.
[CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)
[SIGGRAPH Asia 2023] Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
Core Engine of Singing Voice Conversion & Singing Voice Clone
Deforum extension for AUTOMATIC1111's Stable Diffusion webui
MuseV: Infinite-length and High Fidelity Virtual Human Video Generation with Visual Conditioned Parallel Denoising
Kandinsky 2 — multilingual text2image latent diffusion model
FLUX, Stable Diffusion, SDXL, SD3, LoRA, Fine Tuning, DreamBooth, Training, Automatic1111, Forge WebUI, SwarmUI, DeepFake, TTS, Animation, Text To Video, Tutorials, Guides, Lectures, Courses, ComfyUI, Google Colab, RunPod, Kaggle, NoteBooks, ControlNet, TTS, Voice Cloning, AI, AI News, ML, ML News, News, Tech, Tech News, Kohya, Midjourney, RunPod
A playground to generate images from any text prompt using Stable Diffusion (past: using DALL-E Mini)
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
[IJCV2024] Exploiting Diffusion Prior for Real-World Image Super-Resolution
Just playing with getting VQGAN+CLIP running locally, rather than having to use colab.
A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN. Technique was originally created by https://twitter.com/advadnoun
Contrastive unpaired image-to-image translation, faster and lighter training than cyclegan (ECCV 2020, in PyTorch)
Lora beYond Conventional methods, Other Rank adaptation Implementations for Stable diffusion.
One-step image-to-image with Stable Diffusion turbo: sketch2image, day2night, and more
Lumina-T2X is a unified framework for Text to Any Modality Generation
Project Lyra: Open Generative 3D World Models
A sketch extractor for anime/illustration.
[NeurIPS 2025] Image editing is worth a single LoRA! 0.1% training data for fantastic image editing! Surpasses GPT-4o in ID persistence~ MoE ckpt released! Only 4GB VRAM is enough to run!
Let us democratise high-resolution generation! (CVPR 2024)
Helios: Real Real-Time Long Video Generation Model
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
Official code for "DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps" (Neurips 2022 Oral)
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
Text-to-Image generation. The repo for NeurIPS 2021 paper "CogView: Mastering Text-to-Image Generation via Transformers".
Discovering Interpretable GAN Controls [NeurIPS 2020]
[ECCV 2024] The official implementation of paper "BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion"
A suite of image and video neural tokenizers
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
Official Pytorch Implementation for "TokenFlow: Consistent Diffusion Features for Consistent Video Editing" presenting "TokenFlow" (ICLR 2024)
[ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Generate images from texts. In Russian
[ACM MM 2025] FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
Generative Adversarial Transformers
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.
Unofficial implementation of "Prompt-to-Prompt Image Editing with Cross Attention Control" with Stable Diffusion
Auto1111 extension implementing text2video diffusion models (like ModelScope or VideoCrafter) using only Auto1111 webui dependencies
A clean and readable Pytorch implementation of CycleGAN
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
Inpaint Anything extension performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
[ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
Paint by Example: Exemplar-based Image Editing with Diffusion Models
A family of diffusion models for text-to-audio generation.
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Nodes for better inpainting with ComfyUI: Fooocus inpaint model for SDXL, LaMa, MAT, and various other tools for pre-filling inpaint & outpaint areas.
Self-contained, minimalistic implementation of diffusion models with Pytorch.
PyTorch implementation for SDEdit: Image Synthesis and Editing with Stochastic Differential Equations
[ICLR24] Official implementation of the paper “MagicDrive: Street View Generation with Diverse 3D Geometry Control”
[ICCV 2023 Oral] "FateZero: Fusing Attentions for Zero-shot Text-based Video Editing"
[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Official PyTorch implementation of "VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization" (CVPR 2021)
CogView4, CogView3-Plus and CogView3(ECCV 2024)
[ECCV 2024] PowerPaint, a versatile image inpainting model that supports text-guided object inpainting, object removal, image outpainting and shape-guided object inpainting with only a single model. 一个高质量多功能的图像修补模型,可以同时支持插入物体、移除物体、图像扩展、形状可控的物体生成,只需要一个模型
Text2Room generates textured 3D meshes from a given text prompt using 2D text-to-image models (ICCV2023).
Stable Diffusion implemented from scratch in PyTorch
Stable Diffusion in NCNN with c++, supported txt2img and img2img
Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
[CVPR 2024] PIA, your Personalized Image Animator. Animate your images by text prompt, combing with Dreambooth, achieving stunning videos. PIA,你的个性化图像动画生成器,利用文本提示将图像变为奇妙的动画
[ICLR 2024] SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
Pytorch implementation of MixNMatch
Stable diffusion webui based on diffusers.
[ICLR 2026] A Training-free Iterative Framework for Long Story Visualization
Get Deinked!!
official code repo for paper "CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers"
This project is the implementation of FernRP Package, Include NPR/PBR.
Implementation of Muse: Text-to-Image Generation via Masked Generative Transformers, in Pytorch
Official PyTorch implementation for the paper High-Resolution Virtual Try-On with Misalignment and Occlusion-Handled Conditions (ECCV 2022).
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
An Open-Sourced LLM-empowered Foundation TTS System
침착한 생성모델 학습기
Research of DeepSeek Engram Architecture based on Qwen-3 and Stable Diffusion series.
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Open-source SOTA multi-image editing model
Transform your 3D texturing workflow with the power of generative AI, directly within Blender!
TensorFlow Implementation of Unsupervised Cross-Domain Image Generation
[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
Latent Point Diffusion Models for 3D Shape Generation
Unofficial implementation of YOLO-World + EfficientSAM for ComfyUI
Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
[ICCV 2023] "TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition" (Official Implementation)
Code for APDrawingGAN: Generating Artistic Portrait Drawings from Face Photos with Hierarchical GANs (CVPR 2019 Oral)
Kandinsky 5.0: A family of diffusion models for Video & Image generation
[CVPR 2024] Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models, a no lighting baked texture generative model
[ECCV 2024 Oral 🔥] Arc2Face: A Foundation Model for ID-Consistent Human Faces ------------------------ [ICCVW 2025] ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
CLIP + FFT/DWT/RGB = text to image/video
[CVPR 2024] FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
[CVPR 2021] Anycost GANs for Interactive Image Synthesis and Editing
Simple and readable code for training and sampling from diffusion models
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
AI-powered Text-to-Art Generator - Text2Art.com
Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
🤖🖌️ Generate photo-realistic textures based on source images or (soon) PBR materials. Remix, remake, mashup! Useful if you want to create variations on a theme or elaborate on an existing texture.
An easy to understand TTS / SVS / SVC framework
Personalization for Stable Diffusion via Aesthetic Gradients 🎨
HY-SOAR:Self-Correction for Optimal Alignment and Refinement in Diffusion Models
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
Create images of a given character in different poses
[ICCV 2021] Focal Frequency Loss for Image Reconstruction and Synthesis
Finetune ModelScope's Text To Video model using Diffusers 🧨
[ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
[ECCV 2024] OMG: Occlusion-friendly Personalized Multi-concept Generation In Diffusion Models
[ICML2025, NeurIPS2025 Spotlight] Sparse VideoGen 1 & 2: Accelerating Video Diffusion Transformers with Sparse Attention
AttentionGAN for Unpaired Image-to-Image Translation & Multi-Domain Image-to-Image Translation
DiffuEraser is a diffusion model for video inpainting, which performs great content completeness and temporal consistency while maintaining acceptable efficiency.
HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation
[NeurIPS 2022] Denoising Diffusion Restoration Models -- Official Code Repository
Flash Diffusion — accelerating conditional diffusion models (AAAI 2025 Oral)
Official implementation of OneDiffusion paper (CVPR 2025)
[ICCV 2023] A latent space for stochastic diffusion models
[ICLR 2025] Autoregressive Video Generation without Vector Quantization
SEAN: Image Synthesis with Semantic Region-Adaptive Normalization (CVPR 2020, Oral)
Official repository for the paper "High-Resolution Daytime Translation Without Domain Labels" (CVPR2020, Oral)
[CVPR2024] SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution
A unified framework for easy reinforcement learning in Flow-Matching models
基于Stable Diffusion优化的AI绘画模型。支持输入中英文文本,可生成多种现代艺术风格的高质量图像。| An optimized text-to-image model based on Stable Diffusion. Both Chinese and English text inputs are available to generate images. The model can generate high-quality images in several modern art styles.
Implementation of Paint-with-words with Stable Diffusion : method from eDiff-I that let you generate image from text-labeled segmentation map.
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
The official implementation of our SIGGRAPH 2020 paper Interactive Video Stylization Using Few-Shot Patch-Based Training
Segmind Distilled diffusion
[CVPR 2024 Highlight] MIGC and [TPAMI 2024] MIGC++ (Official Implementation)
[CVPR 2025] RollingDepth: Video Depth without Video Models
[TIP2026] Official codes of CCSRv2 and CCSRv1: Improving the Stability and Efficiency of Diffusion Models for Content Consistent Super-Resolution
Generative Adversarial Text to Image Synthesis / Please Star -->
Yet another PyTorch implementation of Stable Diffusion (probably easy to read)
Official code for the CVPR 2025 paper "SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models."
Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]
ComfyUI adaptation of IDM-VTON for virtual try-on.
Diffusion models of protein structure; trigonometry and attention are all you need!
Official Code for ECCV 2024 paper — One-Shot Diffusion Mimicker for Handwritten Text Generation
A Versatile and Robust SDXL-ControlNet Model for Adaptable Line Art Conditioning
Official implementation of Würstchen: Efficient Pretraining of Text-to-Image Models
T2F: text to face generation using Deep Learning
[CVPR 2022] StyleSwin: Transformer-based GAN for High-resolution Image Generation
[AAAI2024] FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
Train and block edit and save LoRAs directly inside ComfyUI for Z-image, Flux Klein, SDXL, Flux, WAN 2.2, SD 1.5
Tiled Diffusion, MultiDiffusion, Mixture of Diffusers, and optimized VAE
Support for miscellaneous image models. Currently supports: DiT, PixArt, HunYuanDiT, MiaoBi, and a few VAEs.
[NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
Official implementation for "Break-A-Scene: Extracting Multiple Concepts from a Single Image" [SIGGRAPH Asia 2023]
Generate images locally
[WACV'25 Oral] Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
[SIGGRAPH Asia 2024] ReVersion: Diffusion-Based Relation Inversion from Images
Person Image Synthesis via Denoising Diffusion Model (CVPR 2023)
PITI: Pretraining is All You Need for Image-to-Image Translation
Official PyTorch implementation of "BlendGAN: Implicitly GAN Blending for Arbitrary Stylized Face Generation" (NeurIPS 2021)
Z-Image workflow with predefined styles for high-quality image generation and a user-friendly experience. Includes pre-configured versions for GGUF and SAFETENSORS checkpoint formats.
SVG Differentiable Rendering: Generating vector graphics using neural networks. Support: text-to-SVG, Image-to-SVG, SVG Editing.
🧬 Generative modeling of regulatory DNA sequences with diffusion probabilistic models 💨
Diffusion Classifier leverages pretrained diffusion models to perform zero-shot classification without additional training
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models (LLM-grounded Diffusion: LMD, TMLR 2024)
Official Pytorch implementation of "StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance"
Freezing generator for pseudo image translation
StyleShot: A SnapShot on Any Style. 一款可以迁移任意风格到任意内容的模型,无需针对图片微调,即能生成高质量的个性风格化图片!
ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models (ICCV 2021 Oral)
[CVPR 2019 Oral] Multi-Channel Attention Selection GAN with Cascaded Semantic Guidance for Cross-View Image Translation
[ ICLR 2024 ] Official Codebase for "InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists"
A CLI tool/python module for generating images from text using guided diffusion and CLIP from OpenAI.
Official implementation of AsymFlow, pi-Flow, GMFlow
[CVPRW 2026] AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
Source: GitHub API · Curated · Realtime