Base LLMs181 / 440 projects
Base LLM leaderboard featuring the most popular foundational language models on GitHub, including Llama, Qwen, Mistral, and more.
Leading open-weight LLM series by DeepSeek. Mixture-of-Experts architecture, competitive with top proprietary models.
Stability AI's foundational text-to-image model. Open weights, high quality, huge ecosystem of tools.
Meta's open-source LLM family. State-of-the-art performance, available in 8B to 405B parameters.
A generative speech model for daily dialogue.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
Open-source multimodal LLM with vision, speech, and text. Strong performance on mobile and edge devices.
Janus-Series: Unified Multimodal Understanding and Generation Models
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
[NeurIPS 2023] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Solve Visual Understanding with Reinforced VLMs
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。
Chronos: Pretrained Models for Time Series Forecasting
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Repo for BenCao [original name: HuaTuo (华驼)], Instruction-tuning Large Language Models with Chinese Medical Knowledge. 本草(原名:华驼)模型仓库,基于中文医学知识的大语言模型指令微调
Segmentation models with pretrained backbones. Keras and TensorFlow Keras.
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
A fast multimodal LLM for real-time voice
General technology for enabling AI capabilities w/ LLMs and MLLMs
LLM training code for Databricks foundation models
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
A LITE BERT FOR SELF-SUPERVISED LEARNING OF LANGUAGE REPRESENTATIONS, 海量中文预训练ALBERT模型
TimeGPT-1: production ready pre-trained Time Series Foundation Model for forecasting and anomaly detection. Generative pretrained transformer for time series trained on over 100B data points. It's capable of accurately predicting various domains such as retail, electricity, finance, and IoT with just a few lines of code 🚀.
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.
Implementation for MatMul-free LM.
LLM Finetuning with peft
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
[CVPR 2023 Highlight] InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
dLLM: Simple Diffusion Language Modeling
A simple, performant, and scalable Jax LLM!
FinGLM: 致力于构建一个开放的、公益的、持久的金融大模型项目,利用开源开放来促进「AI+金融」。
[ICML'24] Magicoder: Empowering Code Generation with OSS-Instruct
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
[ICLR 2025] LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
Fast Multimodal LLM on Mobile Devices
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Reinforcement Learning
Visual intelligence for your home.
[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
An elegent pytorch implement of transformers
Hypernetworks that adapt LLMs for specific benchmark tasks using only textual task description as the input
Unifying 3D Mesh Generation with Language Models
Grounding DINO 1.5: IDEA Research's Most Capable Open-World Object Detection Model Series
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
GLM-TTS: Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning
中文法律LLaMA (LLaMA for Chinese legel domain)
An open-source educational chat model from ICALK, East China Normal University. 开源中英教育对话大模型。(通用基座模型,GPU部署,数据清理) 致敬: LLaMA, MOSS, BELLE, Ziya, vLLM
Pure Rust implementation of a minimal Generative Pretrained Transformer
Research of DeepSeek Engram Architecture based on Qwen-3 and Stable Diffusion series.
🤖 A PyTorch library of curated Transformer models and their composable components
We introduced a new model designed for the Code generation task. Its test accuracy on the HumanEval base dataset surpasses that of GPT-4 Turbo (April 2024) and GPT-4o.
GLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parameters
[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that can magically restore any image to perfect-4K!
Train a 1B LLM with 1T tokens from scratch by personal
[CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
Hypernetworks that update LLMs to remember factual information
PhoGPT: Generative Pre-training for Vietnamese (2023)
QiZhenGPT: An Open Source Chinese Medical Large Language Model|一个开源的中文医疗大语言模型
A real-time silent speech recognition tool.
The PyTorch implementation of Generative Pre-trained Transformers (GPTs) using Kolmogorov-Arnold Networks (KANs) for language modeling
Generative Representational Instruction Tuning
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
Official implementation of SEED-LLaMA (ICLR 2024).
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
Self-evaluating interview for AI coders
[CVPR 2025] Video Narration as Vocabulary & Video as Long Document
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
The most atomic way to train and inference a GPT in pure, dependency-free C
Train and Infer Powerful Sentence Embeddings with AnglE | 🔥 SOTA on STS and MTEB Leaderboard
(ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models (LLM-grounded Diffusion: LMD, TMLR 2024)
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
[VLDB' 25] ChatTS: LLM for Time Series Understanding and Reasoning
A multimodal Chinese LLM based on LLaMA and Alpaca (VisualCLA).
The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.
Guideline following Large Language Model for Information Extraction
First Open-Source Industry-Specific Model for Semiconductors
[ICML 2026] Let LLMs invent and evolve languages for efficient reasoning.
The Truth Is In There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
Huozhi general large model.
Visual Med-Alpaca is an open-source, multi-modal foundation model designed specifically for the biomedical domain, built on the LLaMa-7B.
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.
Joint speech-language model - respond directly to audio!
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
粋 (Sui) - A programming language optimized for LLM code generation
The resources of LucaOne, including: the model code, training scripts, embedding inference code, and trained checkpoints.
StyleLLM文风大模型:基于大语言模型的文本风格迁移项目。Text style transfer base on Large Language Model. #文字修饰 # 润色 #风格模仿
[ACL2024 Findings] Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
[ICML 2025] A pytorch implementation of the paper "TreeLoRA: Efficient Continual Learning via Layer-Wise LoRAs Guided by a Hierarchical Gradient-Similarity Tree".
A real-time streaming conversational video system that transforms text interactions into continuous, high-fidelity video responses using autoregressive diffusion.
High-performance Image Tokenizers for VAR and AR
Method for Long Context RLMs using verifiable Lambda Calculus
Text2Text Language Modeling Toolkit
电子鹦鹉 / Toy Language Model
Phi-4 for Mac: Locally-run Vision and Language Models for Apple Silicon
Phi-3.5 for Mac: Locally-run Vision and Language Models for Apple Silicon
A comprehensive survey of forging vision foundation models for autonomous driving, including challenges, methodologies, and opportunities.
Contextual Object Detection with Multimodal Large Language Models
Official Repository for PosterGen - CVPR Findings 2026
(CVPR 2025) Code of "Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models"
Foundation models based medical image analysis
Project Tapestry aims to give every nation and participant frontier AI they can call their own — uniting a global consortium to train a shared frontier model from which partners build and own sovereign models aligned to their national, socio-cultural, and industrial needs.
[ICLR 2026] An official implementation of "CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning"
An open-source LLM implementation of RecurrentGPT for interactive generation of arbitrarily long text, used for AI novel writing.
LLaMA-2 in native Go
[ICLR 2026] Official repository of "Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models"
End-to-end Training for Multimodal Recommendation Systems
[EMNLP 2025] LightThinker: Thinking Step-by-Step Compression
Code for the paper: GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
Repository for Chat LLaMA - training a LoRA for the LLaMA (1 or 2) models on HuggingFace with 8-bit or 4-bit quantization. Research only.
Code for DeSTA2.5-Audio, general-purpose LALM
Finetune LLaMA-7B with Chinese instruction datasets
[EMNLP 2025 Main] Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
An out-of-the-box local Web UI for DeepSeek-OCR. Built with FastAPI + Vue.js, it supports PDF/Image uploads, progress tracking, and result visualization with bounding boxes. Easily experience the power of a top-tier OCR model.
[𝗜𝗖𝗠𝗟 𝟮𝟬𝟮𝟲] Dispersion loss counteracts embedding condensation and improves generalization in small language models
[CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs
ZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.
A Text-guided Protein Design Framework, Nat Mach Intell 2025 (https://www.nature.com/articles/s42256-025-01011-z)
[AAAI 2026] Official codebase for "GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning".
Syllable-aware BPE tokenizer for the Amharic language (አማርኛ) – fast, accurate, trainable.
The official repo for "AceCoder: Acing Coder RL via Automated Test-Case Synthesis" [ACL25]
[NeurIPS'25 | CVPR'26] The official repo of OralGPT & MMOral Bench.
PasLLM - LLM inference engine in Object Pascal (synced from my private work repository)
Legend of Elya — N64 game with a real 819K-parameter transformer running on the VR4300 MIPS III CPU. Zelda-style dungeon, AI NPCs, byte-level inference at 60 tok/s. Built with libdragon SDK.
LoRA fine-tune Gemma 4 31B to speak caveman-mode natively. Style: github.com/JuliusBrussee/caveman
A minimal implementation of LLaVA-style VLM with interleaved image & text & video processing ability.
Efficient Infinite Context Transformers with Infini-attention Pytorch Implementation + QwenMoE Implementation + Training Script + 1M context keypass retrieval
Sequential Diffusion Language Model (SDLM) enhances pre-trained autoregressive language models by adaptively determining generation length and maintaining KV-cache compatibility, achieving high efficiency and throughput.
Text classification using LLM and LoRA.
Multi-turn empathetic dialogue model PICA.
[ACL 2024] LangBridge: Multilingual Reasoning Without Multilingual Supervision
Official Implementation of "Reasoning Language Models: A Blueprint"
VL-JEPA (Vision-Language Joint Embedding Predictive Architecture) in MLX
WangChanGLM 🐘 - The Multilingual Instruction-Following Model
Bamboo-7B Large Language Model
[CVPR 2024] Physical Property Understanding from Language-Embedded Feature Fields
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
[ECCV 2026🔥] SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
[ICLR 2024] Official repository for "Vision-by-Language for Training-Free Compositional Image Retrieval"
A Step-by-Step Implementation of Google Veo 3 Architecture from Scratch
Quick Long Video Understanding [TMLR2025]
Image Classification Testing with LLMs
An open-source multimodal large language model based on baichuan-7b.
[Interspeech 2025] DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec
Browser LLM Demo (like ChatGPT). Works locally with JavaScript and WebGPU
Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"
[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding
[ICLR 2025] Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
Multimodal Graph Learning: how to encode multiple multimodal neighbors with their relations into LLMs
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.
Abacus Code LLM.
This repo is text to speech with learnable audio encoder without alignment with transcript reference
A Unified Framework for Benchmarking Generative Electrocardiogram-Language Models (ELMs)
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Visual Instruction Tuning for Qwen2 Base Model
An efficient multi-modal instruction-following data synthesis tool and the official implementation of Oasis https://arxiv.org/abs/2503.08741.
This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"
Tool to automatically generate text descriptions for images using Ollama vision models (LLaVA, Qwen3-VL, Llama Vision)
Anthropic's Claude API. Leading large language model with strong reasoning, coding, and safety features.
Source: GitHub API · Curated · Realtime