Text Models342 / 657 projects
Text models category featuring top LLM projects on GitHub, including base models, reasoning-enhanced models, code-specific models, and embedding models, ranked by Stars.
Leading open-weight LLM series by DeepSeek. Mixture-of-Experts architecture, competitive with top proprietary models.
Meta's open-source LLM family. State-of-the-art performance, available in 8B to 405B parameters.
A generative speech model for daily dialogue.
OpenMMLab Detection Toolbox and Benchmark
Code and documentation to train Stanford's Alpaca models, and generate the data.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
Open-source multimodal LLM with vision, speech, and text. Strong performance on mobile and edge devices.
Janus-Series: Unified Multimodal Understanding and Generation Models
pix2tex: Using a ViT to convert images of equations into LaTeX code.
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
An open source implementation of CLIP.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
An open-source tool-augmented conversational language model from Fudan University
A PyTorch-based Speech Toolkit
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
OpenMMLab Semantic Segmentation Toolbox and Benchmark.
ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.
Easy-to-use image segmentation library with awesome pre-trained model zoo, supporting wide-range of practical tasks in Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, 3D Segmentation, etc.
An implementation of model parallel GPT-2 and GPT-3-style models using the mesh-tensorflow library.
Code for the paper "Jukebox: A Generative Model for Music"
Chinese version of GPT2 training code, using BERT tokenizer.
An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
Open Source Neural Machine Translation and (Large) Language Models in PyTorch
a state-of-the-art-level open visual language model | 多模态预训练模型
Google AI 2018 BERT pytorch implementation
[NeurIPS 2023] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Solve Visual Understanding with Reinforced VLMs
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。
Chronos: Pretrained Models for Time Series Forecasting
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Production First and Production Ready End-to-End Speech Recognition Toolkit
Repo for BenCao [original name: HuaTuo (华驼)], Instruction-tuning Large Language Models with Chinese Medical Knowledge. 本草(原名:华驼)模型仓库,基于中文医学知识的大语言模型指令微调
Kodezi Chronos is a debugging-first language model that achieves state-of-the-art results on SWE-bench Lite (80.33%) and 67% real-world fix accuracy, over six times better than GPT-4. Built with Adaptive Graph-Guided Retrieval and Persistent Debug Memory. Model available Q1 2026 via Kodezi OS.
Easily train your own text-generating neural network of any size and complexity on any text dataset with a few lines of code.
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
A fast multimodal LLM for real-time voice
General technology for enabling AI capabilities w/ LLMs and MLLMs
LLM training code for Databricks foundation models
Efficient AI Backbones including GhostNet, TNT and MLP, developed by Huawei Noah's Ark Lab.
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
An open-source framework for training large multimodal models.
TimeGPT-1: production ready pre-trained Time Series Foundation Model for forecasting and anomaly detection. Generative pretrained transformer for time series trained on over 100B data points. It's capable of accurately predicting various domains such as retail, electricity, finance, and IoT with just a few lines of code 🚀.
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
中文nlp解决方案(大模型、数据、模型、训练、推理)
Turbopilot is an open source large-language-model based code completion engine that runs locally on CPU
This is a Phi Family of SLMs book for getting started with Phi Models. Phi a family of open sourced AI models developed by Microsoft. Phi models are the most capable and cost-effective small language models (SLMs) available, outperforming models of the same size and next size up across a variety of language, reasoning, coding, and math benchmarks
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
Python package to easily retrain OpenAI's GPT-2 text-generating model on new texts
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.
Foundation Architecture for (M)LLMs
Home of CodeT5: Open Code LLMs for Code Understanding and Generation
Implementation for MatMul-free LM.
Rust native ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT2,...)
Chinese-LLaMA 1&2、Chinese-Falcon 基础模型;ChatFlow中文对话模型;中文OpenLLaMA模型;NLP预训练/指令微调数据集
GPT2 for Chinese chitchat/用于中文闲聊的GPT2模型
LLM Finetuning with peft
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
Lingvo
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
dLLM: Simple Diffusion Language Modeling
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
A simple, performant, and scalable Jax LLM!
FinGLM: 致力于构建一个开放的、公益的、持久的金融大模型项目,利用开源开放来促进「AI+金融」。
Lumina-T2X is a unified framework for Text to Any Modality Generation
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
[ICML'24] Magicoder: Empowering Code Generation with OSS-Instruct
Russian GPT3 models.
[ACL 2023] One Embedder, Any Task: Instruction-Finetuned Text Embeddings
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
A Large-scale Chinese Short-Text Conversation Dataset and Chinese pre-training dialog models
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
[NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
[ICLR 2025] LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
中文长文本分类、短句子分类、多标签分类、两句子相似度(Chinese Text Classification of Keras NLP, multi-label classify, or sentence classify, long or short),字词句向量嵌入层(embeddings)和网络层(graph)构建基类,FastText,TextCNN,CharCNN,TextRNN, RCNN, DCNN, DPCNN, VDCNN, CRNN, Bert, Xlnet, Albert, Attention, DeepMoji, HAN, 胶囊网络-CapsuleNet, Transformer-encode, Seq2seq, SWEM, LEAM, TextGCN
Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.
[CVPR 2023] OneFormer: One Transformer to Rule Universal Image Segmentation
中文对话0.2B小模型(ChatLM-Chinese-0.2B),开源所有数据集来源、数据清洗、tokenizer训练、模型预训练、SFT指令微调、RLHF优化等流程的全部代码。支持下游任务sft微调,给出三元组信息抽取微调示例。
Generate images from texts. In Russian
DELTA is a deep learning based natural language and speech processing platform. LF AI & DATA Projects: https://lfaidata.foundation/projects/delta/
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
Fast Multimodal LLM on Mobile Devices
Toolkit for efficient experimentation with Speech Recognition, Text2Speech and NLP
Multimodal-GPT
"Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement" (ICCV 2023 Top-10 Cited 🏆) & (NTIRE 2024 Runner-Up 🏆) & (NTIRE 2025 Winner 🏆) & (NTIRE 2026 Winner 🏆)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Reinforcement Learning
Visual intelligence for your home.
Korean BERT pre-trained cased (KoBERT)
This repository is the official implementation of Disentangling Writer and Character Styles for Handwriting Generation (CVPR 2023)
An Open-sourced Knowledgable Large Language Model Framework.
[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
An elegent pytorch implement of transformers
The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
Hypernetworks that adapt LLMs for specific benchmark tasks using only textual task description as the input
Transformers4Rec is a flexible and efficient library for sequential and session-based recommendation and works with PyTorch.
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
Implementation of various self-attention mechanisms focused on computer vision. Ongoing repository.
Unifying 3D Mesh Generation with Language Models
[NeurIPS 2020] Official code for the paper "DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation". Includes a PyTorch library for deep learning with SVG data.
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
GLM-TTS: Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning
中文法律LLaMA (LLaMA for Chinese legel domain)
[CVPR2022] Geometric Transformer for Fast and Robust Point Cloud Registration
Official Code Release for [SIGGRAPH 2025] RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination
an open-source implementation of sequence-to-sequence based speech processing engine
Pre-training of Deep Bidirectional Transformers for Language Understanding: pre-train TextCNN
An open-source educational chat model from ICALK, East China Normal University. 开源中英教育对话大模型。(通用基座模型,GPU部署,数据清理) 致敬: LLaMA, MOSS, BELLE, Ziya, vLLM
official code repo for paper "CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers"
SOTA Semantic Segmentation Models in PyTorch
Pure Rust implementation of a minimal Generative Pretrained Transformer
A parallel implementation of "graph2vec: Learning Distributed Representations of Graphs" (MLGWorkshop 2017).
An Open-Sourced LLM-empowered Foundation TTS System
TextGAN is a PyTorch framework for Generative Adversarial Networks (GANs) based text generation models.
Implementation of Transformer model (originally from Attention is All You Need) applied to Time Series.
[NeurIPS 2021] You Only Look at One Sequence
Research of DeepSeek Engram Architecture based on Qwen-3 and Stable Diffusion series.
🤖 A PyTorch library of curated Transformer models and their composable components
A Change Detection Repo Standing on the Shoulders of Giants
Large-scale pretrained models for goal-directed dialog
Late Interaction Models Training & Retrieval
SGPT: GPT Sentence Embeddings for Semantic Search
Official Pytorch Code for "Medical Transformer: Gated Axial-Attention for Medical Image Segmentation" - MICCAI 2021
We introduced a new model designed for the Code generation task. Its test accuracy on the HumanEval base dataset surpasses that of GPT-4 Turbo (April 2024) and GPT-4o.
GLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parameters
[ICLR2025] Kolmogorov-Arnold Transformer
[ICLR'23] DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Connectionist Temporal Classification (CTC) decoding algorithms: best path, beam search, lexicon search, prefix search, and token passing. Implemented in Python.
[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that can magically restore any image to perfect-4K!
Train a 1B LLM with 1T tokens from scratch by personal
Keras implementation of BERT with pre-trained weights
A Keras TensorFlow 2.0 implementation of BERT, ALBERT and adapter-BERT.
Hypernetworks that update LLMs to remember factual information
PhoGPT: Generative Pre-training for Vietnamese (2023)
QiZhenGPT: An Open Source Chinese Medical Large Language Model|一个开源的中文医疗大语言模型
CKIP Transformers
A real-time silent speech recognition tool.
[ECCV 2024] InstructIR: High-Quality Image Restoration Following Human Instructions https://huggingface.co/spaces/marcosv/InstructIR
The PyTorch implementation of Generative Pre-trained Transformers (GPTs) using Kolmogorov-Arnold Networks (KANs) for language modeling
Revisiting Pre-trained Models for Chinese Natural Language Processing (MacBERT)
[ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
Generative Representational Instruction Tuning
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
[NAACL 2021] QAGNN: Question Answering using Language Models and Knowledge Graphs 🤖
Cornucopia: Open-source commercial Chinese financial LLM series with an efficient lightweight training framework (Pretraining, SFT, RLHF, Quantize).
Pretrained ELECTRA Model for Korean
Open-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
Next-generation Video instance recognition framework on top of Detectron2 which supports InstMove (CVPR 2023), SeqFormer(ECCV Oral), and IDOL(ECCV Oral))
BERTweet: A pre-trained language model for English Tweets (EMNLP-2020)
Official code for Conformer: Local Features Coupling Global Representations for Visual Recognition
Self-evaluating interview for AI coders
End-to-end ASR/LM implementation with PyTorch
Multi-label Classification with BERT; Fine Grained Sentiment Analysis from AI challenger
[CVPR 2025] Video Narration as Vocabulary & Video as Long Document
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
[CVPR 2023] Official code release of our paper "BiFormer: Vision Transformer with Bi-Level Routing Attention"
Connectionist Temporal Classification (CTC) decoder with dictionary and language model.
The most atomic way to train and inference a GPT in pure, dependency-free C
Train and Infer Powerful Sentence Embeddings with AnglE | 🔥 SOTA on STS and MTEB Leaderboard
Diffusion models of protein structure; trigonometry and attention are all you need!
Implementation of MusicLM, a text to music model published by Google Research, with a few modifications.
[ICLR 2024] Lemur: Open Foundation Models for Language Agents
(ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Meshed-Memory Transformer for Image Captioning. CVPR 2020
BERT for Multitask Learning
[CVPR 2022] StyleSwin: Transformer-based GAN for High-resolution Image Generation
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
Official Pytorch implementation of "OmniNet: A unified architecture for multi-modal multi-task learning" | Authors: Subhojeet Pramanik, Priyanka Agrawal, Aman Hussain
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions (NeurIPS 2025)
VibeVoiceFusion is a full-stack, multi-speaker voice generation web system featuring LoRA fine-tuning, batch generation, and VRAM optimization. Based on Microsoft's VibeVoice (AR + diffusion architecture)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models (LLM-grounded Diffusion: LMD, TMLR 2024)
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
[VLDB' 25] ChatTS: LLM for Time Series Understanding and Reasoning
Korean BART
A multimodal Chinese LLM based on LLaMA and Alpaca (VisualCLA).
[ACL 2022] LinkBERT: A Knowledgeable Language Model 😎 Pretrained with Document Links
The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.
Guideline following Large Language Model for Information Extraction
First Open-Source Industry-Specific Model for Semiconductors
🤗 ParsBERT: Transformer-based Model for Persian Language Understanding
[ICML 2026] Let LLMs invent and evolve languages for efficient reasoning.
The Truth Is In There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
Huozhi general large model.
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
Midi event transformer for symbolic music generation
Joint speech-language model - respond directly to audio!
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
Source: GitHub API · Curated · Realtime