RAG Frameworks330 / 382 projects
RAG framework leaderboard featuring the hottest Retrieval-Augmented Generation framework projects on GitHub.
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.
LlamaIndex is the leading document agent and OCR platform
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
Opiniated RAG for integrating GenAI in your apps 🧠 Focus on your product rather than the RAG. Easy integration in existing products with customisation! Any LLM: GPT4, Groq, Llama. Any Vectorstore: PGVector, Faiss. Any Files. Anyway you want.
[EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation"
Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
A modular graph-based Retrieval-Augmented Generation (RAG) system
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and visual AI workflow orchestration, letting you easily develop and deploy complex question-answering systems without the need for extensive setup or configuration.
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search, and conversational systems.
An open-source RAG-based tool for chatting with your documents.
open-source agentic AI data assistant for the next generation of AI + Data products.
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Private AI platform for agents, assistants and enterprise search. Built-in Agent Builder, Deep research, Document analysis, Multi-model support, and API connectivity for agents.
Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.
A lightweight, lightning-fast, in-process vector database
A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval
A vector index built on TurboQuant, written in Rust with Python bindings
💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified model management, Evaluation, SFT, Dataset Management, Enterprise-level System Management, Observability and more.
🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
High accuracy RAG for answering questions from scientific documents with citations
Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.
SoTA production-ready AI retrieval system. Agentic Retrieval-Augmented Generation (RAG) with a RESTful API.
The AI search platform
Postgres with GPUs for ML/AI apps.
Build ChatGPT over your data, all with natural language
Open-source context retrieval layer for AI agents
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
A suite of tools to develop RAG, semantic search, and other AI applications more easily with PostgreSQL
HelixDB is an OLTP graph-vector database built in Rust on Object Storage.
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
Superduper: End-to-end framework for building custom AI applications and agents.
Neo4j graph construction from unstructured data using LLMs
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense vector, sparse vector, tensor (multi-vector), and full-text.
🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines
The easiest way to use Agentic RAG in any enterprise
RAG (Retrieval Augmented Generation) Framework for building modular, open source applications for production by TrueFoundry
Everything you need to know to build your own RAG application
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
A simple, easy-to-hack GraphRAG implementation
A modular Agentic RAG built with LangGraph — learn Retrieval-Augmented Generation Agents in minutes.
Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).
Interact with your SQL database, Natural Language to SQL using LLMs
企业级 Agentic RAG 智能体 - 全链路覆盖文档解析、多路检索、意图识别、问题重写、会话记忆、MCP 工具调用与深度思考。面向真实业务场景,从 0 到 1 完整工程实现。
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
OpenKB: Open LLM Knowledge Base
[KDD'2026] "VideoRAG: Chat with Your Videos"
RAG Web UI is an intelligent dialogue system based on RAG (Retrieval-Augmented Generation) technology.
Hypergraph is more powerful. Transform unstructured text into structured knowledge with LLMs. Graphs, hypergraphs, and spatio-temporal extractions — with one command.
AI Search & RAG Without Moving Your Data. Get instant answers from your company's knowledge across 100+ apps while keeping data secure. Deploy in minutes, not months.
Jupyter Notebooks to help you get hands-on with Pinecone vector databases
pingcap/autoflow is a Graph RAG based and conversational knowledge base tool built with TiDB Serverless Vector Storage. Demo: https://tidb.ai
The AI-Native Search Database. Best for agent storage, it unifies vector, text, structured, and semi-structured data into a single engine. This all-in-one database makes agents smarter, easier to run, and more stable.
All-in-one platform for search, recommendations, RAG, and analytics offered via API
This repository contains various advanced techniques for Retrieval-Augmented Generation (RAG) systems.
The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
动手学Ollama,CPU玩转大模型部署,在线阅读地址:https://datawhalechina.github.io/handy-ollama/
PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation
Distributed vector search for AI-native applications
A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
Knowhere extracts, parses, and outputs structured chunks ready for AI Agents and RAG.
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
Empowering RAG with a memory-based data interface for all-purpose applications!
⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡
Research project. A Memory solution for users, teams, and applications.
NestJS Helper + AI Chatbot Development
EdegQuake 🌋 High-performance GraphRAG inspired from LightRag written in Rust; Transform documents into intelligent knowledge graphs for superior retrieval and generation
The open-source RAG platform: built-in citations, deep research, 22+ file formats, partitions, MCP server, and more.
ChatWiki 微信公众号的AI知识库工作流Agent平台,RAG大模型AI客服机器人,致力于成为垂直领域的coze、n8n。
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
[ACL2026] "MiniRAG: Making RAG Simpler with Small and Open-Sourced Language Models"
A community-driven collection of RAG (Retrieval-Augmented Generation) frameworks, projects, and resources. Contribute and explore the evolving RAG ecosystem.
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Fast BM25 search in Python, powered by Numpy and Numba
The official implementation of RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
This repository provides an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering. It uses sophisticated graph based algorithm to handle the tasks.
WikiChat is an improved RAG. It stops the hallucination of large language models by retrieving data from a corpus.
This repository contains examples for customers to get started using the Amazon Bedrock Service. This contains examples for all available foundational models
Prismer Cloud
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
Embeddable vector database for Go with Chroma-like interface and zero third-party dependencies. In-memory with optional persistence.
A @ClickHouse fork that supports high-performance vector search and full-text search.
Retrieval Augmented Generation (RAG) framework and context engine powered by Pinecone
Lite & Super-fast re-ranking for your search & retrieval pipelines. Supports SoTA Listwise and Pairwise reranking based on LLMs and cross-encoders and more. Created by Prithivi Da, open for PRs & Collaborations.
Build fast and accurate GenAI apps with GraphRAG SDK at scale 🌟
Parsing-free RAG supported by VLMs
Running Llama 2 and other Open-Source LLMs on CPU Inference Locally for Document Q&A
Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs
RAG-Fusion: multi-query generation + Reciprocal Rank Fusion for better retrieval-augmented generation. Includes evaluation harness with NFCorpus/BEIR.
Empower Large Language Models (LLM) using Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) for knowledge intensive tasks
🧠 纯原生 Python 实现的 RAG 框架 | FAISS + BM25 混合检索 | 支持 Ollama / SiliconFlow | 适合新手入门学习
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
📚 Process PDFs, Word documents and more with spaCy
RAG Time: A 5-week Learning Journey to Mastering RAG
ID-based RAG FastAPI: Integration with Langchain and PostgreSQL/pgvector
Epsilla is a high performance Vector Database Management System
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.
Build your own serverless AI Chat with Retrieval-Augmented-Generation using LangChain.js, TypeScript and Azure
An LLM-powered advanced RAG pipeline built from scratch
RAG for Local LLM, chat with PDF/doc/txt files, ChatPDF. 纯原生实现RAG功能,基于本地LLM、embedding模型、reranker模型实现,支持GraphRAG,无须安装任何第三方agent库。
Use late-interaction multi-modal models such as ColPali in just a few lines of code.
Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes
Full-text and semantic search on any Postgres
End-to-end RAG system design, evaluation, and optimization. 极客时间RAG训练营,RAG 10大组件全面拆解,4个实操项目吃透 RAG 全流程。RAG的落地,往往是面向业务做RAG,而不是反过来面向RAG做业务。这就是为什么我们需要针对不同场景、不同问题做针对性的调整、优化和定制化。魔鬼全在细节中,我们深入进去探究。
Train Models Contrastively in Pytorch
This repository contains the implementation of AutoSchemaKG, a novel framework for automatic knowledge graph construction that combines schema generation via conceptualization.
Single-file memory layer for AI agents, sub mili-second RAG on Apple Silicon. Metal Optimized On-Device. No Server. No API. One File. Pure Swift
Framework for enhancing LLMs for RAG tasks using fine-tuning.
Fast, streaming indexing, query, and agentic LLM applications in Rust
Ingest files for retrieval augmented generation (RAG) with open-source Large Language Models (LLMs), all without 3rd parties or sensitive data leaving your network.
This NVIDIA RAG blueprint serves as a reference solution for a foundational Retrieval Augmented Generation (RAG) pipeline.
🔥 Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation 🔥. Our toolkit integrates 40 pre-retrieved benchmark datasets and supports 7+ retrieval techniques, 24+ state-of-the-art Reranking models, and multiple RAG methods.
An Educational Project (step by step) to teach how to build a production-ready app for RAG application.
Chat with multiple PDFs locally
RAGLight is a modular framework for Retrieval-Augmented Generation (RAG). It makes it easy to plug in different LLMs, embeddings, and vector stores, and now includes seamless MCP integration to connect external tools and data sources.
The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.
An application for running LLMs locally on your device, with your documents, facilitating detailed citations in generated responses.
Easy-to-Use RAG Framework; CCF AIOps International Challenge 2024 Top3 Solution; CCF AIOps 国际挑战赛 2024 季军方案
The all-in-one RWKV runtime box with embed, RAG, AI agents, and more.
This repository provides programs to build Retrieval Augmented Generation (RAG) code for Generative AI with LlamaIndex, Deep Lake, and Pinecone leveraging the power of OpenAI and Hugging Face models for generation and evaluation.
GraphRAG4OpenWebUI integrates Microsoft's GraphRAG technology into Open WebUI, providing a versatile information retrieval API. It combines local, global, and web searches for advanced Q&A systems and search engines. This tool simplifies graph-based retrieval integration in open web environments.
A NodeJS RAG framework to easily work with LLMs and embeddings
Fullstack "Chat with your PDFs" RAG (Retrieval Augmented Generation) app built fully on Cloudflare
Context layer platform in your infrastructure
[EMNLP'25 findings] This is the official repo for the paper, HiRAG: Retrieval-Augmented Generation with Hierarchical Knowledge.
Summarize and query from a lot of heterogeneous documents. Any LLM provider, any filetype, advanced RAG, advanced summaries, scriptable, etc
A full-stack demo showcasing a local RAG (Retrieval Augmented Generation) pipeline to chat with your PDFs.
The Supabase of AI era. A modular, open-source backend for building AI-native software — designed for knowledge, not static data.
RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics etc. Built-in image/audio generation with dynamic loading generators. Live chat deployment. Built-in block based graphical language. Prompt versioning and much more...
Dataset and benchmark for RAG on company internal documents.
RAG-GPT, leveraging LLM and RAG technology, learns from user-customized knowledge bases to provide contextually relevant answers for a wide range of queries, ensuring rapid and accurate information retrieval.
Advanced Retrieval-Augmented Generation (RAG) through practical notebooks, using the power of the Langchain, OpenAI GPTs ,META LLAMA3 ,Agents.
A multithreaded 🕸️ web crawler that recursively crawls a website and creates a 🔽 markdown file for each page, designed for LLM RAG
MindSQL: A Python Text-to-SQL RAG Library simplifying database interactions. Seamlessly integrates with PostgreSQL, MySQL, SQLite, Snowflake, and BigQuery. Powered by GPT-4 and Llama 2, it enables natural language queries. Supports ChromaDB and Faiss for context-aware responses.
Official Python SDK for the Pinecone vector database
RAG (Retrieval-augmented generation) ChatBot that provides answers based on contextual information extracted from a collection of Markdown files.
Open Source LLM toolkit to build trustworthy LLM applications. TigerArmor (AI safety), TigerRAG (embedding, RAG), TigerTune (fine-tuning)
Super performant RAG pipelines for AI apps. Summarization, Retrieve/Rerank and Code Interpreters in one simple API.
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models, multimodal agents, speech, vector DB, and RAG.
Adaptive Chunking: automatically select the best chunking method per document for RAG. Accepted at LREC 2026.
Program that lets you ask questions about your documents, audio, and video files.
A LLM RAG system runs on your laptop. 大模型检索增强生成系统,可以轻松部署在笔记本电脑上,实现本地知识库智能问答。
Self-hosted, OpenAI-compatible AI gateway for private RAG, natural-language data access, and tool-calling agents.
Metrics to evaluate the quality of responses of your Retrieval Augmented Generation (RAG) applications.
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.
The RAG Experiment Accelerator is a versatile tool designed to expedite and facilitate the process of conducting experiments and evaluations using Azure Cognitive Search and RAG pattern.
VeritasGraph — open-source Knowledge Graph & GraphRAG framework on GitHub. Build multi-hop reasoning, ontology-aware retrieval, and verifiable attribution over your own data. Nodes, edges, RDF, linked-data — runs locally or in the cloud.
qKnow is an open-source agent platform for enterprise knowledge intelligence and industry AI. It provides knowledge graphs, RAG, and bot building, encapsulating industry AI into standardized modules for rapid deployment. Supporting unified integration of documents and data, it helps enterprises quickly build knowledge hubs and AI solutions.
A modern desktop application for exploring, managing, and analyzing vector databases
Local-first, zero-key semantic code search for large and custom codebases — hybrid vector + keyword retrieval with symbol-aware chunking. Usable as a CLI, Python library, REST API, or web UI.
Cut LLM costs by up to 80% and unlock sub-millisecond responses with intelligent semantic caching.A drop-in, provider-agnostic LLM proxy written in Go with sub-millisecond response
High performance embedded vector database
The Kubernetes operator for K8ssandra
Graph RAG with pure vector search, achieving SOTA performance in multi-hop reasoning scenarios.
Python client for Antarys vector database, optimized for large-scale vector operations with built-in caching, parallel processing, and dimension validation.
Open-source protocol suite standardizing LLM, Vector, Graph, and Embedding infrastructure across LangChain, LlamaIndex, AutoGen, CrewAI, Semantic Kernel, and MCP. 3,330+ conformance tests. One protocol. Any framework. Any provider.
A lightweight, production-ready RAG (Retrieval Augmented Generation) library in Go.
Self-hosted AI RAG + MCP Platform
A simple, easy-to-hack Vector Database
RAFT, or Retrieval-Augmented Fine-Tuning, is a method comprising of a fine-tuning and a RAG-based retrieval phase. It is particularly suited for the creation of agents that realistically emulate a specific human target.
RapidFire AI: Rapid AI Customization from RAG to Fine-Tuning
A framework for fine-tuning retrieval-augmented generation (RAG) systems.
Official code of the ACL 2025 paper "SimGRAG: Leveraging Similar Subgraphs for Knowledge Graphs Driven Retrieval-Augmented Generation"
Build Enterprise RAG (Retriver Augmented Generation) Pipelines to tackle various Generative AI use cases with LLM's by simply plugging componants like Lego pieces.
A diverse, simple, and secure all-in-one LLMOps platform
Local AI-powered document search and editing with first-in-class hybrid retrieval, LLM answers, WebUI, REST API and MCP support for AI clients.
Self-hosted RAG platform for AI document search across GitHub, Notion, Google Drive, local files, and web sources with citations.
Opinionated sample on how to build/deploy a RAG web app on AWS powered by Amazon Bedrock and PGVector (on Amazon RDS)
Self-hostable RAG platform - document ingestion, embedding, and vector search behind a simple REST API
OasisDB: A minimal and lightweight vector database
A resource-efficient C++ vector index engine built for low-RAM production workloads
Selfhost modern LLM stacks. Run the whole fleet from your terminal
Microsoft Foundry Local - Workshop - Build AI apps locally with Foundry Local - RAG, agents, and multi-agent workflows in Python, JS, and C#. No cloud required.
面向 PRD、业务规则、SOP、流程文档和产品截图的证据型 RAG 知识库。
Skardi is an agent data plane that gives AI agents data autonomy.
Handling 10M+ docs using RAG with zero hallucinatons
Official Implementation for the ACL 2026 paper "SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression"
Enterprise RAG ecosystem managing 32000+ semantic chunks. Features hybrid parsing (LlamaParse/PyMuPDF) and 256-dim MRL embeddings for 512MB RAM environments
VideoDB Python SDK
RAG-powered documentation assistant that converts, processes, and provides semantic search capabilities for Odoo's technical documentation. Supports multiple Odoo versions with an interactive chat interface powered by LLM models.
High-performance vector search for Browser, Node, and Edge
A local LLM chatbot with RAG for PDF input files
The open, machine-readable standard for the Somali language — orthography, grammar, terminology, translation, and AI/benchmark resources, as a versioned, citable RFC-style catalog.
Vectorless, Reasoning-Based Retrieval-Augmented Generation (RAG)
Source: GitHub API · Curated · Realtime