Data Pipelines & ETL74 / 1052 projects
Data pipeline leaderboard featuring the hottest AI data processing and ETL pipeline projects on GitHub for data preparation.
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
A high-performance observability data pipeline.
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
An orchestration platform for the development, production, and observation of data assets.
Always know what to expect from your data.
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
The Open Source Feature Store for AI/ML
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Where data access meets operational intelligence
Build data pipelines, the easy way 🛠️
pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).
A system for agentic LLM-powered data processing and ETL
The fastest ⚡️ way to build data pipelines. Develop iteratively, deploy anywhere. ☁️
Compare tables within or across databases
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Open-source inference server and production cluster for all the models your agent needs.
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
🔮 Instill Core is a full-stack AI infrastructure tool for data, model and pipeline orchestration, designed to streamline every aspect of building versatile AI-first applications
A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
Feathr – A scalable, unified data and AI engineering platform for enterprise
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Scalable data pre processing and curation toolkit for LLMs
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.
The open document intelligence platform for builders and hackers - DMS for the agentic world
A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.
Open-source ETL and ELT tool on DuckDB. Low-code visual data pipelines or SQL: 364 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents.
A comprehensive guide to building a modern data warehouse with SQL Server, including ETL processes, data modeling, and analytics.
Smarter data pipelines for audio.
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.
Community-driven, simple, yet powerful framework for fast, cost-effective distributed Compute over Data.
A scalable general purpose micro-framework for defining dataflows. THIS REPOSITORY HAS BEEN MOVED TO www.github.com/dagworks-inc/hamilton
Google Cloud Dataflow provides a simple, powerful model for building both batch and streaming parallel data processing pipelines.
Python framework for building efficient data pipelines. It promotes modularity and collaboration, enabling the creation of complex pipelines from simple, reusable components.
All-in-one text de-duplication
Optimus is an easy-to-use, reliable, and performant workflow orchestrator for data transformation, data modeling, pipelines, and data quality management.
Know your data better!Datavines is Next-gen Data Observability Platform, support metadata manage and data quality.
The open source, end-to-end computer vision platform. Label, build, train, tune, deploy and automate in a unified platform that runs on any cloud and on-premises.
Advanced and Fast Data Transformation in R
🍁 Sycamore is an LLM-powered search and analytics platform for unstructured data.
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
Augmentation pipeline for rendering synthetic paper printing, faxing, scanning and copy machine processes
Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.
Lineage metadata API, artifacts streams, sandbox, API, and spaces for Polyaxon
Production-ready data processing made easy and shareable
FeatHub - A stream-batch unified feature store for real-time machine learning
Python based Open Source ETL tools for file crawling, document processing (text extraction, OCR), content analysis (Entity Extraction & Named Entity Recognition) & data enrichment (annotation) pipelines & ingestor to Solr or Elastic search index & linked data graph database
Radient turns many data types (not just text) into vectors for similarity search, RAG, regression analysis, and more.
Production-grade PostgreSQL extension to execute arbitrary SQL in background worker processes — with async execution, autonomous transactions, cookie-protected handles, cancellation, progress reporting, and observability.
Construct a modern data stack and orchestration the workflows to create high quality data for analytics and ML applications.
Transform your pythonic research to an artifact that engineers can deploy easily.
Fast, minimal, and ergonomic zero-copy Apache Arrow tabular data implementation in Rust, with PyO3, Python bindings, NdArray, XArray, shared-memory and low-latency SIMD for data-engineering and columnar workloads. Bridge to Polars and Arrow-rs over FFI effortlessly.
Data trust for Python. Validate, clean, and profile DataFrames before everything else.
Prosto is a data processing toolkit radically changing how data is processed by heavily relying on functions and operations with functions - an alternative to map-reduce and join-groupby
🦆 Batch data pipeline with Airflow, DuckDB, Delta Lake, Trino, MinIO, and Metabase. Full observability and data quality.
A schema-aware Scala library for data transformation
Multi-modal speech separation task data generation script on LRS3 data set.
Prism is the easiest way to develop, orchestrate, and execute data pipelines in Python.
Data streaming runtime focused on performance, consistency, and extensibility. Write plugins in Rust or WASM and process data with data guarantees.
Beneath is a serverless real-time data platform ⚡️
Materials for the Deploy and Monitor ML Pipelines with Python, Docker and GitHub Actions workshop at the PyData NYC 2024 conference
PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it
mloda.ai - Open Data Access for AI and ML. Plugin-based. Traceable. Framework-agnostic.
Data transformation framework for LinkML data models
Ontology-driven Linked Data processor and server for SPARQL backends. Apache License.
A Declarative framework for Building, Maintaining, and Analyzing Graph Data
Synthetic test data that hits the numbers you declare, exactly. Multi-table with verified foreign-key integrity, deterministic, no model in the data path. Python + MCP server.
Tools for filtering and cleaning parallel and monolingual corpora for machine translation and other natural language processing tasks.
Source: GitHub API · Curated · Realtime