G
goodreads_etl_pipeline
san089/goodreads_etl_pipeline
An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.
★1.5kstars
Python
MIT
Updated: 1d ago
📋 Project at a Glance
Tap to expand
What's this?A open-source Data & Infrastructure project, built with Python, focusing on airflow and airflow-dag
Who made it?Maintained by san089 team, 1.5K⭐ on GitHub, #864 out of 3133 in Data & Infrastructure
Why does it exist?As the Data & Infrastructure landscape evolved, the san089 team identified the need for better airflow solutions. goodreads_etl_pipeline was created to simplify airflow-dag workflows.
What can it do?Key use cases: apache-airflow, apache-spark, data-engineering
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/san089/goodreads_etl_pipeline
🔗 github.com/san089/goodreads_etl_pipeline
Topics
airflowairflow-dagapache-airflowapache-sparkdata-engineeringdata-engineering-pipelinedata-lakedata-migrationemr-clusteretl-frameworketl-jobetl-pipelinegoodreads-data-pipelinelivypythonredshifts3schedulersparkwarehouse