T
trafilatura
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
★6.4kstars
Python
Apache-2.0
Updated: Today
📋 Project at a Glance
Tap to expand
What's this?A open-source Data & Infrastructure project, built with Python, focusing on article-extractor and corpus-builder
Who made it?Maintained by adbar team, 6.4K⭐ on GitHub, #292 out of 3133 in Data & Infrastructure
Why does it exist?The adbar team recognized that existing article-extractor tools in Data & Infrastructure were hard to use. trafilatura was designed to make corpus-builder more accessible.
What can it do?Key use cases: corpus-tools, crawler, html-to-markdown
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/adbar/trafilatura | 官网 https://trafilatura.readthedocs.io
🔗 github.com/adbar/trafilatura | 官网 https://trafilatura.readthedocs.io
Topics
article-extractorcorpus-buildercorpus-toolscrawlerhtml-to-markdownhtml2textllmnews-aggregatornews-crawlernlpragreadabilityrss-feedscrapingteitext-cleaningtext-extractiontext-miningtext-preprocessingweb-scraping