P
PaddleOCR
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
★87.2kstars
Python
Apache-2.0
Updated: Today
📋 Project at a Glance
Tap to expand
What's this?Models project leveraging ai4science and chineseocr, built with Python, open-source
Who made it?Maintained by PaddlePaddle team, 87.2K⭐ on GitHub, #6 out of 3201 in Models
Why does it exist?With the growing demand for ai4science and chineseocr in Models, PaddlePaddle created PaddleOCR as a streamlined solution.
What can it do?Key use cases: document-parsing, document-translation, kie
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/PaddlePaddle/PaddleOCR | 官网 https://www.paddleocr.com
🔗 github.com/PaddlePaddle/PaddleOCR | 官网 https://www.paddleocr.com
Topics
ai4sciencechineseocrdocument-parsingdocument-translationkieocrpaddleocr-vlpdf-extractor-ragpdf-parserpdf2markdownpp-ocrpp-structurerag