Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
PythonPaddlePaddle/PaddleOCR

PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

85.5/100
89.1KForks: 11.3K
View on GitHubHomepage →
Loading report...

Similar Projects

ade-cli

81

The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal

Python2.4K

ParseBench

77

ParseBench - A Document Parsing Benchmark for AI Agents

Python565

MinerU

94

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python79.4K

paperless-ngx

93

A community-supported supercharged document management system: scan, index and archive all your documents

Python44.9K
Back to List