PaddleOCR is an open-source OCR toolkit that converts PDFs and images into structured data (JSON/Markdown) for LLMs. It includes PP-OCRv6 for multilingual text recognition (100+ languages), PaddleOCR-VL for document parsing, and PP-StructureV3 for layout-aware conversion. Supports CPU, GPU, and various AI accelerators.
OPEN SOURCE, OPEN TO EVERYONE
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
✳Human-curated · Discover open source4083discovered
A little curiosity. A world of open source.
THE FIRST COLLECTIONTopic: pdf-parser清除
paddleocrPaddlePaddle
ADDEDmineruopendatalab
ADDEDMinerU is an open-source document parsing engine that converts PDF, images, DOCX, PPTX, and XLSX into structured Markdown/JSON for LLM, RAG, and agent workflows. It supports VLM+OCR dual engines, 109-language OCR, and private/offline deployment.
dolphinbytedance
ADDEDByteDance's open-source document image parsing model (ACL 2025) that turns PDFs and photographed pages into structured Markdown/JSON, covering layout, reading order, text, tables, formulas, and code.