PaddleOCR is an open-source OCR toolkit that converts PDFs and images into structured data (JSON/Markdown) for LLMs. It includes PP-OCRv6 for multilingual text recognition (100+ languages), PaddleOCR-VL for document parsing, and PP-StructureV3 for layout-aware conversion. Supports CPU, GPU, and various AI accelerators.
OPEN SOURCE, OPEN TO EVERYONE
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
✳Human-curated · Discover open source4109discovered
A little curiosity. A world of open source.
THE FIRST COLLECTIONTopic: document-parsing清除
paddleocrPaddlePaddle
ADDEDdoclingdocling-project
ADDEDDocling parses diverse document formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, audio, and more, with advanced PDF understanding, export to Markdown/JSON, OCR, VLM support, and integrations with LangChain, LlamaIndex, and other AI frameworks for RAG and agentic workflows.