PaddleOCR is an open-source OCR toolkit that converts PDFs and images into structured data (JSON/Markdown) for LLMs. It includes PP-OCRv6 for multilingual text recognition (100+ languages), PaddleOCR-VL for document parsing, and PP-StructureV3 for layout-aware conversion. Supports CPU, GPU, and various AI accelerators.
OPEN SOURCE, OPEN TO EVERYONE
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
✳Human-curated · Discover open source4109discovered
A little curiosity. A world of open source.
THE FIRST COLLECTIONTopic: ai4science清除
paddleocrPaddlePaddle
ADDEDclio-agentiowarp
ADDEDCLIO Agent is an autonomous AI system for scientific data management at scale. Built on the IOWarp platform, it orchestrates multi-expert tiers with FastMCP support for HDF5/Parquet/CSV, and features a Bubbletea terminal UI. It integrates with major LLMs (Ollama, OpenAI, Anthropic) and serves as the intelligence layer (CEI) of the IOWarp ecosystem.
mineruopendatalab
ADDEDMinerU is an open-source document parsing engine that converts PDF, images, DOCX, PPTX, and XLSX into structured Markdown/JSON for LLM, RAG, and agent workflows. It supports VLM+OCR dual engines, 109-language OCR, and private/offline deployment.