About this project

PDF Craft is a Python library aimed at scanned books and academic or technical documents. It extracts page content and organizes body text, chapters, tables of contents, footnotes, tables, formulas and images so the result can be edited and read more easily. Main capabilities - PDF to Markdown, with text and image asset files. - PDF to EPUB, including book metadata and a table of contents. - Translation during conversion, or translation of an existing EPUB, with translation-only and bilingual output modes. - Translated PDF output: extracted text is translated and written back onto the source pages. - Reusable extraction files (`.pcex`) that can be saved once and reused later for rendering, translation or processing on another machine. Usage paths - Online app for trying the workflow in a browser. - Python with remote OCR, for developers who do not want to run OCR models locally; requires Python, Poppler, and a compatible OCR service URL and credentials. - Python with local OCR, for developers with their own NVIDIA GPU; requires Python, Poppler, CUDA, sufficient VRAM and model files. Quick start Install with `python -m pip install pdf-craft`, then configure an OCR vendor (for example a DeepSeek OCR-compatible service) and call `convert_pdf_to_markdown` or `convert_pdf_to_epub`. An `AsyncPDFCraft` facade is available for async applications; the synchronous facade should not be called from a running event loop. The README notes that example endpoints are placeholders and must be replaced with a working service. OCR and requirements The project supports DeepSeek OCR, DeepSeek OCR 2 and Unlimited OCR, each with local and remote configurations. The standard installation supports remote OCR; local OCR requires the `pdf-craft[local]` extra, a matching CUDA-enabled PyTorch build, sufficient VRAM and model files, which download from Hugging Face by default or can be loaded locally. Language support depends on the processing stage: text recognition depends on the OCR model, and translation depends on the translator and text LLM. The EPUB `lan` parameter currently offers `zh` and `en`. Documentation and feedback The README links guides for installation, OCR backends, PDF conversion and translation, EPUB translation, API reference, the `.pcex` format and troubleshooting. Issues and pull requests are welcomed; for conversion problems the project asks for package version, OCR configuration type, error logs and a shareable minimal file, with credentials and private content removed first. The project is released under the MIT license, and third-party dependencies and selected OCR models retain their own licenses.