About this project

Unstructured is an open-source pre-processing library for unstructured data. It ingests and partitions images and text documents — PDFs, HTML, Word documents, emails, and many other formats — and turns them into structured outputs suitable for language model workflows. Core capabilities - A `partition` function auto-detects file type and routes to format-specific partitioners (e.g. `partition_pdf`, `partition_text`). - Output is a list of document elements that can be serialized to strings or structured formats. - Modular functions and connectors form a system for data ingestion and pre-processing, intended to be adaptable across platforms. - The README also describes an MCP server offering (Unstructured Transform) that exposes document processing to agents, handling 60+ file types with parsing, enrichment, chunking and embedding inside an agent session. Installation and usage - Install via PyPI: `pip install unstructured` for plain text, HTML, XML, JSON and emails; `pip install "unstructured[all-docs]"` for all document types; or targeted extras such as `unstructured[docx,pptx]`. - Optional system dependencies include libmagic-dev (file type detection), poppler-utils (images/PDFs), tesseract-ocr (OCR, with tesseract-lang for extra languages), and libreoffice (MS Office docs); pandoc is bundled via pypandoc-binary. - Docker images are published for x86_64 and Apple silicon; images are tagged by commit hash, version, and `latest`, and can also be built locally with `make docker-build`. - Local development uses `uv` for dependency management; `make install` runs `uv sync --locked --all-extras --all-groups`. Optional pre-commit hooks and `make check`/`make tidy` support formatting and linting. Documentation and community The project points to docs.unstructured.io for quick start, core functionality, connectors, concepts and integrations. It provides a security policy, a bug report template (with `scripts/collect_env.py` for environment info), and links to a company website, full API documentation, and a separate batch-processing repository (unstructured-ingest). Telemetry The README states that lightweight analytics are sent by default: a library-load ping on import and a best-effort local attempt per top-level public partition call. Runtime events reportedly contain package version, normalized platform/Python/architecture values, fixed-enum partition characteristics, and aggregate final-element counts as URL query parameters, with no request body and no document content, filenames, paths, URLs, MIME values, exception details, credentials, or persistent identifiers. Delivery is described as non-blocking, without redirects, retries, response-body download, or queue, and not consulting proxy or netrc settings. Opting out is possible by setting `DO_NOT_TRACK` or `SCARF_NO_ANALYTICS` to any non-empty value before importing or partitioning.