About this project
Unstructured is an open-source pre-processing library for unstructured data. It ingests and partitions images and text documents — PDFs, HTML, Word documents, emails, and many other formats — and turns them into structured outputs suitable for language model workflows.
Core capabilities
- A `partition` function auto-detects file type and routes to format-specific partitioners (e.g. `partition_pdf`, `partition_text`).
- Output is a list of document elements that can be serialized to strings or structured formats.
- Modular functions and connectors form a system for data ingestion and pre-processing, intended to be adaptable across platforms.
- The README also describes an MCP server offering (Unstructured Transform) that exposes document processing to agents, handling 60+ file types with parsing, enrichment, chunking and embedding inside an agent session.
Installation and usage
- Install via PyPI: `pip install unstructured` for plain text, HTML, XML, JSON and emails; `pip install "unstructured[all-docs]"` for all document types; or targeted extras such as `unstructured[docx,pptx]`.
- Optional system dependencies include libmagic-dev (file type detection), poppler-utils (images/PDFs), tesseract-ocr (OCR, with tesseract-lang for extra languages), and libreoffice (MS Office docs); pandoc is bundled via pypandoc-binary.
- Docker images are published for x86_64 and Apple silicon; images are tagged by commit hash, version, and `latest`, and can also be built locally with `make docker-build`.
- Local development uses `uv` for dependency management; `make install` runs `uv sync --locked --all-extras --all-groups`. Optional pre-commit hooks and `make check`/`make tidy` support formatting and linting.
Documentation and community
The project points to docs.unstructured.io for quick start, core functionality, connectors, concepts and integrations. It provides a security policy, a bug report template (with `scripts/collect_env.py` for environment info), and links to a company website, full API documentation, and a separate batch-processing repository (unstructured-ingest).
Telemetry
The README states that lightweight analytics are sent by default: a library-load ping on import and a best-effort local attempt per top-level public partition call. Runtime events reportedly contain package version, normalized platform/Python/architecture values, fixed-enum partition characteristics, and aggregate final-element counts as URL query parameters, with no request body and no document content, filenames, paths, URLs, MIME values, exception details, credentials, or persistent identifiers. Delivery is described as non-blocking, without redirects, retries, response-body download, or queue, and not consulting proxy or netrc settings. Opting out is possible by setting `DO_NOT_TRACK` or `SCARF_NO_ANALYTICS` to any non-empty value before importing or partitioning.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.