About this project
Unstract is an open-source platform designed to turn unstructured documents (PDFs, images, scans, spreadsheets, presentations) into structured JSON data using large language models (LLMs). It allows users to define extraction schemas with natural language prompts, then deploy the extraction as a REST API or an ETL pipeline. Key features include a Prompt Studio for schema definition, API deployment, ETL pipeline integration, an MCP server for AI agents, and an n8n node for automation workflows. It supports multiple LLM providers (OpenAI, Anthropic, Bedrock, Gemini, Ollama, etc.), vector databases (Qdrant, Pinecone, Weaviate, etc.), and text extractors (LLMWhisperer, Unstructured.io, LlamaIndex Parse). The platform can be run locally via Docker Compose with a simple script, and offers cloud/enterprise options with additional features like dual-LLM verification, human-in-the-loop review, SSO, and compliance certifications. It is built with a React frontend, Django backend, Celery worker, and FastAPI platform service, using PostgreSQL, Redis, and RabbitMQ. The project is licensed under AGPL-3.0 and welcomes contributions.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.