About this project

dlt (data load tool) is an open-source Python library that automates data extraction, transformation, and loading. It is designed as a library you drop into existing Python code rather than a managed platform. **Extraction sources** include REST APIs (with declarative pagination, filtering, and flattening), SQL databases (with automatic reflection of tables and types), files in cloud buckets (S3, GCS, Azure — CSV, JSONL, Parquet), and DataFrames (pandas, Polars, PyArrow with zero-copy Arrow support). **Loading** supports 20+ destinations including DuckDB, BigQuery, Snowflake, Postgres, Redshift, Databricks, Athena, ClickHouse, MotherDuck, Iceberg, Delta Lake, and generic filesystems. Destination changes require only updating the `destination` string. Key capabilities: - **Schema inference and typing**: Automatically infers schema from source data and maps types across destinations. - **Schema contracts**: Three modes — `evolve` (adapt schema), `freeze` (reject mismatches), `discard` (drop offending rows/columns). - **Incremental loading**: Built-in support for loading only new or changed records using cursors or timestamps. - **Merge/upsert strategies**: Declarative primary-key based merge via the `write_disposition="merge"` option. - **Nested data normalization**: Automatically flattens and normalizes nested structures. - **Ibis integration**: Loaded tables can be lifted into Ibis expressions for composable Python queries that compile to SQL in the destination dialect. - **Secrets management**: Credentials injected from `secrets.toml` or environment variables. - **Pipeline reconnection**: Use `dlt.attach()` to reconnect to a pipeline by name and read data back via a `Dataset` API that supports `.df()`, `.arrow()`, and `.to_ibis()`. dlt is built with LLM and coding-agent compatibility in mind, offering declarative, human-readable primitives suitable for AI-assisted pipeline generation. It supports Python 3.10 through 3.14. Installation: `pip install dlt` with optional extras such as `[duckdb]`, `[bigquery]`, `[sql_database]`, and `[hub]`.