About this project
dlt (data load tool) is an open-source Python library that automates data extraction, transformation, and loading. It is designed as a library you drop into existing Python code rather than a managed platform.
**Extraction sources** include REST APIs (with declarative pagination, filtering, and flattening), SQL databases (with automatic reflection of tables and types), files in cloud buckets (S3, GCS, Azure — CSV, JSONL, Parquet), and DataFrames (pandas, Polars, PyArrow with zero-copy Arrow support).
**Loading** supports 20+ destinations including DuckDB, BigQuery, Snowflake, Postgres, Redshift, Databricks, Athena, ClickHouse, MotherDuck, Iceberg, Delta Lake, and generic filesystems. Destination changes require only updating the `destination` string.
Key capabilities:
- **Schema inference and typing**: Automatically infers schema from source data and maps types across destinations.
- **Schema contracts**: Three modes — `evolve` (adapt schema), `freeze` (reject mismatches), `discard` (drop offending rows/columns).
- **Incremental loading**: Built-in support for loading only new or changed records using cursors or timestamps.
- **Merge/upsert strategies**: Declarative primary-key based merge via the `write_disposition="merge"` option.
- **Nested data normalization**: Automatically flattens and normalizes nested structures.
- **Ibis integration**: Loaded tables can be lifted into Ibis expressions for composable Python queries that compile to SQL in the destination dialect.
- **Secrets management**: Credentials injected from `secrets.toml` or environment variables.
- **Pipeline reconnection**: Use `dlt.attach()` to reconnect to a pipeline by name and read data back via a `Dataset` API that supports `.df()`, `.arrow()`, and `.to_ibis()`.
dlt is built with LLM and coding-agent compatibility in mind, offering declarative, human-readable primitives suitable for AI-assisted pipeline generation. It supports Python 3.10 through 3.14.
Installation: `pip install dlt` with optional extras such as `[duckdb]`, `[bigquery]`, `[sql_database]`, and `[hub]`.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.