About this project
SDV (Synthetic Data Vault) is a Python library for creating tabular synthetic data. It learns patterns from real data using a range of machine learning models and reproduces them in generated data.
Key capabilities described in the README:
- Multiple synthesizers, from classical statistical methods (GaussianCopula) to deep learning (CTGAN).
- Support for single tables, multiple connected tables, and sequential tables.
- Metadata-driven modeling: column data types, primary keys and relationships are described in metadata.
- Evaluation and visualization: compare synthetic vs real data with quality reports and column plots.
- Preprocessing, anonymization options, and logical constraints to encode business rules.
- Demo datasets and tutorials via Colab notebooks, plus documentation and a community forum.
Typical workflow: download a demo dataset and metadata, instantiate a synthesizer such as GaussianCopulaSynthesizer, fit it on real data, then sample synthetic rows. A quality report computes an overall score plus column shape and column pair trend breakdowns.
The project is part of The Synthetic Data Vault Project, originally created at MIT's Data to AI Lab in 2016 and later developed by DataCebo. It is distributed under the Business Source License and installable via pip or conda.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.