About this project
Splink is a Python library for probabilistic record linkage (entity resolution), enabling deduplication and linking of datasets without unique identifiers. It implements the Fellegi-Sunter model with customizations for improved accuracy. Key features include high speed (linking ~1M records on a laptop in ~1 minute), support for fuzzy matching and term frequency adjustments, scalability via DuckDB (Python) or Spark/PostgreSQL (for 100M+ records), unsupervised learning (no training data required), and interactive visualizations for model diagnostics. Works best with multi-column, low-correlation input data (e.g., name, DOB, city for persons). Used in government, academia, and private sector. Offers backend-specific installations and extensive documentation with tutorials and examples.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.