About this project
EconML is a Python package for estimating heterogeneous treatment effects from observational data via machine learning. It was designed and built as part of the ALICE project at Microsoft Research, combining state-of-the-art machine learning techniques with econometrics to bring automation to complex causal inference problems.
Core purpose: measure the causal effect of one or more treatment variables T on an outcome variable Y, controlling for features X and W, and how that effect varies as a function of X. Methods work with observational (non-experimental or historical) datasets. For causal interpretation, some methods assume no unobserved confounders, while others assume access to an instrument Z that affects treatment T but has no direct effect on outcome Y. Most methods provide confidence intervals and inference results.
Stated design goals include implementing recent techniques at the intersection of econometrics and machine learning, maintaining flexibility in modeling effect heterogeneity (random forests, boosting, lasso, neural nets) while preserving causal interpretation and often offering valid confidence intervals, using a unified API, and building on standard Python packages for machine learning and data analysis.
Implemented estimation families shown in the README include:
- Double Machine Learning (LinearDML, SparseLinearDML, NonParamDML, CausalForestDML) with OLS or bootstrap confidence intervals.
- Dynamic Double Machine Learning for panel data (DynamicDML).
- Orthogonal Random Forests (DMLOrthoForest, DROrthoForest).
- Meta-learners: XLearner, SLearner, TLearner.
- Doubly Robust learners: LinearDRLearner, SparseLinearDRLearner, ForestDRLearner.
- Instrumental variable methods: OrthoIV, NonParamDMLIV, LinearDRIV, SparseLinearDRIV, ForestDRIV, LinearIntentToTreatDRIV.
Interpretability support includes a tree interpreter of the CATE model (SingleTreeCateInterpreter), a policy interpreter (SingleTreePolicyInterpreter), and SHAP values for the CATE model.
Model selection and validation: an RScorer for causal model selection, scoring and weighted ensembling; built-in first-stage model selection that works with sklearn model selection classes such as LassoCV or GridSearchCV, or a list of candidate models.
Inference: when enabled, an InferenceResults object provides p-values, z-statistics, standard errors and confidence intervals; a summary() method is available for linear parametric CATE models, plus effect_inference summaries at sample and population level.
Policy learning: direct policy learning from observational data using a doubly robust method for offline policy learning, predicting a recommended treatment.
Installation is via pip install econml from PyPI; source installation is documented for developers. The project maintains an active release history and documentation at pywhy.org/EconML, with tests and documentation generation instructions for contributors.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.