Applied Econometrics
14 articles
Matching in Python: Balance Does Not Prove Validity
Propensity-score matching returns 2.21 for a planted effect of 2.0 with a balance table that passes the 0.10 rule. The overlap diagnostic shows 6% of treated units have no comparable control; trimming recovers 2.04.
Difference-in-Differences in Python: When TWFE Misleads
With staggered adoption and heterogeneous effects, two-way fixed effects returns 1.01 where the planted average is 1.60, and a group-time estimator with clean controls recovers 1.60.
Synthetic control in Python: read the pre-fit before the gap
A hands-on synthetic-control workflow: inspect constraints and rank, validate on held-out pre-treatment periods, and treat fit as necessary rather than sufficient.
Regression discontinuity in Python: effects at the cutoff
An overconstrained global shortcut returns a clean, plausible 1.8 where the effect planted at the cutoff is 0.75. A local fit recovers about 0.75 and makes the identifying assumptions visible.
Validating a Double Machine Learning Estimate
Double machine learning in Python: why a naive plug-in reads a true effect of 1.0 as 0.55, how cross-fitting recovers 0.97, and the confounder it still cannot detect.
Using difference-in-differences in practice
When difference-in-differences fits, which assumptions matter, and how it breaks, shown in worked cases with study-level material status.
Instrumental Variables in Python: Strength Is Not Validity
2SLS recovers a planted effect of 2.0 where OLS reads 2.79, but only if the exclusion restriction holds. A small direct path biases 2SLS to 2.62 while the first-stage F stays 2051, uncatchable in-sample.
How do we know an AI's estimator does what we meant?
AI-generated econometric code can run without error and still be wrong. A routine to verify it: spec the low-visibility choices, plant a known truth, and read the code against its source.
Rolling DiD with Few Units: Article-Reported Simulation
A guide to rolling difference-in-differences with few units. Simulation estimates and the Python rebuild are article-reported, not reproduced.
Regression vs. language-model readouts for survey prediction
An article-reported local comparison of logistic regression and language-model readouts. Scripts, provenance, and matched outputs are not public.
Prediction-powered inference for AI-assisted survey estimation
A method guide with an article-reported HINTS imputation example and a separate regression-adjustment simulation. The local results are not publicly reproduced.
A regression view of steering vectors
An article-reported local experiment comparing steering-vector estimators. The scripts, advance document, and matched outputs are not publicly reproduced.
Well-Executed but Not Important
Publication scope and citation patterns across four health-economics journals, and the limits of using citations to infer editorial importance.
Cycling Through Bad Ideas Faster
An article-reported Medicaid branding example that illustrates iterative critique; the project results and timings are not publicly reproduced.