When the parallel-trends test fails on one lead, what's left?

Worked example referenced by AI for Applied Researchers: Documentation

An article-reported two-way fixed-effects difference-in-differences exercise with Goodman-Bacon and leave-one-out diagnostics; the project panel and outputs are not public.

The article describes a two-way fixed effects (TWFE) specification on a 51-state panel: state and year fixed effects, SNAP take-up as the outcome, 41 adopters, and 10 never-adopters. The panel and output are not public. The instructional question remains useful: do the pre-treatment trends look parallel?

The article reports an F-test p-value of 0.041 with one testable lead. That numerical result is not publicly reproduced.

A rejected diagnostic creates two common temptations: soften the result as "marginal," or treat the entire effect as undefined. Neither response is warranted. The reported test rejects at the 5% level, but one lead cannot diagnose the full parallel-trends path. Its power against substantively relevant alternatives has not been calculated here. The rest of the diagnostic stack can narrow what remains credible.

The setting in one paragraph

Broad-Based Categorical Eligibility (BBCE) lets a state align SNAP categorical eligibility with receipt of Temporary Assistance for Needy Families (TANF)-funded benefits or services. In practice, a state can waive the SNAP asset test and raise the gross-income limit from 130% of the federal poverty guideline up to 200%. The first states adopted around 2000; by 2011, 41 states had adopted BBCE.1 The published estimates are a +5.9% TWFE coefficient and a +15.3% Callaway-Sant'Anna difference-in-differences (CSDD) estimate on log SNAP participation per capita.2 The difference between those published estimators motivates this diagnostic exercise.

The article specifies TWFE with state and year fixed effects and state-clustered standard errors. It reports an Integrated Public Use Microdata Series USA extract of the American Community Survey (IPUMS USA ACS) covering 2005 to 2016, with 612 state-year observations across 51 states, 41 ever-treated states, and 10 never-treated states. These sample counts and constructed outcomes have not been matched to a public extract. Bertrand, Duflo, and Mullainathan provide the external methodological reference for cluster-count concerns.3

BBCE rollout is staggered, and Wang et al. report materially different TWFE and CSDD estimates on this policy. That published comparison makes estimator choice part of the research question. The next validation step would apply both estimators to one frozen panel and add a BBCE-adjusted eligibility denominator.4

What TWFE actually returns on this panel

What does TWFE return under the article's stated panel specification?

The article reports a +1.37 percentage-point TWFE coefficient on SNAP take-up, a state-clustered standard error of 0.0102, a p-value of 0.18, and a 95% interval of [-0.63, +3.37] percentage points. It also reports a 0.410 baseline and a 3.35% relative comparison. These values are article-reported, not publicly reproduced.

For the log-SNAP-per-capita specification, the article reports +5.81% with a 95% interval of [-2.44%, +14.05%] and a p-value of 0.17. The published TWFE comparison is +5.9%.2 Numerical proximity is not a replication claim here because this site's panel and output are unavailable.

Why isn't the right reading "p > 0.05; not significant"?

Figure status: the event-study chart formerly shown here was removed because no public coefficient table or run record supports it.

The article-reported point estimate is numerically close to Wang's published TWFE estimate on one outcome and differs on another. That proximity is a comparison to investigate, not evidence that this site's exercise replicated the published analysis.

The article-reported p-value exceeds 0.05, so the stated estimate is not statistically distinguishable from zero under its specification. Numerical proximity to a published point estimate cannot substitute for reproducing the panel, confidence interval, covariate set, and outcome definition. A matched rerun must explain any difference in uncertainty rather than assume it comes from the sample window or controls.

The two outcome scales create a second discrepancy. The take-up-rate point estimate (+3.35% of baseline) and the log-pc point estimate (+5.81%) differ by about 2.5 percentage points. Same data, same method, two outcomes. What's going on?

A piece of the proposed explanation is denominator construction. The article says its take-up denominator uses an American Community Survey Public Use Microdata Sample (ACS PUMS) screen at 130% of the federal poverty level (FPL), while BBCE can raise the gross-income limit. The published analysis attributes approximately 11.5% of its participation increase to eligibility above 130% FPL.2 Whether this site's unreproduced denominator cleanly separates that channel cannot be verified from the public record.

The article defines its log-per-capita specification as the full SNAP caseload divided by state population. If reproduced, that outcome would combine the eligibility-expansion and take-up channels, while the stated 130% FPL outcome would target a narrower population.

The two outcome definitions answer different policy questions. A public rerun must reconstruct both denominators before comparing the article-reported coefficients.

The parallel-trends test, on one lead

The article reports a pre-period F-test p-value of 0.041 from a single lead at event time -2. The teaching point is that one lead provides weak information about the full parallel-trends assumption; the project-specific p-value remains unreproduced.

With one lead, the F-test is a test of one pre-treatment coefficient, not the full pre-treatment path. This page does not include a power calculation, so it cannot label the test powerful or powerless against violations of a specified magnitude. Roth (2022) shows why failing to reject parallel trends should not be confused with establishing that parallel trends hold, and why conditioning on a passed pre-test can distort subsequent inference.5

The article reports an event-time -2 coefficient of -0.0096 with SE 0.0046, opposite in sign to its reported TWFE coefficient. That sign check is a useful diagnostic to reproduce; it is not a verified result on this page.

We have a rejection on a single lead and a pre-period coefficient with the opposite sign to the reported TWFE coefficient. Where does that leave us? The F-test alone cannot resolve the parallel-trends question. We need evidence from the full pre-treatment path and the rest of the diagnostic stack.

Placebo: assign treatment two years too early

The first convergent check: assign a fake BBCE adoption date two years before the actual adoption date for each ever-treated state, restrict the sample to pre-actual-treatment observations plus the never-treated controls, and re-estimate TWFE. Under valid parallel trends, the placebo coefficient should be indistinguishable from zero.

The article reports a +0.78 percentage-point placebo estimate, SE 0.0087, p = 0.37, and 95% interval [-0.009, +0.025]. Those values and the interpretation that follows are article-reported; the placebo output is not public.

The placebo result alone does not establish parallel trends. Its reported confidence interval also includes effects that may be substantively relevant. Without the public panel and output, it does not resolve whether the rejected lead reflects a pre-trend, sampling variation, or another specification issue.

Leave-one-out: does any single state drive the result?

Second convergent check: drop each of the 41 ever-treated states one at a time and re-estimate TWFE on the remaining 50-state panel, containing 40 ever-treated and 10 never-treated states. If the result is sensitive to a single state's inclusion, the design is fragile.

The article reports a leave-one-out range of [+0.0108, +0.0159], median +0.0139, and interquartile range [+0.0133, +0.0150] across 41 iterations. No iteration table or saved output is public.

If reproduced, that range would bound the influence of deleting one treated state at a time. Whether a roughly -21% to +16% change is substantively stable should be judged against a declared threshold, not the fact that all iterations retain the same sign.

Figure status: the leave-one-out chart formerly shown here was removed because the iteration output is not public.

Goodman-Bacon: where does the TWFE coefficient come from?

The third check decomposes the TWFE coefficient into its component comparisons. Goodman-Bacon (2021) expresses that coefficient as the weighted average of all 2x2 difference-in-difference comparisons that the panel makes implicitly.6 Each comparison gets a weight that reflects how much identifying variation it contributes to the OLS coefficient, not how much of the sample it represents.

The conceptual buckets are:

  • Treated vs never-treated
  • Earlier-treated vs later-treated, with the later cohort as the not-yet-treated control
  • Later-treated vs earlier-treated, with the earlier cohort serving as an already-treated control

An earlier version attached weights and effect values to these buckets, but it mixed group means and weighted contributions and the numbers did not sum to the reported +0.0137 TWFE coefficient. Those numeric decomposition claims and charts have been withdrawn.

A valid release must publish every 2x2 component estimate τj and weight wj, verify that the weights sum to 1, and verify that Σj wjτj equals the TWFE coefficient. For a grouped table, the group weight is the sum of its component weights, the group mean is the within-group weighted mean, and the group contribution is group weight multiplied by group mean. The group contributions, not the unweighted group means, must reconstruct the coefficient.

The sample window can change decomposition weights because the weights depend on cohort shares, treatment timing, residualized treatment variation, and retained periods. A matched rerun should hold the treatment history fixed and compare nested windows. Without a valid component table, this page cannot attribute any decomposition difference to the sample window.

What the convergent checks establish

The article-reported diagnostics define a proposed audit sequence: inspect the one-lead pre-test, run a placebo, inspect the leave-one-out range, and produce a decomposition that exactly reconstructs the TWFE coefficient. The specific p-values and ranges on this page remain unreproduced; the earlier decomposition values have been withdrawn.

What does this tell us? A diagnostic stack can reveal more than a single pre-test, but it cannot make an unavailable analysis defensible by assertion. The design status remains article-reported until the panel, code, and outputs are public. A heterogeneity-robust estimator would be the appropriate next comparison.4

Figure status: the cross-study comparison chart formerly shown here was removed because this site's estimate is not publicly reproduced.

Two further things this specification does not handle

The first limit is the eligibility-expansion channel. The article's proposed 130% FPL denominator may omit participation above that threshold. Wang's published appendix provides the 11.5% decomposition cited here; this site's outcome construction and implied 88.5% interpretation are not publicly reproduced. A BBCE-adjusted denominator is a future robustness specification.

The second limit is BBCE reversion. The article reports two reversions and says it uses a time-varying treatment indicator. That coding choice and its post-reversion comparison role should be audited in a public treatment file before being described as canonical for this application.

Why show the conventional estimator

A conventional TWFE estimate provides a useful teaching anchor when it appears beside a heterogeneity-robust estimate and a decomposition of the identifying comparisons.

This page specifies that comparison but does not yet provide the public project package required to inspect it.

TWFE has a clear interpretation under simultaneous adoption and homogeneous effects. BBCE adoption is staggered, and Wang et al.2 report a materially larger heterogeneity-robust estimate in their published analysis. This article therefore treats TWFE as a teaching anchor, not a verified lower bound or recommended headline estimate.

What remains to reproduce

The article reports the following validation targets. None is reproduced by a matching public panel, script, and saved output:

  • A two-way fixed-effects (TWFE) coefficient of +1.37 percentage points on the stated take-up outcome, with p = 0.18.
  • A TWFE coefficient of +5.81% on log Supplemental Nutrition Assistance Program (SNAP) participation per capita, compared with a published +5.9% benchmark.
  • A pre-period test with p = 0.041 on one testable lead.
  • A placebo estimate of +0.78 percentage points with p = 0.37.
  • A reported leave-one-out range of [+0.0108, +0.0159], about 21% below to 16% above the article's +0.0137 headline coefficient.
  • A complete Goodman-Bacon component table whose weights sum to 1 and whose weighted component estimates reconstruct the TWFE coefficient.

A credible rerun would freeze the treatment history, outcome denominator, exclusion rules, and state-year panel; reproduce each target above; then estimate a heterogeneity-robust comparison on the same rows. Until then, the page demonstrates the diagnostic sequence and makes no causal claim about Broad-Based Categorical Eligibility (BBCE).

The next methods comparison is the Callaway-Sant'Anna estimator on that frozen panel.

See also the literature review step for setting a published benchmark and the documentation step for tracing a result from raw data to saved output.

References

  1. Ganong, P., & Liebman, J. B. (2018). The decline, rebound, and further rise in SNAP enrollment: Disentangling business cycle fluctuations and policy changes. American Economic Journal: Economic Policy, 10(4), 153-176.
  2. Wang, X., Valizadeh, P., Nayga, R. M., Bryant, H. L., & Fischer, B. L. (2026). Broad-based categorical eligibility policy and SNAP participation. Journal of Policy Analysis and Management, 45(1), e70063.
  3. Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How much should we trust differences-in-differences estimates? The Quarterly Journal of Economics, 119(1), 249-275.
  4. Callaway, B., & Sant'Anna, P. H. C. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200-230.
  5. Roth, J. (2022). Pretest with caution: Event-study estimates after testing for parallel trends. American Economic Review: Insights, 4(3), 305-322.
  6. Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. Journal of Econometrics, 225(2), 254-277.