When we run a difference-in-differences analysis, what does due diligence look like? At minimum, we check parallel trends. If the p-value is high, we're good, right?
Let's take a look at an article-reported case where the parallel-trends test returned p = 0.9997, yet additional checks changed the conclusion. The numerical results have not been publicly reproduced.
The Setup
Say we want to know whether bank branch closures affect SNAP (food stamp) participation. The hypothesis makes sense: when banks close, residents lose access to services that facilitate benefit delivery. Transaction costs rise. Enrollment might fall.
The article frames bank-branch closures from 2010 to 2020 as a potential natural experiment and tracks county SNAP participation before and after closures. Whether the design supports causal inference depends on the identifying assumptions tested below.
The article reports a 1,408-county sample, with some counties experiencing closures and others not, analyzed with the Callaway-Sant'Anna (2021) estimator. The estimator is designed for staggered treatment timing; the sample and implementation in this case are not publicly reproduced.
The article reports a 0.47 percentage-point decline in SNAP participation following bank closures. That estimate is an unreplicated association at this stage, so it should not be translated into affected-household counts or treated as a causal effect.
So far, so good. But can we trust this estimate?
The Standard Checks
Let's run through the usual diagnostics.
Parallel trends test: The article reports p = 0.9997 for the joint test of pre-treatment coefficients. A high p-value does not validate the assumption. The following values have not been publicly reproduced:
| Event Time | Coefficient | Standard Error |
|---|---|---|
| e = -3 | -0.056 pp | 0.620 |
| e = -2 | -0.030 pp | 0.582 |
| e = -1 | +0.000 pp | 0.549 |
In the article-reported output, these coefficients are statistically indistinguishable from zero. That result may reflect either parallel pre-trends or low power to detect violations.
Estimator comparison: The article reports -0.47 percentage points from Callaway-Sant'Anna and -0.50 from two-way fixed effects. Agreement can be informative, but both estimates may share the same identification problem and neither has been publicly reproduced here.
Plausible dynamics: The effect builds over time, as we'd expect if bank closures created persistent barriers. Counties don't suddenly drop SNAP participation; they drift downward over several years.
At this point, the causal claim looks solid. But here's where things get interesting.
A Warning Sign
One robustness check gives us pause: the fake timing test. The idea is to artificially shift treatment backward by two years and re-estimate. If parallel trends truly hold, this placebo should produce a null result.
The article reports that it does not: the fake-timing test has p = 0.04. That value is not publicly reproduced.
What does this mean? A significant placebo suggests that "pre-treatment" periods (under the fake timing) still show negative effects. That pattern is consistent with pre-existing trends, not parallel trends.
We could dismiss this as a fluke. The main parallel trends test passed overwhelmingly. One alternative test failing doesn't necessarily invalidate everything.
But let's dig deeper.
The Power Problem
Look again at the article-reported pre-treatment coefficients. Their standard errors imply confidence intervals spanning more than 2 percentage points, although those values have not been publicly reproduced.
The reported treatment estimate is -0.47 percentage points. A pre-trend of similar magnitude would fall well within the article-reported confidence interval around zero.
Here's the thing: the parallel trends test can't actually detect violations of the size that would matter.
This is the statistical power problem. A "passing" parallel trends test can mean two things:
- Parallel trends genuinely hold, or
- The test lacks power to detect violations
With the article-reported three pre-treatment periods and standard errors of 0.5 to 0.6, low power is plausible. A non-rejection tests whether the data reject parallel trends; it does not show that parallel trends hold.
That high reported p-value does not validate the assumption. Given the reported standard errors, the test may lack power to detect violations large enough to matter.
This Is Where Sensitivity Analysis Comes In
Rambachan and Roth (2023) offer a framework that doesn't assume parallel trends hold exactly. Instead, it parameterizes potential violations through a parameter M:
- M = 0: Parallel trends assumed exactly (the standard assumption)
- M = 1: Violations can be as large as the maximum observed pre-treatment coefficient movement
- Breakdown M: The smallest M where the identified set includes zero
The idea here is to ask: how large do violations need to be before our conclusion changes?
The article reports the following bounds for different values of M. They are retained to teach interpretation but have not been publicly reproduced:
| M | ATT Bounds | 95% CI | Excludes Zero? |
|---|---|---|---|
| 0 | [-0.47, -0.47] | [-0.91, -0.03] | Yes |
| 0.25 | [-0.49, -0.45] | [-0.93, -0.01] | Yes |
| 0.35 | — | — | Breakdown |
| 0.50 | [-0.51, -0.42] | [-0.95, +0.01] | No |
| 1.0 | [-0.56, -0.38] | [-0.99, +0.06] | No |
The article reports a breakdown M of 0.35.
In the article-reported output, the estimate is robust only to violations 35% as large as the maximum observed pre-trend movement. At M = 0.5, the reported confidence interval includes zero; at M = 1, the reported interval also includes positive values. These are provisional results.
A useful interpretive benchmark is whether the result survives violations at least as large as the observed pre-period movement, or M greater than 1. By that benchmark, the article-reported result is fragile.
One More Test: County-Specific Trends
Sensitivity analysis tells us the result is fragile. But it doesn't tell us what's actually happening. Is there selection into treatment?
Let's try adding county-specific linear time trends to our specification. This absorbs pre-existing trajectories. If the treatment effect is real, it should be identified off deviations from each county's own trend. If the effect is driven by pre-trends, it should disappear.
| Specification | ATT | SE | p-value |
|---|---|---|---|
| Baseline (County + Year FE) | -0.47 | 0.22 | 0.036 |
| With County Trends | +0.003 | 0.016 | 0.87 |
In the article-reported county-trend specification, the estimate changes from -0.47 to +0.003 and the p-value from 0.036 to 0.87. This descriptive contrast is not publicly reproduced.
The article interprets this attenuation as consistent with pre-existing county trajectories. The specification alone does not prove that explanation, and the output is not publicly reproduced.
What's Going On Here?
The pattern now makes sense. Bank closures don't happen randomly. They happen in counties experiencing economic decline, population loss, reduced commercial activity. These same forces also reduce SNAP participation: fewer eligible residents, out-migration of low-income families, changing local economies.
In the article's interpretation, the design captured an association between bank closures and SNAP declines but could not separate a closure effect from pre-existing trajectories.
The combined diagnostic pattern is consistent with a low-power parallel-trends test and a problematic identification strategy. Because the analysis is not publicly reproduced, the case should be read as an instructional example rather than a verified empirical finding.
So What Can We Take Away?
Passing is not validation. A high p-value on a parallel trends test provides some evidence, but it's not proof. When pre-treatment periods are limited and standard errors are large, the test can't detect meaningful violations. We should report the power of our parallel trends tests alongside p-values.
Sensitivity analysis should be standard. Rambachan-Roth bounds show how conclusions change under deviations from parallel trends. In this article-reported case, the provisional breakdown M is 0.35.
Unit-specific trends are informative. Adding county-specific trends is a demanding specification. A large change in the estimate is a warning that pre-existing trajectories may matter; it does not by itself prove which mechanism generated the baseline association.
Listen to the warning signs. The article-reported fake-timing test has p = 0.04. If reproduced, that discrepancy would warrant investigation rather than dismissal.
The Revised Conclusion
Given the article-reported diagnostics, the defensible wording is an association rather than a causal claim. The following is the article's qualified summary, not a publicly reproduced estimate:
"Bank closures are associated with a 0.5 percentage point reduction in SNAP participation rates. However, sensitivity analysis reveals that treated counties were already on declining SNAP trajectories prior to bank closures. Causality cannot be established with the current identification strategy."
The wording is appropriately cautious, but the numerical association itself remains provisional until reproduced.
A Practical Checklist
For anyone running difference-in-differences analyses, here's what to check:
- How many pre-treatment periods do we have? Three or fewer often means low power. Be cautious about interpreting passing parallel trends tests.
- How large are the pre-treatment standard errors? If confidence intervals around pre-treatment coefficients are wider than the treatment effect, the test can't detect violations that matter.
- What is the Rambachan-Roth breakdown M? Interpret it relative to substantively plausible violations and the observed pre-period movement. The article uses 0.5 and 1 as practical warning benchmarks, not universal decision rules.
- Does the effect survive unit-specific trends? This is a demanding check, but informative. If the effect disappears, pre-existing trajectories may explain the result.
- Do alternative diagnostic tests agree? When the joint test passes but fake timing or placebo tests fail, investigate the discrepancy.
Sensitivity analysis doesn't make causal inference harder. It makes it honest.
This checklist is the worked basis for the quality assurance step of our Start Here guide, where parallel-trends, placebo, and leave-one-out diagnostics decide whether an estimate is ready to report.
References
Callaway, B., & Sant'Anna, P. H. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200-230.
Rambachan, A., & Roth, J. (2023). A more credible approach to parallel trends. Review of Economic Studies, 90(5), 2555-2591.
Public Materials
Article only. No matching public analysis script, frozen input manifest, run record, or saved output is currently linked for the bank-closure sample, estimates, p-values, sensitivity bounds, placebo test, or county-trend specification.
This project serves as the basis for an interactive methods lab at CAPHE (California Association of Public Health Economists): Understanding the Limits of Parallel Trends Tests. The lab includes an interactive Rambachan-Roth slider that lets readers explore how sensitivity bounds expand as M increases.
Suggested Citation
Cholette, V. (2025, December 10). Parallel-trends sensitivity: An article-reported case. Too Early To Say. https://tooearlytosay.com/research/methodology/parallel-trends-sensitivity/Copy citation