Parallel-Trends Sensitivity: An Article-Reported Case

A tutorial built around article-reported bank-closure estimates showing why a high p-value can mislead and how sensitivity bounds probe fragility. The results have not been publicly reproduced.

When we run a difference-in-differences analysis, what does due diligence look like? At minimum, we check parallel trends. If the p-value is high, we're good, right?

Let's take a look at an article-reported case where the parallel-trends test returned p = 0.9997, yet additional checks changed the conclusion. The numerical results have not been publicly reproduced.


The Setup

Say we want to know whether bank branch closures affect SNAP (food stamp) participation. The hypothesis makes sense: when banks close, residents lose access to services that facilitate benefit delivery. Transaction costs rise. Enrollment might fall.

The article frames bank-branch closures from 2010 to 2020 as a potential natural experiment and tracks county SNAP participation before and after closures. Whether the design supports causal inference depends on the identifying assumptions tested below.

The article reports a 1,408-county sample, with some counties experiencing closures and others not, analyzed with the Callaway-Sant'Anna (2021) estimator. The estimator is designed for staggered treatment timing; the sample and implementation in this case are not publicly reproduced.

The article reports a 0.47 percentage-point decline in SNAP participation following bank closures. That estimate is an unreplicated association at this stage, so it should not be translated into affected-household counts or treated as a causal effect.

So far, so good. But can we trust this estimate?


The Standard Checks

Let's run through the usual diagnostics.

Parallel trends test: The article reports p = 0.9997 for the joint test of pre-treatment coefficients. A high p-value does not validate the assumption. The following values have not been publicly reproduced:

Article-reported pre-treatment coefficients; not publicly reproduced
Event Time Coefficient Standard Error
e = -3 -0.056 pp 0.620
e = -2 -0.030 pp 0.582
e = -1 +0.000 pp 0.549

In the article-reported output, these coefficients are statistically indistinguishable from zero. That result may reflect either parallel pre-trends or low power to detect violations.

Estimator comparison: The article reports -0.47 percentage points from Callaway-Sant'Anna and -0.50 from two-way fixed effects. Agreement can be informative, but both estimates may share the same identification problem and neither has been publicly reproduced here.

Plausible dynamics: The effect builds over time, as we'd expect if bank closures created persistent barriers. Counties don't suddenly drop SNAP participation; they drift downward over several years.

At this point, the causal claim looks solid. But here's where things get interesting.


A Warning Sign

One robustness check gives us pause: the fake timing test. The idea is to artificially shift treatment backward by two years and re-estimate. If parallel trends truly hold, this placebo should produce a null result.

The article reports that it does not: the fake-timing test has p = 0.04. That value is not publicly reproduced.

What does this mean? A significant placebo suggests that "pre-treatment" periods (under the fake timing) still show negative effects. That pattern is consistent with pre-existing trends, not parallel trends.

We could dismiss this as a fluke. The main parallel trends test passed overwhelmingly. One alternative test failing doesn't necessarily invalidate everything.

But let's dig deeper.


The Power Problem

Look again at the article-reported pre-treatment coefficients. Their standard errors imply confidence intervals spanning more than 2 percentage points, although those values have not been publicly reproduced.

The reported treatment estimate is -0.47 percentage points. A pre-trend of similar magnitude would fall well within the article-reported confidence interval around zero.

Here's the thing: the parallel trends test can't actually detect violations of the size that would matter.

This is the statistical power problem. A "passing" parallel trends test can mean two things:

  1. Parallel trends genuinely hold, or
  2. The test lacks power to detect violations

With the article-reported three pre-treatment periods and standard errors of 0.5 to 0.6, low power is plausible. A non-rejection tests whether the data reject parallel trends; it does not show that parallel trends hold.

That high reported p-value does not validate the assumption. Given the reported standard errors, the test may lack power to detect violations large enough to matter.


This Is Where Sensitivity Analysis Comes In

Rambachan and Roth (2023) offer a framework that doesn't assume parallel trends hold exactly. Instead, it parameterizes potential violations through a parameter M:

  • M = 0: Parallel trends assumed exactly (the standard assumption)
  • M = 1: Violations can be as large as the maximum observed pre-treatment coefficient movement
  • Breakdown M: The smallest M where the identified set includes zero

The idea here is to ask: how large do violations need to be before our conclusion changes?

The article reports the following bounds for different values of M. They are retained to teach interpretation but have not been publicly reproduced:

Article-reported Rambachan-Roth sensitivity bounds; not publicly reproduced
M ATT Bounds 95% CI Excludes Zero?
0 [-0.47, -0.47] [-0.91, -0.03] Yes
0.25 [-0.49, -0.45] [-0.93, -0.01] Yes
0.35 Breakdown
0.50 [-0.51, -0.42] [-0.95, +0.01] No
1.0 [-0.56, -0.38] [-0.99, +0.06] No

The article reports a breakdown M of 0.35.

In the article-reported output, the estimate is robust only to violations 35% as large as the maximum observed pre-trend movement. At M = 0.5, the reported confidence interval includes zero; at M = 1, the reported interval also includes positive values. These are provisional results.

A useful interpretive benchmark is whether the result survives violations at least as large as the observed pre-period movement, or M greater than 1. By that benchmark, the article-reported result is fragile.


Sensitivity analysis tells us the result is fragile. But it doesn't tell us what's actually happening. Is there selection into treatment?

Let's try adding county-specific linear time trends to our specification. This absorbs pre-existing trajectories. If the treatment effect is real, it should be identified off deviations from each county's own trend. If the effect is driven by pre-trends, it should disappear.

Article-reported baseline and county-trend specifications; not publicly reproduced
Specification ATT SE p-value
Baseline (County + Year FE) -0.47 0.22 0.036
With County Trends +0.003 0.016 0.87

In the article-reported county-trend specification, the estimate changes from -0.47 to +0.003 and the p-value from 0.036 to 0.87. This descriptive contrast is not publicly reproduced.

The article interprets this attenuation as consistent with pre-existing county trajectories. The specification alone does not prove that explanation, and the output is not publicly reproduced.


What's Going On Here?

The pattern now makes sense. Bank closures don't happen randomly. They happen in counties experiencing economic decline, population loss, reduced commercial activity. These same forces also reduce SNAP participation: fewer eligible residents, out-migration of low-income families, changing local economies.

In the article's interpretation, the design captured an association between bank closures and SNAP declines but could not separate a closure effect from pre-existing trajectories.

The combined diagnostic pattern is consistent with a low-power parallel-trends test and a problematic identification strategy. Because the analysis is not publicly reproduced, the case should be read as an instructional example rather than a verified empirical finding.


So What Can We Take Away?

Passing is not validation. A high p-value on a parallel trends test provides some evidence, but it's not proof. When pre-treatment periods are limited and standard errors are large, the test can't detect meaningful violations. We should report the power of our parallel trends tests alongside p-values.

Sensitivity analysis should be standard. Rambachan-Roth bounds show how conclusions change under deviations from parallel trends. In this article-reported case, the provisional breakdown M is 0.35.

Unit-specific trends are informative. Adding county-specific trends is a demanding specification. A large change in the estimate is a warning that pre-existing trajectories may matter; it does not by itself prove which mechanism generated the baseline association.

Listen to the warning signs. The article-reported fake-timing test has p = 0.04. If reproduced, that discrepancy would warrant investigation rather than dismissal.


The Revised Conclusion

Given the article-reported diagnostics, the defensible wording is an association rather than a causal claim. The following is the article's qualified summary, not a publicly reproduced estimate:

"Bank closures are associated with a 0.5 percentage point reduction in SNAP participation rates. However, sensitivity analysis reveals that treated counties were already on declining SNAP trajectories prior to bank closures. Causality cannot be established with the current identification strategy."

The wording is appropriately cautious, but the numerical association itself remains provisional until reproduced.


A Practical Checklist

For anyone running difference-in-differences analyses, here's what to check:

  1. How many pre-treatment periods do we have? Three or fewer often means low power. Be cautious about interpreting passing parallel trends tests.
  2. How large are the pre-treatment standard errors? If confidence intervals around pre-treatment coefficients are wider than the treatment effect, the test can't detect violations that matter.
  3. What is the Rambachan-Roth breakdown M? Interpret it relative to substantively plausible violations and the observed pre-period movement. The article uses 0.5 and 1 as practical warning benchmarks, not universal decision rules.
  4. Does the effect survive unit-specific trends? This is a demanding check, but informative. If the effect disappears, pre-existing trajectories may explain the result.
  5. Do alternative diagnostic tests agree? When the joint test passes but fake timing or placebo tests fail, investigate the discrepancy.

Sensitivity analysis doesn't make causal inference harder. It makes it honest.

This checklist is the worked basis for the quality assurance step of our Start Here guide, where parallel-trends, placebo, and leave-one-out diagnostics decide whether an estimate is ready to report.


References

Callaway, B., & Sant'Anna, P. H. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200-230.

Rambachan, A., & Roth, J. (2023). A more credible approach to parallel trends. Review of Economic Studies, 90(5), 2555-2591.


Public Materials

Article only. No matching public analysis script, frozen input manifest, run record, or saved output is currently linked for the bank-closure sample, estimates, p-values, sensitivity bounds, placebo test, or county-trend specification.

This project serves as the basis for an interactive methods lab at CAPHE (California Association of Public Health Economists): Understanding the Limits of Parallel Trends Tests. The lab includes an interactive Rambachan-Roth slider that lets readers explore how sensitivity bounds expand as M increases.

Suggested Citation

Cholette, V. (2025, December 10). Parallel-trends sensitivity: An article-reported case. Too Early To Say. https://tooearlytosay.com/research/methodology/parallel-trends-sensitivity/
Copy citation

Frequently asked questions

What does the parallel-trends assumption require?

That, absent treatment, the treated and comparison groups would have followed the same outcome trend. We can't observe this directly, so we test it on pre-treatment periods and probe how sensitive results are to violations.

What if a pre-treatment lead fails the parallel-trends test?

One failing lead doesn't automatically invalidate the design, but it shifts the burden to sensitivity analysis: we bound how large a trend violation would have to be to overturn the estimate. The numerical case in this article is article-reported and has not been publicly reproduced.

Are the bank-closure estimates publicly reproduced?

No. The sample, estimates, p-values, sensitivity bounds, and county-trend results are article-reported, and no matching public script, frozen input manifest, run record, or saved output is currently linked.