What County Comparisons Teach Us About Measurement
If we want to understand food access across California, we might start by comparing counties. In this article-reported case, applying one methodology to all 58 counties produces a large spread: Merced (index 0.595) is reported at 2.3 times San Francisco (0.260). 1 The article also reports mobility-desert rates from under 5% to over 40%. These values have not been publicly reproduced.
But this variation raises a question we need to take seriously: Are we measuring differences in food access, or differences in something else entirely?
By construction, a high score summarizes high values on the five index components. The score does not independently establish that residents face a particular barrier, identify what caused the score, or determine whether an intervention is warranted.
The index includes poverty, renter share, minority share, sprawl, and food-access measures, so relationships between those inputs and the final score are partly mechanical. A raw county comparison can describe the composite but cannot estimate how much any factor caused the difference or whether policy produced it. Here is how to examine that interpretation problem more closely.
The Statewide Picture
The article reports a master analysis of 9,039 residential census tracts across all 58 California counties. 2 The following article-reported summary has not been publicly reproduced:
| Metric | Value |
|---|---|
| Mean vulnerability index | 0.318 |
| Median vulnerability index | 0.308 |
| Standard deviation | 0.084 |
| Range | 0.119 - 0.797 |
The vulnerability index combines five components: food access (25%), poverty rate (25%), renter percentage (20%), minority percentage (15%), and sprawl index (15%). Each is normalized 0-1, with higher scores indicating greater vulnerability.
Vulnerability Distribution
Evidence note: the counts and percentages in this table are article-reported and not publicly reproduced.
| Level | Tracts | Percentage |
|---|---|---|
| Low (0-0.25) | 1,826 | 20.2% |
| Moderate (0.25-0.50) | 6,916 | 76.5% |
| High (0.50-0.75) | 294 | 3.3% |
| Very High (0.75-1.0) | 3 | 0.03% |
Within the article-reported output, most tracts fall in the moderate range and only a small number fall in the highest range. Without public inputs and output, those distributional claims should be treated as provisional.
County Rankings: Top and Bottom
The article reports the following 10 highest-scoring counties:
| Rank | County | Vulnerability Index | Food Desert % |
|---|---|---|---|
| 1 | Merced | 0.595 | 28.4% |
| 2 | Alpine | 0.531 | 50.0% |
| 3 | Tulare | 0.489 | 31.2% |
| 4 | Kern | 0.476 | 27.8% |
| 5 | Imperial | 0.471 | 33.3% |
| 6 | Madera | 0.458 | 25.0% |
| 7 | Kings | 0.455 | 35.7% |
| 8 | Fresno | 0.447 | 22.1% |
| 9 | Colusa | 0.442 | 40.0% |
| 10 | Glenn | 0.438 | 37.5% |
The article reports the following 10 lowest-scoring counties:
| Rank | County | Vulnerability Index | Food Desert % |
|---|---|---|---|
| 49 | Marin | 0.275 | 8.3% |
| 50 | Santa Clara | 0.273 | 5.2% |
| 51 | Contra Costa | 0.272 | 11.4% |
| 52 | Orange | 0.271 | 9.8% |
| 53 | Santa Cruz | 0.269 | 12.5% |
| 54 | Ventura | 0.268 | 10.2% |
| 55 | San Mateo | 0.268 | 4.7% |
| 56 | Placer | 0.264 | 14.3% |
| 57 | El Dorado | 0.261 | 18.8% |
| 58 | San Francisco | 0.260 | 2.1% |
Variables Associated with the Reported Variation
The article reports a Central Valley cluster among higher-scoring counties and Bay Area and coastal counties among lower-scoring counties. That reported pattern motivates the structural-context lesson, but the ranking output is not publicly reproduced.
Component 1: Income and Poverty
The article reports lower median incomes and higher poverty rates for its Central Valley grouping:
| Region | Median Income | Poverty Rate |
|---|---|---|
| Central Valley counties | $58,000 | 18.2% |
| Bay Area counties | $115,000 | 8.4% |
The article reports a correlation of r = -0.71 between income and the index. Because poverty is assigned 25% of the index weight, part of that relationship is mechanical. The reported correlation has not been publicly reproduced.
Component 2: Geography and Sprawl
The article reports lower density and more sprawling development for its Central Valley grouping:
| Region | Median Pop Density | Sprawl Index |
|---|---|---|
| Central Valley | 120/sq mi | 8.4 |
| Bay Area | 2,400/sq mi | 1.2 |
The specification uses sprawl as 15% of the index and therefore assigns different scores to counties with different sprawl values. The cross-sectional comparison does not estimate a causal effect of sprawl or density on transit service.
Component 3: Retail Distribution
The specification assigns 25% of the index to its food-access component. The article reports geographic differences in store density and dispersion, but those values and their relationship to development patterns have not been publicly reproduced.
Component 4: Demographics
Minority percentage and renter rate carry a combined 35% of the index weight. Any association between those inputs and the composite is therefore partly built into the formula; this comparison does not estimate their causal effects or the effect of county wealth.
The Interpretation Challenge
Because the index directly includes demographic and structural measures, a raw ranking cannot isolate policy performance. The article-reported Merced and San Francisco comparison illustrates the interpretive problem:
| Metric | Merced | San Francisco |
|---|---|---|
| Vulnerability index | 0.595 | 0.260 |
| Median income | $52,000 | $119,136 |
| Population density | 155/sq mi | 18,600/sq mi |
| Car ownership | 91% | 65% |
The reported Merced and San Francisco values differ on the index, income, density, and car-ownership measures shown here. Those descriptive differences do not establish resident-level hardship, the performance of either county’s food policy, or conditions in a specific neighborhood.
Treating the index difference as a food-policy effect would conflate the composite outcome with its correlated inputs. The comparison does not show which inputs determine a baseline or what either county’s score would have been under another policy.
Improving Cross-County Comparisons
Several approaches can make cross-county descriptions more comparable. None of them, by itself, converts this cross-section into a causal design:
1. Compare Similar Counties
Grouping counties by structural characteristics allows more meaningful comparison:
Rural Central Valley: Merced, Madera, Kings, Fresno, Tulare
- Compare index values within this peer group
- Narrow observed differences in measured structure while retaining the possibility of unmeasured confounding
Bay Area suburban: Alameda, Contra Costa, Santa Clara
- Use a different descriptive benchmark from the Central Valley grouping
- Compare with other suburban regions while reporting remaining differences
2. Use Residualized Metrics
A statistical model can describe variation left unexplained by the structural variables included in that model:
Residual = Actual vulnerability - Predicted vulnerability (from income, density, etc.)
A positive residual means the observed index exceeds that model’s fitted value; a negative residual means it falls below. A residual does not isolate policy-attributable variation, because omitted variables, measurement choices, and model specification remain in it.
This approach reframes the descriptive question as "how does the observed score compare with this model’s fitted score?" It does not answer why the difference exists.
3. Track Within-County Variation
County averages can mask substantial within-county heterogeneity. The article reports Los Angeles County tracts from 0.15 to 0.72 and a county average of 0.38; those values have not been publicly reproduced.
Tract-level analysis within counties preserves more geographic variation than a county average. Whether a pattern is actionable requires direct evidence about needs, programs, and feasible responses.
4. Focus on Change Over Time
Comparing the same county at different time points removes time-invariant county differences from a simple change calculation. If a hypothetical Merced index moves from 0.59 to 0.55, the composite changed; the comparison alone does not show which input changed or why.
Longitudinal data can support designs that address timing and stable differences, but a before-and-after change is not itself evidence of a policy effect. Other concurrent changes and trends still require an identification strategy.
What Raw Rankings Do Tell Us
Despite the interpretation challenges, county rankings provide useful information:
They rank the composite as specified. A high county score identifies a high value on this particular weighted index. It does not independently verify resident-level need or justify a funding formula, but it can identify places for follow-up with direct measures.
They display regional patterns. The article-reported Central Valley cluster describes counties with similar index positions. The clustering is consistent with shared or correlated regional features, but cross-sectional data cannot establish that any listed feature caused it.
They provide a benchmark for tracking the same measure. A county can monitor whether its index changes over time relative to a defined comparison group. That tracks the composite, not an independently validated improvement in conditions or a policy effect.
They identify candidates for deeper investigation. Counties that show vulnerability levels substantially different from what their structural factors predict are candidates for case study research to understand what unmeasured factors distinguish them. Such research requires additional data and methods beyond cross-sectional comparison.
The Aggregation Problem
County-level analysis has a deeper problem: aggregation obscures tract-level variation. This is related to the "ecological fallacy" in social science research, where relationships observed at the aggregate level may not hold at the individual level. 3
Consider two hypothetical counties, each with 100 tracts:
County A: All tracts have vulnerability index 0.35
- County average: 0.35
- Gini coefficient: 0 (perfect equality)
County B: 50 tracts at 0.20, 50 tracts at 0.50
- County average: 0.35
- Gini coefficient: approximately 0.214 under the standard uncorrected formula
Both counties have an average of 0.35. County A assigns that value to every tract, while County B contains two equally sized groups at 0.20 and 0.50. The Gini differs because it summarizes that within-county dispersion.
Reporting only the average would hide this distributional difference. The example does not show that either county needs a particular intervention; choosing between universal and targeted programs requires direct measures of need, program eligibility, costs, and expected effects.
Within-County Variation in California
The article reports the following within-county means and Gini coefficients. They have not been publicly reproduced:
| County | Mean Vulnerability | Gini Coefficient |
|---|---|---|
| Los Angeles | 0.38 | 0.24 |
| San Diego | 0.34 | 0.19 |
| San Francisco | 0.26 | 0.11 |
| Fresno | 0.44 | 0.18 |
In the article-reported output, Los Angeles has the highest within-county inequality among the counties shown, with tracts reported from 0.15 to 0.72; San Francisco tracts cluster closer to the reported county mean. These comparisons are provisional pending reproduction.
Cost-of-Living Sensitivity Analysis
One reasonable objection to cross-county comparisons: shouldn't we adjust for differences in cost of living? A dollar goes further in Merced than in San Francisco. Perhaps vulnerability rankings would change substantially if we accounted for regional price differences.
The article reports a sensitivity test that merged 2023 Regional Price Parities (RPP) from the Bureau of Economic Analysis with county-level vulnerability scores. 4 RPP measures local prices relative to the national average (100 = national average). In California, values range from 98.4 (Kings County/Hanford-Corcoran metro) to 118.2 (San Francisco Bay Area counties).
We applied a partial adjustment that scales the income-sensitive portion of vulnerability by regional price levels:
Adjusted vulnerability = Original × (0.75 + 0.25 × county_RPP / state_avg_RPP)
This formula scales the entire original index by a factor whose RPP term receives a 25% weight. It does not isolate and price-adjust only the poverty component, so it should be read as a coarse sensitivity transformation rather than a component-specific correction.
Article-reported result: minimal impact on rankings; not publicly reproduced.
| Metric | Value |
|---|---|
| Maximum rank change | 3 positions |
| Counties unchanged | 28 of 58 (48%) |
| Correlation (original vs. adjusted) | 0.998 |
In the article-reported output, Merced and San Francisco remain at the highest and lowest ends of the ranking, and Imperial, Tulare, and Kings each shift by three positions. These findings are provisional pending public reproduction.
Why might the article report little rank movement? The adjustment changes only one weighted component, and several index inputs may be correlated with local prices. Those mechanical features could limit movement, but the public materials do not include a reproducible model with which to test that explanation.
The article-reported within-state result cannot establish how an adjustment would behave in another state or specification. Differences in Regional Price Parities (RPPs) describe price levels; they do not, by themselves, predict the direction or size of a ranking change.
Interpretive Implications
As a methodological illustration, the article-reported values support cautions about what this county-level index can claim. Even after reproduction, a descriptive ranking would require additional evidence before it could guide resource allocation or a specific program:
Do not infer a prescribed intervention from the ranking. A high score can flag a county for follow-up, but this comparison does not show that its score stems from geography or income distribution and does not estimate the effects or costs of mobile markets, demand-responsive transit, or SNAP outreach.
Interpret county rankings as weighted descriptions, not policy effects. The raw rankings incorporate poverty, renter share, minority share, sprawl, and food access by construction. Cross-sectional county comparisons cannot distinguish policy choices, structural constraints, or unmeasured factors, and the index alone does not verify resident-level barriers.
Do not use the index alone for resource allocation or causal attribution. The composite locates high index values, not individual eligibility or independently measured need. Allocation decisions require direct evidence and explicit value judgments; policy-effect claims require an identification strategy.
Look within counties for variation in similar structural contexts. The article reports example scores of 0.65 in south Los Angeles and 0.18 in Santa Monica. Those values are provisional, but the general caution remains: county averages can obscure heterogeneity.
Track change over time without treating change as attribution. A 0.05 decline means the weighted index fell by that amount. It does not show which condition improved or whether a policy caused the change; those conclusions require component-level analysis and an identification strategy.
Limitations
Index construction choices matter. The weights (25/25/20/15/15) are defensible but not unique. Different weights would change rankings.
Structural factors are imperfectly measured. Income, density, and sprawl capture some structural variation but not all. Unmeasured factors (political culture, historical investment patterns, geographic constraints) also matter.
County boundaries are arbitrary. Census tracts near county borders may have more in common with neighboring tracts in other counties than with distant tracts in the same county.
Cross-sectional data can't establish causation. We observe correlation between structure and vulnerability. We cannot determine from this data whether changing structural factors would change vulnerability.
Data and Methods
Data sources:
- Grocery distances: Calculated from population-weighted tract centroids
- Demographics: ACS 2019-2023 5-year estimates
- Sprawl metrics: Calculated from tract area and population
Vulnerability index components:
| Component | Weight | Source |
|---|---|---|
| Food access (distance) | 25% | Grocery distance analysis |
| Poverty rate | 25% | ACS Table S1701 |
| Renter percentage | 20% | ACS Table B25003 |
| Minority percentage | 15% | ACS Table B03002 |
| Sprawl index | 15% | Calculated (area/population) |
County rankings: Based on mean tract-level vulnerability index within each county.
Public Materials
Article only. The related public folder contains documentation and placeholders, but no matching analysis script, frozen input manifest, run record, or saved output that derives the 58-county and 9,039-tract results reported here.
Notes
[1] The article reports an index calculated for 9,039 California census tracts after applying density and population filters. The population, filters, and resulting index are not publicly reproduced. ↩
[2] The article states that the analysis was conducted in November 2025 using ACS, grocery-store, USDA SNAP-retailer, and Cal-ITP data. No frozen input manifest or run record is public. ↩
[3] Robinson, W. S. (1950). "Ecological correlations and the behavior of individuals." American Sociological Review, 15(3), 351-357. https://doi.org/10.2307/2087176. Classic paper on the ecological fallacy and interpretation of aggregate-level data. ↩
[4] Bureau of Economic Analysis. (2024). "Regional Price Parities by Metropolitan Statistical Area." Table MARPP. https://www.bea.gov/data/prices-inflation/regional-price-parities-state-and-metro-area. RPP measures relative price levels for all goods and services consumed, with 100 representing the national average. Non-metro California counties assigned state average RPP (112.6). ↩
Tags: #FoodSecurity #California #CountyAnalysis #VulnerabilityIndex #SpatialAnalysis #Methodology #Comparison
Next in this series: Building a residualized accessibility index that describes variation left unexplained by a stated structural model.
Suggested Citation
Cholette, V. (2025, November 23). County rankings and policy context: An article-reported case. Too Early To Say. https://tooearlytosay.com/research/methodology/county-comparison-methods/Copy citation