The article reports a 22-fold range in Santa Clara County crime rates, from 23 per 100,000 in Los Altos Hills to 498 in San Jose, compared with a county value of 381. It also reports different employment estimates at county and PUMA levels. These results are retained as a geographic-measurement case, but they have not been publicly reproduced.
This is not a California quirk. Jacob Kaplan, who maintains the authoritative guide to FBI crime data, is direct: "County-level UCR data should not be used for research." [1] The infrastructure fails in three ways: multi-county agencies distribute crimes by population rather than location, imputation adds noise, and the entire system rests on an assumption of uniform crime distribution within jurisdictions. Crime does not distribute uniformly.
Does Precision Change the Answer?
Chalfin and McCrary found that measurement error in police-employment data made older estimates of the effect of police on crime too small by about a factor of five [2] . That result illustrates attenuation from a mismeasured regressor, but it does not quantify the error in the geographic crime measure used here. Geographic aggregation introduces a different problem: averaging can conceal local variation and can change an estimated association.
To examine that sensitivity, we can allocate jurisdiction-reported crime to PUMAs and compare it with a county-level measure. This creates an approximation of place-associated exposure, not a direct measure of crime at each resident's address.
Mapping Crime to Community
Here is the problem: crime data comes from police jurisdictions, but Census microdata locates people in PUMAs (Public Use Microdata Areas, geographic units of about 100,000 people). These geographies do not align. Berkeley PD reports crime for Berkeley. The Census reports employment for PUMA 101. How do we connect them?
We start with shapefiles. The Census publishes geometric boundaries for both California cities and PUMAs. We overlay them and calculate exactly how much of each city falls in each PUMA.
Take Berkeley. The article reports that its geometric intersection assigns 72% of the city to PUMA 101 and 28% to PUMA 114, then allocates reported crime using those shares. The split is illustrative and has not been publicly reproduced.
The article reports applying this approach to 462 police agencies and 275 PUMAs. The method is the reusable lesson; those counts remain provisional.
The Precision Test
Same population. Same outcomes. Same controls. Different crime measures.
Setup:
Evidence note: the sample size, years, estimates, standard errors, and p-values below are article-reported and not publicly reproduced.
- 817,000 working-age California adults (2018-2022)
- Outcomes: employment, labor force participation, hours worked
- Fixed effects: county + year
County-Level Crime (Standard Approach):
| Outcome | Coefficient (crime scale undocumented) | SE | p-value |
|---|---|---|---|
| Employment | -0.005 | (0.004) | 0.21 |
| Labor Force | -0.003 | (0.002) | 0.13 |
| Hours | -0.25 | (0.13) | 0.06 |
In the article-reported county-level specification, each p-value exceeds 0.05. That is a null result for this specification, not evidence that crime has no labor-market effect.
PUMA-Level Crime (Spatial Crosswalk):
| Outcome | Coefficient (crime scale undocumented) | SE | p-value |
|---|---|---|---|
| Employment | +0.005 | (0.002) | 0.01 |
| Labor Force | +0.003 | (0.002) | 0.05 |
| Hours | +0.23 | (0.10) | 0.02 |
The article reports p-values at or below 0.05 in the PUMA-level specification. But the archived article does not state whether crime entered as a rate per 100, per 1,000, per 100,000, a standardized value, or a transformation such as a logarithm. The coefficient scale is therefore not interpretable from the page. The contrast is useful for teaching how geography can change an estimate, but it is not a publicly reproduced result.
Unmasking the Variation
Two mechanisms:
Less aggregation. In the article-reported example, a county average collapses a reported 22-fold range into one number. More granular measures can preserve within-county variation, though whether they reduce total measurement error must be tested.
Different identifying variation. With PUMA crime and county fixed effects, the coefficient uses within-county differences as well as changes over time. County fixed effects remove time-invariant county characteristics, and year fixed effects remove shocks common to all counties. They do not remove county-specific shocks that change over time, and adjacent PUMAs need not share the same local conditions. The remaining variation can therefore include both neighborhood crime exposure and unmeasured local factors.
This design uses more granular identifying variation than the county-level specification. It does not, by itself, resolve omitted-variable bias or establish causality.
Density, Not Causality
The article reports coefficients of +0.005 for employment and +0.23 for weekly hours. If employment is coded from zero to one, +0.005 corresponds to 0.5 percentage points for a one-unit change in the crime regressor. Because that crime unit is not documented, neither coefficient has a defensible substantive interpretation from the page alone. The values are not publicly reproduced and should not be interpreted causally.
One plausible explanation for a positive coefficient is omitted-variable bias, including economic density. Vibrant economic hubs can generate both legal employment opportunities and opportunities for crime. This article does not test that mechanism.
The methodological insight is narrower: changing geographic resolution can reveal associations that aggregation may obscure. Whether it does so in this case remains provisional until the results are reproduced.
The Shift Toward Granularity
The economics of crime has moved toward granular data. Hjalmarsson, Machin, and Pinotti document the shift "from aggregated data (e.g., at the US state level) to highly disaggregated data." [3] Ihlanfeldt's Atlanta work used census tracts to identify neighborhood effects that metro analysis missed [4] .
For crime-labor research specifically, the article did not document a systematic, current literature search sufficient to establish that no published study uses PUMA-level crime with labor outcomes. Treat that as a research question, not a verified novelty claim.
Limits of the Data
Spatial allocation is imperfect. This method assumes uniform crime within cities. Reality is more concentrated.
PUMAs are still large. About 100,000 people each. Tract-level would be more precise, but ACS microdata does not identify tracts.
Correlation is not causation. County fixed effects help, but time-varying confounders remain possible. The positive sign suggests omitted variables or selection rather than a causal benefit of crime.
Why This Matters for Policy
For research: Null findings in county-level crime studies might reflect measurement choices. Before concluding "no effect," researchers should test whether geographic aggregation changes the estimate.
For policy: This article-reported case is not sufficient evidence that neighborhood crime-reduction programs have underestimated labor-market benefits. It identifies a measurement question that a reproducible, causal design would need to test.
For data infrastructure: A spatial crosswalk can be implemented with tools such as pygris and GeoPandas, but the allocation rule still needs validation against agency service areas and within-city crime location. Publishing those choices would make geographic sensitivity easier to audit.
Public Materials
Article only. The related public folder contains documentation and placeholders, but no matching analysis script, frozen input manifest, run record, or saved output that derives the geographic crosswalk, descriptive range, sample, or regression results reported here.
Technical Notes
| Element | Detail |
|---|---|
| Crime data | CA DOJ CJSC, 2018-2022 |
| Labor data | ACS 5-year PUMS |
| Sample | 817,000 adults ages 25-64 |
| Crosswalk | 462 agencies to 275 PUMAs via TIGER/Line spatial join |
| Estimation | WLS, clustered SEs at PUMA |
-
Kaplan, J. (2024). Uniform Crime Reporting (UCR) Program Data: A Practitioner's Guide. Chapter 10: County-Level UCR Data. https://ucrbook.com/county-level-ucr-data.html ↩︎
-
Chalfin, A. & McCrary, J. (2018). Are U.S. Cities Underpoliced? Theory and Evidence. Review of Economics and Statistics, 100(1), 167-186. Author-hosted paper. ↩︎
-
Hjalmarsson, R., Machin, S., & Pinotti, P. (2024). Crime and the Labor Market. In C. Dustmann & T. Lemieux (Eds.), Handbook of Labor Economics (Vol. 5, pp. 679-759). Elsevier. ↩︎
-
Ihlanfeldt, K. (2007). Neighborhood Drug Crime and Young Males' Job Accessibility. Review of Economics and Statistics, 89(1), 151-164. ↩︎
Suggested Citation
Cholette, V. (2025, December 28). Crime geography precision: An article-reported case. Too Early To Say. https://tooearlytosay.com/research/methodology/crime-geography-precision/Copy citation