The USDA keeps a database of every store in the country that accepts SNAP benefits. In California alone, about 28,800 stores appear in this list. Researchers use this data all the time: the USDA's main report on food deserts has been cited in over 480 studies, and this data feeds into health rankings for nearly all 3,143 U.S. counties. The County Health Rankings program uses it as the only source for "Limited Access to Healthy Foods" in about 3,100 areas. A 2020 study called the Food Access Research Atlas "the most complete food environment tool" that is "widely used by policy makers and researchers."
In national studies, some errors in the data are expected. As long as the errors are random, they balance out across the whole country. But when you study a single county, every gap matters. One missing store could make a neighborhood look like a food desert when it isn't.
As part of our series on food security in Silicon Valley, we needed to look at food access. Starting with the USDA made sense. But the quality of the data varies a lot across counties. In one county, the database listed 67% more stores than actually exist. In another, it only captured 65% of actual stores. Same state, same data sources, same methods: a 2.6× gap in data quality that nobody warns you about.
This matters because retail-listing errors can change which neighborhoods are classified as low access. A 2013 South Carolina study compared USDA and CDC low-access measures built from two commercial store listings with measures built from an eight-county field census. The resulting tract classifications varied with the data source (Ma et al., 2013). That study did not audit the USDA SNAP retailer database, so it supports the general measurement-risk point rather than the California match rates reported on this page.
This raises a question: how close can we get to a real answer when studying local areas? To explore EBT acceptance across California, we can use a three-step method that combines USDA records, Google Maps data, and chain store logic. The article reports a test across 7 counties: Santa Clara, Sacramento, San Francisco, San Diego, Contra Costa, Orange, and Alameda. Those exact results remain under reconciliation.
How Much Does USDA Data Quality Vary?
Let's start by comparing the USDA database to what's actually on the ground. We used Google Maps as our "ground truth" for existing stores, then matched those against USDA's list of SNAP retailers. The results were striking.
| Metric | Sacramento County | Contra Costa County |
|---|---|---|
| Actual stores (Google Maps) | 627 | 927 |
| USDA SNAP retailers | 1,047 | 600 |
| USDA/Google ratio | 1.67 | 0.65 |
| Match rate | 81.8% | 66.3% |
The article reports a Sacramento ratio of 1.67 and a Contra Costa ratio of 0.65, a 2.6× range. These counts and ratios are not currently reproduced by a public output.
It also reports an association between stores per 100,000 people and the USDA/Google ratio of r = -0.828 (p < 0.05). With seven counties and no matching public output, treat that estimate as provisional.
Counties with more stores have worse USDA data.
This is the "Retail Density Paradox." Dense urban areas with lots of food stores are harder to verify using USDA data alone. Rural areas with fewer stores have more complete USDA records.
The catch: High match rates may show good USDA data quality, not good food access. Comparing match rates across regions without knowing the USDA/Google ratio can lead to wrong conclusions.
Three-Step Method
Given these data quality problems, a method for checking EBT acceptance that doesn't rely only on the USDA database would be useful. Here's a three-step approach where each step adds checking power but also adds work.
Step 1: Match Stores to USDA Records
The first step is simple: take actual stores and try to match them to the USDA database. Match on both location (are they close enough?) and name (do the store names look alike?).
Why 200 meters? The distance limit needs to catch real matches (the same store with slightly different map pins) without catching false ones (two different stores near each other). The article reports that tests of 100m, 200m, and 300m favored 200m while keeping false matches under 2%. That threshold and error rate remain under reconciliation.
Why 50% name match? Store names in different databases rarely match exactly. The article proposes a 50% name threshold to catch common variations, but the threshold and example scores have not been reproduced publicly.
The combined score: The article proposes weighting name match at 60% and distance at 40%, with at least 70% required to call a match. These settings are provisional until a reproducible validation output is posted.
What to expect: Step 1 alone matches 18-50% of stores, depending on the county's USDA/Google ratio. Sacramento (ratio 1.67) hits 50% from just this step. Contra Costa (ratio 0.65) only reaches 18%. Low Step 1 rates aren't a method problem; they're a data quality signal telling you the USDA database is missing stores in your area.
Step 2: Use Chain Names to Generate Candidates
A chain match is a clue for follow-up, not verification. USDA states that a SNAP permit is valid only for the location on record and that each store location must be separately authorized. One Safeway in the retailer file therefore cannot verify a different Safeway location.
What shared infrastructure can tell us: Large chains may share point-of-sale equipment, processors, backend systems, and training. That can make a chain name useful for prioritizing checks, but it does not replace a location-level permit or a current match in the USDA retailer file.
What chain-name logic looks like: Normalize variants such as "Safeway," "Safeway #1234," and "Safeway Store" to a common candidate name. Then check each location against a store-level authorization record. Locations found only through the name rule remain candidates rather than verified stores.
Starting with about 50 major chains (Walmart, Target, Kroger stores, Albertsons stores, etc.) and growing to over 150 as regional chains emerge works well. Each entry captures name variations: "7-Eleven" and "7-11" and "Seven Eleven" all map to the same chain.
Something to watch for: Looking at chain matches within a single county will miss regional chains that have no USDA records in that specific area.
The article uses ExtraMile as an example: it reports 50 Santa Clara locations with no county match and 15 matches elsewhere in the seven-county file. That pattern can identify locations for follow-up, but it does not establish that all 50 Santa Clara locations were authorized.
The fix: Build a regional candidate-name list, then run location-level authorization checks in each county. The article reports that its broader name rule reclassified 303 stores (+4.3 points) and changed Santa Clara from 59% to 70%; those figures are under reconciliation and should be interpreted as provisional classifications, not verified permits.
What the article reports across counties: Steps 1 and 2 classify 66-82% of stores. The figures are not publicly reproduced, and the chain-name portion requires location-level confirmation before it can be called verification.
Step 3: Manual Checking (Optional)
After Steps 1 and 2, 20-30% of stores remain marked "unknown." These are mostly small stores that don't belong to any chain. What to do with them?
One option is manual checking: actually look up whether each store accepts EBT. A few approaches:
- Website checking: 2–3 minutes per store, but only about 40% of small stores have websites with payment info
- Photo checking: Google Maps street views sometimes show "We Accept EBT" signs
- Phone checking: Calling stores works but takes 5–10 minutes per store
A pilot in San Francisco with 40 small stores found that 67% of "unknown" stores actually do accept EBT. They just don't appear in the USDA database.
When Step 3 is worth it:
- Your area has a USDA/Google ratio below 0.8 (meaning the database is really lacking)
- You're in an urban area with many small grocers
- Your analysis is high-stakes and you need exact numbers, not ranges
What to Check Before You Start
Before diving into this process, it helps to know what you're working with. Four metrics tell you which approach makes sense and what results to expect.
1. USDA/Google Ratio (Data Quality)
This is the most important number. Divide the count of USDA SNAP retailers in your area by the count of food stores you find via Google Maps.
A ratio above 1.0 means USDA has more records than actual stores (some may be closed or duplicates). A ratio below 1.0 means the database is missing stores. Given the 2.6× range across California counties, don't assume your area looks like anyone else's.
2. Store Density (How Hard It Will Be)
Count how many food stores exist per 100,000 people. This tells you about your retail landscape:
- High density (>70 per 100k): Expect more small stores, ethnic grocers, and specialty shops. Chain logic won't carry you as far. This will be harder.
- Low density (<50 per 100k): Chains likely dominate. Walmart, Safeway, and Kroger stores are probably most of your list. Chain logic will get you most of the way there.
3. Border Check (Data Pollution)
The first San Diego query included stores in Mexico because the search geometry crossed the international border. The draft previously gave conflicting shares for that contamination, so the exact count and percentage are withdrawn pending a matched output.
The method should clip results to the study boundary before matching. The article reports that geographic filtering increased the San Diego match rate, but the exact before-and-after values are under reconciliation.
If your area borders another country: Check for this first. It will save you days of confusion.
4. Chain Presence (Method Fit)
What share of your stores belong to chains you can spot? This sets how far Step 2 will get you.
- High chain presence (>60%): Chain logic can prioritize likely matches and reduce the manual-review queue. It does not verify an unmatched location; location-level authorization still needs to be established.
- Low chain presence (<40%): Many small stores means many unknowns. Plan time for manual checking if you need exact numbers.
Things That Went Wrong
Research rarely goes as planned. Let's take a look at some of those moments in this process.
1. Missing Regional Chains
Building the chain list while working in Northern California centered on familiar names to this area: Safeway, Lucky, Raley's, Trader Joe's. For southern counties, we were missing regional chains like:
- Vons (42 stores): A major Albertsons brand that doesn't exist in NorCal
- Albertsons (33 stores): Uses different branding down south
- Northgate González (23 stores): A Latino grocery chain focused in SoCal
That's 98 stores from just three chains, all marked "unknown" because we hadn't added them to our list.
The lesson: Your chain list is local knowledge. When you move to a new area, look at your most common "unknown" store names. If the same name appears 20+ times, it's probably a chain you haven't added yet.
2. Including Stores That Can't Accept SNAP
The article reports 32 BevMo! locations with no USDA matches. Alcohol cannot be purchased with SNAP, but an alcohol-focused store name alone does not prove that a location is ineligible; store eligibility depends on its qualifying food inventory or sales and its authorization status. These locations should remain unmatched until a store-level check resolves them.
Including such stores makes match rates look worse than they should be.
The fix: Check eligibility and authorization at the location level before excluding a store. Names associated with alcohol, tobacco, or pet retail can flag records for follow-up, but they cannot establish ineligibility on their own. The article reports that a name-based screen flagged 337 stores (5%); without location-level checks, that count is a review queue, not a verified exclusion set.
3. Checking Franchise Logic
Ownership structure does not remove the location-level authorization requirement. A franchise may share payment infrastructure with a national brand, but each store still needs its own SNAP authorization.
The article reports that multiple 7-Eleven locations appeared in its USDA matches. That observation can support a candidate-name rule; it cannot assign authorization to unmatched locations or justify unsourced confidence percentages.
What We Learned
The article reports applying this method to 7,023 stores in 7 California counties. The sample and the results below remain under reconciliation.
1. USDA Data Quality Varies 2.6× Across Regions
This is the main finding. Don't assume the USDA database is complete (or incomplete) everywhere. Sacramento's 1.67 ratio and Contra Costa's 0.65 ratio are totally different data quality stories, and there's no way to know which you're dealing with until you check.
What this means for research: Match rates can't be compared across studies without knowing the data quality behind them. A study reporting "80% verified" in one area and "60% verified" in another might be measuring database gaps, not real differences in food access.
What this means for food access findings: If you're using USDA data to spot food deserts or measure EBT access, your findings are only as good as your local data. In a county with a 0.65 ratio, the USDA database is missing about a third of actual stores. If you figure "distance to nearest EBT store" using only USDA data, you'll think distances are farther than they are because you're missing nearby stores. A neighborhood might look like a food desert when it actually has a small grocery around the corner that USDA doesn't know about.
On the flip side, in a county with a 1.67 ratio, USDA lists more stores than actually exist. Some of these are likely closed shops that haven't been removed from the database. If you count EBT stores per person, you'll count too many. A neighborhood might look well-served when several of those "stores" no longer exist.
2. The Retail Density Paradox
It is reasonable to expect counties with more stores to be easier to verify: more data, more matches. The opposite is true.
The article reports a negative association between food-store density and its USDA coverage measure (r = -0.828). The small county sample and missing reproducible output do not support calling that relationship established.
Why this matters for policy: A high match rate might show good data quality (like Sacramento), not good food access. Reading match rates without knowing this paradox can lead to backward conclusions.
A concrete example: Sacramento County has 41.5 food stores per 100,000 people and an 81.8% match rate. Contra Costa County has 82.9 stores per 100,000 people and a 66.3% match rate. If you looked only at match rates, you might think Sacramento has better EBT access. But Sacramento has half as many stores per person. The high match rate reflects USDA's good coverage of a sparse retail landscape, not lots of food access. Contra Costa actually has twice the store density; we just can't verify as many because the database hasn't kept up with a more active market.
3. Cross-County Names Expand the Candidate List
The article reports that a regional name list reclassified 303 stores and added 4.3 points to its provisional match rate. Those classifications are under reconciliation and require location-level authorization checks.
The practical point: If you're working with multiple counties or areas, pool your chain evidence first. Build the list regionally, then apply it everywhere. Small chains with regional presence are invisible when you analyze counties one at a time.
4. Border Areas Need Special Steps
The initial San Diego query included Mexican stores because its search area crossed the border. Earlier drafts gave inconsistent percentages for that contamination, so the exact share is withdrawn until the query output and study boundary are reconciled.
If your area borders another country: Set up geographic filtering before anything else. The few minutes of setup will save you days of puzzling over weird results.
Limits
Let's check in about what this method can and can't do.
What it does:
- Steps 1 and 2 classify 66–82% of stores; locations found only by chain-name logic remain candidates until matched to a location-level SNAP authorization record.
- Works well in suburban areas with lots of chain stores
- Gives you a way to understand your specific area's data quality
What it doesn't fix:
- 20–30% of stores stay "unknown" after automated steps
- Manual checking (Step 3) takes a lot of time
- Chain-name logic prioritizes follow-up but cannot establish SNAP authorization for a different location.
- Results depend heavily on your area's USDA data quality, which you can't control
What remains uncertain:
- Of the 29.1% unknown stores in this data, the San Francisco pilot suggests about 65% actually accept EBT. But that's one pilot in one city.
- Some unknowns are new stores that haven't made it into the USDA database yet.
- Some are chains not yet identified because they're regional to areas not studied.
Reporting ranges and owning the uncertainty is more honest than picking a single number. The article-reported 66–82% range is a provisional classification range, not a verified-store range, because the chain-name portion still requires location-level confirmation.
Method Notes
Data Sources:
- Google Maps API for actual store locations
- USDA SNAP Retailer Locator for authorized retailers
- US Census shapefiles for county borders
- ACS data for population (store density math)
Article-reported settings:
- 200m distance limit: Tested against typical urban store spacing
- 50% name match: Handles "Safeway" vs. "Safeway #1234"
- 70% combined score: Balances catching true matches vs. false ones
- Cross-county chain logic: +303 stores, +4.3 points
Article-reported sample: 7,023 stores across 7 California counties; not publicly reproduced.
References
Cates, S., Cortés, A., Guthrie, J., Gupta, S., Jayaraman, A., & Yeh, M. A. (2019). Scanner Capability Assessment of SNAP Authorized Small Retailers. U.S. Department of Agriculture, Food and Nutrition Service.
Ma, X., Battersby, S. E., Bell, B. A., Hibbert, J. D., Barnes, T. L., & Liese, A. D. (2013). Variation in low food access areas due to data source inaccuracies. Applied Geography, 45, 131–137. https://doi.org/10.1016/j.apgeog.2013.08.014.
Public Policy Institute of California. (2024). California's Nutrition Safety Net.
USDA Economic Research Service. (2021). Food Access Research Atlas Documentation.
USDA Food and Nutrition Service. SNAP Retailer Notice: Permits. Each store location must be separately authorized.
Related materials: pinned GitHub repository folder. This folder does not derive or validate the seven-county results reported in this article.
Suggested Citation
Cholette, V. (2025, October 12). The retail density paradox: Why more stores mean worse data. Too Early To Say. https://tooearlytosay.com/research/methodology/ebt-verification-methodology/Copy citation