Kosali Simon’s account of the May 8, 2026 NBER Applications of AI in Healthcare panel quotes David Bradford, editor of the Wiley journal Health Economics, asking whether editors will need to focus more explicitly on whether a technically sound paper is important or interesting.1 This article uses well-executed but not important as shorthand for that editorial question. The panel account motivates the analysis; it does not establish that AI has already changed acceptance decisions.
An article’s importance is some combination of four things: the contribution’s marginal lift over what the field already knows; the practical reach of the insight for policy or clinical decision-making; how the finding relates to the field’s accumulated theoretical and empirical understanding; and whether the question is timely for the journal’s current portfolio. It is harder for an editor to explain than a missing identifying assumption, harder for a referee to translate into actionable suggestions, and harder for an author to learn from an email than from a conversation.
The published record can show what appeared and what later received citations. It cannot show rejected manuscripts, desk decisions, referee recommendations, editorial capacity, or the counterfactual publication path. Citation counts are an engagement proxy, not a direct measure of scientific or policy importance. The analysis below therefore describes associations inside a selected sample of published articles; it does not recover an editorial acceptance rule.
What journals have actually published
The article reports a corpus of 2,493 articles published from 2015 through 2022 across four health-economics field journals, assembled from Crossref metadata with abstract supplementation from PubMed, Semantic Scholar, and OpenAlex. The four journals are Journal of Health Economics (JHE, Elsevier, since 1982), Health Economics (HE, Wiley, since 1992), American Journal of Health Economics (AJHE, Chicago / ASHEcon, since 2015), and International Journal of Health Economics and Management (IJHEM, Springer, since 2014).2
Public-material status: The linked repository contains a methods note, derived CSV summaries, requirements, and selected classification and analysis scripts. It does not currently include the Crossref or PubMed acquisition scripts, the raw abstract corpus, figure-generation scripts, or a Makefile. The full acquisition and end-to-end analysis described in this article therefore are not one-step reproducible from the public repository.
We measure two things about each article. The first is substantive topic, by keyword match in title and abstract (managed care, maternal health, mental health, and so on). The second is contribution scope, classified by an LLM reading each abstract.
Topic shares changed over time
JHE has the longest archive; its topic-share changes are shown below.

Percentage-point change per decade in JHE topic share, 2000-2010 vs 2015-2025. Source: Crossref metadata for ISSN 0167-6296, 854 research articles in 2000-2010 and 1,051 in 2015-2025. Topics tagged by keyword presence in titles. Multiple topics per article are allowed.
The journal has shifted toward mental health, child health, maternal health, and COVID-related work, and toward coverage-policy topics (ACA, Medicaid). The retreat is from the topics that dominated the 1990s and early 2000s: insurance demand, managed care, and smoking. AI and machine learning is essentially absent from JHE titles so far.
A caveat on the topic measurement: publication dates follow submissions and editorial decisions with a lag. The cited studies document long publication processes in economics, but this corpus does not observe submission or acceptance dates and cannot assign a fixed 12-to-24-month lag to these four journals.34 Publication-year topic shares should therefore not be read as a contemporaneous measure of editorial decisions.
Across all four field journals in the most recent decade, the topical portfolios differ as follows:

Each panel shows a single journal’s top-five topics ranked by within-journal share of articles 2015-2025. Ranking rather than raw cross-journal levels makes the portfolios comparable even though abstract-availability differences across publishers affect detection rates. IJHEM is the most concentrated portfolio (top 3 topics are over 24 percent each). AJHE includes smoking/tobacco in its top five, which the others do not. Medicaid is in the top three at every journal except JHE, where it ranks fourth.
What kind of contribution
Topic identifies a paper’s subject. Scope identifies its contribution type. We can classify each abstract into one of three non-hierarchical bins based on this rubric:
| Scope | What the paper does | Signal phrases in the abstract |
|---|---|---|
| Calibration | Adds a credible estimate to an established literature. Applies a known method to a setting that fits a known template. Effect is incremental; results often consistent with prior literature. | “We apply X to Y,” “Extends to,” “Standard model,” “Consistent with prior literature,” “We estimate (existing parameter)” |
| Identification | Introduces a credible identifying strategy for a question prior work has measured noisily. The identifying variation is the lift. | “First credible quasi-experimental evidence on,” “Exploits [reform / natural experiment / IV],” “Difference-in-differences using [credible variation]” |
| Reframing | Opens or reshapes a line of inquiry. Combines causal design with new data, new method, or new question. Implications apply beyond the immediate setting. | “First credibly causal estimate of [important effect],” “Novel linkage of [data sources],” “Provides first multi-country evidence,” “Reshapes [established debate]” |
The labels are deliberately non-hierarchical. Calibration is not a lesser activity than Identification; field journals are supposed to host calibration work, and the cumulative empirical literature in any subfield rests on it. The labels name what the paper does without ranking its worth.
Three diagnostic questions a reader can apply to their own paper:
- Does the identifying variation exist before this paper?
If yes (an existing reform, an existing dataset, an existing instrument) and the paper applies it to one more setting, it is classified as Calibration. If no (the paper constructs the variation, links new data, or designs the experiment), it is classified as at least Identification.
- Does the answer change how someone in an adjacent subfield would frame their next paper?
If yes, at least Identification, possibly Reframing. If the answer adds to a known table of estimates, Calibration.
- Are the implications confined to one setting/policy/country, or do they generalize?
If implications generalize across health-economics subfields, and the design is credibly causal, and the data is novel or unusually well-suited, the paper is classified as Reframing.
The empirical distribution of scope across the four journals:

Share of 2015-2022 published articles by contribution scope, four health-economics field journals. n = 783 (JHE), 177 (AJHE), 1,358 (HE), 175 (IJHEM). Classification by an LLM reading the abstract; the abstract corpus is supplemented from PubMed, Semantic Scholar, and OpenAlex to reach 79-99 percent coverage per journal. Methodology writeup linked above.
JHE and AJHE have the largest shares labeled Identification in this classifier. HE is closer to an even split, and IJHEM has the largest Calibration share. Reframing labels are rare in every journal. These are abstract-based labels in the published sample, not verified measures of contribution quality or editorial intent.
Citation patterns in the published sample
Across the four journals over 2015-2022 (n=2,493), the raw citation distribution differs from the publication distribution. The table is descriptive and does not adjust its means or shares for citation exposure, although every included article has at least three years. Ratios use the underlying unrounded shares; the displayed percentages are rounded.
| Scope | Publication share | Citation share | Ratio | Mean cites |
|---|---|---|---|---|
| Calibration | 35% | 18% | 0.51x | 11 |
| Identification | 63% | 78% | 1.23x | 27 |
| Reframing | 2% | 5% | 2.14x | 46 |
Within this published sample, articles labeled Calibration make up a larger share of publications than citations, while articles labeled Reframing make up a larger share of citations than publications. Differences in article age, classification, unobserved quality, journal selection, and citation practices can all contribute to that pattern.
Holding journal, year, and 18 topic fixed effects constant in an OLS regression of log(citations + 1) on scope dummies (n=2,493, R² = 0.36, HC1-robust standard errors), the fitted coefficients are:
- The Identification label is associated with an exponentiated coefficient of 1.91 relative to Calibration. (95% CI [1.60, 2.27], p<0.001)
- The Reframing label is associated with an exponentiated coefficient of 2.26 relative to Calibration. (95% CI [1.60, 3.20], p<0.001)
Because the outcome is log(citations + 1), these exponentiated coefficients describe multiplicative differences on the citations + 1 geometric scale. They do not mean 91 percent or 126 percent higher arithmetic mean citations. The regression is observational, conditional on its included fixed effects, and does not show that a scope label causes citations. Abstract rhetoric may also affect both the classifier label and later visibility.
The 35-percent publication share and 18-percent citation share for Calibration is a descriptive engagement gap. It is not evidence that those papers were unimportant to editors, technically adequate, or over-published. Those claims require submission and decision data that this study does not observe.
What this can say about a rising methods floor
The cited studies document different pieces of a possible capacity story: submission growth at one management journal, detected AI-assisted writing across a broad journal sample, pandemic-era submission pressure in agricultural economics, a simulation of business-school publishing, and concerns in health publishing.56789 They do not estimate a common editorial capacity threshold, a response for these four health-economics journals, or the causal effect of AI on acceptance. The claim that a system can absorb 30 percent but not a doubling is not supported by this evidence and is not used here.
The defensible implication is a research agenda rather than a forecast. If AI lowers the cost of preparing technically competent manuscripts and raises submissions, which manuscript types increase? Do desk-rejection criteria change? Do review times, revision rounds, or acceptance rates move? Answering those questions requires journal-level submission and decision records before and after adoption, not publication and citation data alone.
The scope regression reports one descriptive association: published abstracts labeled Identification or Reframing are associated with higher citations + 1 on the fitted multiplicative scale. It does not tell authors to imitate “first credible” phrasing, because the classifier may be measuring rhetoric as well as substance.
In the JHE title corpus, the AI and machine-learning keyword tag changes by +0.08 percentage points per decade from a near-zero base. Without submission and acceptance dates, the published series cannot determine whether more AI papers are currently in an editorial pipeline or predict their future share.
Questions for training the next generation of economists
PhD training already asks students to defend identifying assumptions, run robustness analyses, and build replication packages. If AI makes some drafting and coding tasks cheaper, programs may need to make question selection and contribution assessment more explicit as well. This corpus does not measure training, skill abundance, or hiring and publication criteria. It motivates the question rather than answering it.
The open training question is how programs teach question selection alongside identification and execution. If technically competent drafts become cheaper to produce, students may face more competition on contribution and relevance. This article has no acceptance data with which to establish an old or new publishing equilibrium, so that proposition remains a hypothesis to test.
A conditional opportunity
If AI lowers some drafting and coding costs, researchers can choose to spend the saved time on question selection, design, and falsification. Our companion piece treats faster rejection of weak hypotheses as one possible use of that capacity. Whether this changes manuscript quality or editorial decisions is an empirical question, not a settled consequence of AI adoption.
The analysis supports one practical discipline for applied economists: state the contribution and its policy relevance explicitly, then support both with the design and evidence. The citation analysis describes the published record; it does not supply a formula for acceptance or importance.
Limitations
The most consequential limitation remains unresolved: the classifier cannot cleanly separate papers that do identifying work from papers whose abstracts use identifying-work conventions. Abstracts using “first credibly causal” or “first to” framing are classified as Identification 88 percent of the time, while Calibration-labeled abstracts use that framing at a 1.5 percent rate. Without a blinded human validation sample and an error model, the direction and magnitude of misclassification bias in the scope coefficients remain unidentified.
Field journals are one tier of the health-economics publication market. Top contributions are often submitted to or published in general-interest journals (AER, QJE, JPE, ReStud, AEJ:Applied/Policy) before or instead of field journals. The scarcity of Reframing papers (2 to 5 percent across the four journals) and the absence of the “Foundational” tier in our classification reflect selection into the field-journal market specifically, distinct from the field’s overall production of path-opening work.
Citation count is an imperfect measure of engagement and an even less direct measure of importance. A controversial paper can accumulate citations for criticism, and recent papers have had less time to accumulate citations. Year fixed effects adjust the regression for cohort-level differences, but the raw share table is not exposure-adjusted and within-year exposure still varies. The article reports that median quantile regression and trimming the top 1 percent preserve the coefficient signs and significance while attenuating magnitudes by 20-30 percent. Those checks do not convert the associations into causal effects or editorial criteria.
The public METHODS.md documents the reported acquisition approach, classifier specification, regression specification, and limitations. It can be used to inspect the stated methods, but it is not a substitute for the missing acquisition scripts, raw corpus, fitted outputs, and figure-generation code.
Public Materials
The journal-topic-shares repository currently provides derived CSV summaries and selected scripts for classification, merging, and analysis. It also includes fetch scripts for Semantic Scholar and OpenAlex. It does not provide Crossref or PubMed fetch scripts, a complete raw or processed article-level corpus, figure-generation scripts, or a one-command build target. The posted files support inspection of parts of the workflow; they do not presently reproduce every figure and estimate in this article from source acquisition through final output.
References
- Simon, K. (2026, May 18). Practical advice for using AI in health (& other) economics research: Summary of NBER panel from May 8th 2026. Frankly, the counterfactual was worse (Substack). Writeup of moderated panel at NBER Applications of AI in Healthcare meeting, Cambridge, MA, May 8, 2026; panelists Kosali Simon, Scott Cunningham, David Bradford, and Coady Wing. https://franklythecounterfactual.substack.com/p/practical-advice-for-using-ai-in
- Crossref. (2026). Crossref REST API and journal metadata. https://api.crossref.org/journals/. Retrieved May 26, 2026 (n = 3,020 research articles in ISSN 0167-6296, 1982-2026; 4,098 in ISSN 1057-9230, 1992-2026; 296 in ISSN 2332-3493, 2015-2026; 253 in ISSN 2199-9023, 2014-2026). Abstract supplementation via NCBI E-utilities (eutils.ncbi.nlm.nih.gov), Semantic Scholar Graph API (api.semanticscholar.org), and OpenAlex (api.openalex.org), retrieved May 25-26, 2026.
- Ellison, G. (2002). The slowdown of the economics publishing process. Journal of Political Economy, 110(5), 947-993. https://doi.org/10.1086/341868
- Card, D., & DellaVigna, S. (2013). Nine facts about top journals in economics. Journal of Economic Literature, 51(1), 144-161. https://doi.org/10.1257/jel.51.1.144
- Gartenberg, C., Hasan, S., Murray, A., & Pierce, L. (2026). More versus better: Artificial intelligence, incentives, and the emerging crisis in peer review. Organization Science, 37(3). https://doi.org/10.1287/orsc.2026.ed.v37.n3
- He, Y., & Bu, Y. (2026). Academic journals’ AI policies fail to curb the surge in AI-assisted academic writing. Proceedings of the National Academy of Sciences, 123(9), e2526734123. https://doi.org/10.1073/pnas.2526734123
- Biondi, B., Barrett, C. B., Mazzocchi, M., Ando, A., Harvey, D., & Mallory, M. (2021). Journal submissions, review and editorial decision patterns during initial COVID-19 restrictions. Food Policy, 105, 102167. https://doi.org/10.1016/j.foodpol.2021.102167
- Jiang, S. (2025). Tenure under pressure: Simulating the disruptive effects of AI on academic publishing. arXiv preprint arXiv:2509.16925. https://arxiv.org/abs/2509.16925
- Arzilli, G., Di Maggio, E., De Angelis, L., Baglivo, F., Savoia, E., Privitera, G. P., & Rizzo, C. (2025). A surge of AI-driven publications: The impact on health professionals and potential mitigating solutions. Frontiers in Public Health, 13, 1680630. https://doi.org/10.3389/fpubh.2025.1680630
Cite this article
Cholette, V. (2026, May 25). Well-executed but not important. Too Early To Say. https://tooearlytosay.com/research/methodology/well-executed-but-not-important/