What We Mistake for AI Capability

A hypothesis about how task tolerance and specification shape perceived AI capability, with a controlled comparison still to run.

AI-generated slides can be impressive. They may reduce formatting work and help distill information into a sequence. Before attributing that output to deep model understanding, we need to separate at least three factors: model capability, the precision of the specification, and how much variation the task permits.

A model’s training data may include good and bad decks, conference talks, and corporate pitches. Training data may encode patterns associated with effective presentations. But rhetoric depends on context. The model lacks task-specific information about whether this deck is for methodologists who want coefficient plots or policymakers who need one takeaway per slide, whether the talk is 12 minutes or an hour, or whether the goal is to teach, persuade, or present a job-market paper. That context comes from the prompt, not the training data.

Without task-specific specification, a model can generate output that satisfies generic presentation conventions without being tailored to the local audience or purpose. When the context is stable, that generic output may appear tailored because the requirements do not change. The missing specification becomes evident when the context changes and the prior output no longer meets the task requirements.

The Variation Tolerance Test

Our impression of how capable an AI model is at a task depends on how much variation we tolerate in the output.

When the acceptable range is wide (a slide deck that looks professional, follows field conventions, uses readable fonts), the model can satisfy the acceptance criteria across many outputs. We may call this “good output” and credit the model.

When the acceptable range is narrow, failures can become easier to see. Consider an economics diagram where indifference curves must be convex, the budget constraint must pivot correctly, and the production-possibility frontier must bow outward with a specified curvature. An uncited LinkedIn anecdote described an economist supplying examples, rules, and project context without obtaining a convincing result. The model, version, prompt, scoring rule, and outputs are unavailable here, so the anecdote motivates a test rather than supplying evidence.

These examples do not hold the model, prompt, or task constant, so they cannot identify why performance differs. The proposed mechanism is that a wider acceptable set makes success easier to observe, while a narrower target makes consequential errors easier to detect.

The testable hypothesis is that assessments of AI capability partly reflect a task’s tolerance for variation. A task with many acceptable outputs may receive higher ratings than one with a narrow answer set, even when both use the same frozen model.

That pattern could reflect wider standards, model capability, specification quality, or an interaction among them. A controlled design is needed to separate the explanations.

What We Are Actually Measuring

When someone says “the model is amazing at slides” and someone else says “the model is terrible at economics diagrams,” both may be accurately describing their outputs. Neither observation by itself identifies whether the cause is the model, the specification, the scoring rule, or the task.

This matters because the mistake (crediting the model’s “deep understanding” when the task was actually underspecified) prevents us from asking the more useful questions. What do we actually know about our audience? What level of technical detail serves this room? What should each slide accomplish? What teaching purpose does this diagram serve, and what should the student see first?

These are questions a production workflow can neglect. Practitioners sometimes report that generative AI reduces formatting time, but this article does not measure that change. If a workflow does save execution time, one use of the recovered capacity is upstream design: What makes a presentation effective in a specific context? What makes a diagram clear for teaching? What makes any visual work for its audience?

The model can also help researchers work through those questions. We could use it to reflect on our audience, to articulate what makes a diagram work for teaching trade theory versus presenting a research result. We could use it to generate more precise specifications, if we use it for that instead of assuming it already knows the answer.

The Specification Is the Skill

The hypothesis may extend beyond slides and diagrams: task tolerance, specification precision, and model capability can jointly shape an evaluator’s judgment. The present article does not estimate their separate contributions.

Specification is one skill worth testing: can a researcher state the target precisely enough that the output satisfies the required criteria? That requires articulating standards that may previously have remained tacit during manual production. We have called this the copy-paste ceiling: at some point, the workflow demands more than accepting model output. It demands knowing what we want and why.

Michael Polanyi had a phrase for this: “we can know more than we can tell.” He called it tacit knowledge in 1966. When a workflow requires instructions for a model, some of that knowledge must be made explicit. If production costs fall, specification can become a larger share of the remaining work. The linked From Methodology to Code article reports one instructional example, but it does not provide a controlled estimate of time savings or isolate specification precision as the cause.

The Experiment This Claim Needs

A controlled test would freeze the model version and task inputs, cross high and low specification detail with broad and narrow scoring tolerances, and score outputs against a preregistered answer key. Repeating that factorial design across slides, diagrams, prose, and code would estimate the contribution of specification, tolerance, task family, and their interactions. Blinded evaluators and saved prompts, outputs, runtimes, and corrections would make the result inspectable.

Until that comparison is run, the practical contribution is a design proposition: treat “the model is capable” as a claim that needs a task definition, an acceptance region, and a comparison condition.

Whether this makes AI more useful or less depends on whether we recognize the shift. Too early to say.

For a deeper look at how this applies in empirical research, where specification precision can affect whether agents produce reliable code or plausible-looking errors, see What Agents Actually Do (And What They Don’t).

Suggested Citation

Cholette, V. (2026, March 2). What we mistake for AI capability. Too Early To Say. https://tooearlytosay.com/research/methodology/what-we-mistake-for-ai-capability/
Copy citation