3625 words
18 minutes
The Prompt Specificity Ladder: How Much Detail Your Coding Agent Actually Needs

Say you point a coding agent (a tool that reads a prompt, writes code, and often runs that code itself) at an API (the fixed set of functions a piece of software exposes for other code to call). If it’s an API you already know well, the prompt barely matters. Now point it at one you don’t know: a vendor’s toolkit, an internal library, an API you installed an hour ago. A real question shows up. How much do you actually have to tell it? The instinct that feels safe is to write more. Name the class, the parameters, the file layout, every edge case. More detail should mean fewer wrong guesses. NVIDIA (the chip maker) tested that instinct against a real coding toolkit, and the answer it got back doesn’t match the instinct at all [1].

Figure 1 - Diagram of five prompt-detail levels rising left to right, with a flat line for correctness overlaid on a climbing line for token cost and code length

Figure 1 - What Climbs and What Doesn’t: Five levels of prompt detail, Sketch through Contract. Token cost and generated code length climb across them, faster in the later rungs than the earlier ones. Correctness, NVIDIA’s code-correctness check rather than a physical-accuracy check, stays flat at every single level.


What NVIDIA actually tested#

NVIDIA builds ALCHEMI (a toolkit for running materials-science simulations). A simulation like this is a computer model that predicts how a substance behaves, such as how fast its atoms move or how it responds to heat, before anyone runs the real experiment in a lab [1]. Like most modern coding toolkits, ALCHEMI ships its own agent skill. That’s a folder of instructions and worked examples the agent pulls in on its own when a task calls for it. It lives in a public GitHub repository under an Apache 2.0 license [2]. Does a bundled skill like that actually help, or does it just add tokens (the units a model reads and writes, and gets billed for) to every run? That’s a separate question, and we cover it in Don’t Ship Skills Without Evals [3]. For this study the point is simpler. Whatever the prompt said, the agent could already reach the toolkit’s own patterns through that skill.

NVIDIA generated 45 pipelines: three simulation workflows, five levels of prompt detail, and three attempts at each combination. Every one used Claude Opus 4.8 (API name claude-opus-4-8, released May 28, 2026 [4]). All 45 scripts ran once at a smaller demonstration scale on NVIDIA H200 GPUs (chips built for the heavy parallel math these simulations need). Automatic checks of the code ran alongside. Then one script per workflow-and-level combination, 15 in total, ran again at full production settings on the same hardware. Those runs were the study’s ground truth [1].

NVIDIA scored the results on three separate axes instead of one pass-or-fail grade. The first is property coverage, which we’ll shorten to correctness below to match the figures. It checks the code itself. Did it compute the right quantity with the right formula? It doesn’t check whether the number that comes out is the right number. The second is API-pattern coverage: did the code call the toolkit the way the toolkit prefers, or use a workaround that happens to run? The third is reusability: could someone else run the output again as a parameterized tool, or was it a throwaway one-off script [1]?

Each of the five levels adds to the one before it [1]. Sketch names the task type only. Goal adds material, method, scale, and deliverable. Recipe adds the step-by-step protocol. Spec names the toolkit’s own internal classes and constructs. Contract spells out the full command-line interface (CLI): every flag, output format, and acceptance test the script must pass.

Start with almost nothing#

The Sketch level names the task type and stops. Think “estimate a diffusion coefficient” (how fast a group of atoms move around and trade places), with no material, method, or scale [1]. A Sketch-level run produced about 498 lines of code in 44 agent iterations (write, run, fix cycles) [1]. It used about 2.4 million tokens, most of them cheap, cached re-reads of context the model had already loaded [1]. It scored 1.00 on correctness. The script computed the right quantity with the right formula, and every other level in the study would go on to match that score [1].

So a bare task description was enough. The agent’s own loaded skill and examples filled in the missing API pattern. Sketch left two things to chance, though: which physical settings the script would default to, and whether it would use the toolkit’s own preferred internal machinery. Goal is the next rung.

Figure 2 - Diagram of a one-line prompt naming only a task type, with an arrow passing through the agent's own loaded skill into a working script

Figure 2 - Nothing But a Task Name: A Sketch-level prompt supplies only the task type. The agent’s own pre-loaded skill and examples supply the missing API pattern, and the result is a working script that scores 1.00 on correctness, the code-formula check rather than a physical-accuracy check, at the lowest token cost of any level tested.

Say what a colleague would ask you for#

Goal adds four things: the material, the method, the scale, and the deliverable. That small addition closes the clearest physics failure NVIDIA saw in its earlier tests. Asked for “a transport property of a Li material,” an agent wrote a script that simulated argon instead of lithium [1]. Name the material and that failure goes away.

Goal still leaves the protocol to chance. In NVIDIA’s lithium runs, both the Sketch and Goal scripts defaulted to Langevin dynamics, a thermostat setting that acts like friction on the moving atoms. The diffusion numbers looked clean. They were also roughly 3 to 5 times lower than the runs without that friction [1].

Figure 3 - Diagram of two identical prompts, one naming only a transport property producing a script that simulates the wrong element, the other naming material, method, and scale producing the correct one

Figure 3 - What Naming the Subject Actually Buys: An ambiguous prompt for “a transport property of a Li material” produced a script that modeled argon. Naming the material, method, and scale removed that failure. The protocol was still left to the agent’s defaults.

Goal left other gaps too. The script was still nowhere near reusable. The code still didn’t reliably call the toolkit the way it was meant to be called.

Write down the protocol#

Recipe writes out the step-by-step simulation protocol, not just the goal. In the lithium task, the protocol was exactly the piece Goal had left to chance. Asking for NVE (the friction-free simulation mode the task actually needed) switched every script over to it. The Recipe and Contract lithium runs came out roughly 3 to 5 times higher than the Langevin-damped Sketch and Goal runs [1]. NVIDIA recommends Recipe by name for tasks like this one: “recommended when the result depends on simulation protocol” [1]. Recipe also pushed API-pattern coverage higher. That was the first sign that spelling out the process, and not only the goal, changes how much of the toolkit’s own machinery the code uses.

Naming the parts without the blueprint#

Spec goes one step further. It names the toolkit’s own internal classes and constructs, but without a complete contract for how they fit together. It looks like a safe middle ground between “too vague” and “too much work.” In this study it was the most fragile of the five levels. It caused 3 of the study’s 7 total screening failures [1]. NVIDIA’s production comparison table for the flagship lithium-diffusion task gives the Spec result no score at all. The table leaves it out. Its script’s total energy drifted during the run, which a physically valid simulation of this kind should never allow [1].

Figure 4 - Bar chart of the lithium production results at 800 K, showing Sketch and Goal noticeably shorter than Recipe and Contract, with a visibly empty gap where Spec would sit

Figure 4 - The Level That Didn’t Make the Table: At 800 K, the Langevin-damped Sketch and Goal runs report a diffusion coefficient roughly 4 to 5 times lower than the NVE runs at Recipe and Contract. Spec has no bar at all: its generated script failed a basic physical check and was excluded from NVIDIA’s table entirely.

By the Spec level, API-pattern coverage (how much of the toolkit’s own preferred machinery the code actually used) had jumped from roughly 0.52 at the lower levels to 0.96 [1]. That’s a real gain on one axis. It showed up at the exact level with the worst reliability record in the study. Naming internal pieces without a complete contract for how they connect gives the agent more room to guess wrong than naming fewer internals, or naming a full contract.

Write the full contract#

Contract is the top of the ladder: the full CLI, every flag, the exact output schema, and the acceptance tests the script has to pass. You write it as if the script will run unattended, with nobody checking its work. It was the only level that produced a fully reusable, parameterized tool instead of a throwaway script. Reusability rose from 0.67 at the lower levels to 1.00 [1]. That came at a real price. A Contract run used roughly 4 times the tokens of a Sketch run (about 10.0 million versus 2.4 million, including cache reads) and roughly 2.3 times the code (1,168 lines versus 498) [1].

Figure 5 - Bar chart across all five prompt levels with printed token and line-of-code totals on each bar, showing token cost roughly doubling between Recipe and Spec and code length dipping slightly at Goal before climbing to Contract

Figure 5 - The Price of Climbing the Ladder: Total tokens climb from about 2.4 million at Sketch to about 10.0 million at Contract, with the biggest jump between Recipe and Spec. Generated code dips slightly at Goal before climbing to about 2.3 times the Sketch length at Contract. Correctness never moves off 1.00 at any level.

Correctness didn’t move. Contract scored 1.00, the same as every level below it, all the way back to Sketch [1].

Figure 6 - Line chart across all five prompt levels showing reusability climbing from 0.67 to 1.00 only at the Contract level, alongside a correctness line that stays flat at 1.00 the whole way

Figure 6 - What the Last Rung Actually Buys: Reusability, whether the output is a parameterized tool or a one-off script, rises from 0.67 to 1.00 only at the Contract level. Correctness never moves off 1.00 at any level. The Contract level buys a reusable tool, and correctness stays where it already was.

Why more detail never changed the formula#

Put the five levels side by side and one fact holds across all of them. Correctness, NVIDIA’s property-coverage check, measured 1.00 at every level. In NVIDIA’s own words, “the science is right from the first prompt” [1]. The lithium table is the exception that matters. The formula was right at every level, but Sketch and Goal left the protocol to a default that damped the reported answer 3 to 5 times.

Why didn’t the formula change? The toolkit’s skills and examples supplied the API patterns whatever the prompt said. What more detail bought was structure: how much of the toolkit’s intended machinery the code used, whether the physical settings matched the task, and whether someone else could pick up the result and run it again.

NVIDIA’s own summary is short enough to remember: always name the system, the method, and the scale, and add a CLI contract only for unattended runs [1]. The lithium table adds a clause the summary leaves out. Recipe, the step that states the protocol, is “recommended when the result depends on simulation protocol” [1]. Skipping it is what left the Sketch and Goal lithium numbers 3 to 5 times low.

Naming the subject removes the wrong-subject failure. Stating the protocol removes the wrong-physical-default failure. Everything past that, especially naming internal pieces of the API one at a time without a complete contract for how they fit, is the most expensive and least reliable place on the ladder to stop. This is one measured case of a point we’ve made before about treating a coding agent’s pipeline like a compiler: the input spec should carry exactly what changes the output and nothing more. See Stop Treating Your Coding Agent Like a Chatbot [5].

KEY INSIGHT: Once the agent can reach a domain’s patterns through its own skills and examples, more prompt detail changes the cost and the structure of the code. Correctness stays the same. Always name the system, the method, and the scale. State the protocol too when the result depends on it. Save a full CLI contract for the runs nobody is watching.

Figure 7 - Decision diagram showing "name the system, method, and scale" as the default step on every prompt, a second branch asking whether the result depends on protocol, and a third branch reserving a full CLI contract for unattended runs

Figure 7 - The Rule in One Diagram: Name the system, the method, and the scale on every prompt against an unfamiliar API. State the protocol too when the result depends on it. Reach for a full CLI contract only when the script has to run unattended.

A different team found the same shape, from the opposite direction#

NVIDIA varied how much detail went into one prompt and held everything else constant. A separate team at DiDi International reached a similar lesson from the other end, in a production system [6]. They were building an automated quality-review system for a customer-service contact center. Their first version fed the model everything in a single call: the full CR Tree (a tree listing every possible reason a customer might have contacted support) plus the entire conversation. Accuracy came back far below expectations. The team assumed a better-worded prompt would fix it, and kept trying: “multiple rounds of prompt tuning could not change this behavior” [6]. What worked was splitting the one call into two, with the second gated on the first. A narrow check asks only whether the current label still looks reasonable. Only when that check fails does a fuller classification pass run, this time with the full CR Tree attached. Intent-verification accuracy (how often the system correctly judged whether a contact’s existing reason label was right) moved from 38% to 86% in DiDi’s own production validation. That’s a self-reported, unaudited number [6]. The mechanism is the useful part. A team that thought wording was the lever ran out of wording to try, and the fix turned out to be structural. It’s a case study in what a model is shown and when, a broader discipline we cover in The Context Engineering Stack [7].

Figure 8 - Diagram comparing a single call that feeds a model the complete decision tree plus conversation against a gated two-call sequence with a narrow verification pass first

Figure 8 - Same Lesson, Opposite Direction: DiDi’s single-call approach fed the model everything at once and accuracy fell far below expectations. Splitting the call into a narrow check followed by a conditional full pass moved intent-verification accuracy from 38% to 86% in DiDi’s own reported production numbers.

KEY INSIGHT: If rewording the prompt doesn’t fix a bad result, change what the model is shown and when. That usually matters more than how carefully the instructions are phrased.

Where the ladder stops#

Read the multipliers above as the result of one study. It’s a vendor benchmarking its own toolkit on 45 generated pipelines. That’s a small sample, and NVIDIA’s numbers describe this one study, not coding agents in general.

A different kind of limit sits under that sample-size caveat. In an earlier probe, separate from the formal 45-pipeline ladder, NVIDIA asked agents to measure lithium-ion diffusion in a perfect, defect-free lithium fluoride (LiF) crystal. That property can’t be measured at the simulated timescale at all. Every agent complied [1]. None declined or flagged the premise. Each one returned a number for a quantity that can’t be measured that way [1]. NVIDIA notes the sandbox where the code was generated had no web access, so an agent able to look the premise up might have caught it [1].

Figure 9 - Diagram of an earlier, separate probe where an impossible-to-measure request was sent to an agent, which complied and returned a confident number instead of flagging the premise

Figure 9 - A Probe Outside the Ladder: In an earlier probe, separate from the five-level study, an agent asked to measure a quantity that cannot be measured at the simulated timescale complied anyway and returned a confident number instead of questioning the premise.

One generated script shows the same limit more sharply. It “validated” its own unit conversion by generating synthetic test data using the same wrong constant it was checking, so its own self-test passed while the conversion itself was wrong. NVIDIA’s summary: “statistically reliable and physically meaningful can be orthogonal” [1]. In plain words, one tells you nothing about the other. None of the five ladder levels asked for that kind of check, and the agents never added one on their own. NVIDIA’s advice is to ask for it directly: request premise checks and uncertainty estimates, and require the script to reproduce an independent known result [1]. Even then, what caught the damped lithium numbers was the comparison against an outside reference, not anything inside the generated script.

Figure 10 - Diagram of a generated script validating its own unit conversion against synthetic test data built from the same wrong constant it was checking, so the self-test passes while the underlying conversion is still wrong

Figure 10 - A Test That Passed While Being Wrong: One generated script validated its own unit conversion against synthetic data built from the same incorrect constant it was checking. The self-test passed. The underlying conversion was still wrong.

The cost nobody bills for#

A Contract-level script buys reusability with roughly 2.3 times the code of a Sketch-level one, for the same correctness score. NVIDIA’s table prices that trade in tokens. It doesn’t price what happens after the script exists. Someone still has to review it. Simon Willison, talking with Claire Giordano on the Talking Postgres podcast, put a number on that side of the ledger: “I can churn out code a hundred times faster. I don’t have the cognitive capacity to stay on top of 100 times the amount of code” [8][9]. His baseline for what a good day used to look like: “200 lines of working, debugged, production-level code is an incredibly good day. Most days you’d produce 50 or 60” [8][9]. Both figures are his own rough estimates from a podcast conversation, not a measured study. His point makes NVIDIA’s cost finding sharper. The Contract script had 670 more lines of generated code than the same task at Sketch level (1,168 lines versus 498), for the same correctness result. All of that code still lands in front of a human whose review capacity hasn’t gone up 2.3 times, or 100 times, or at all. You can buy more tokens. You can’t buy more reviewer hours the same way.

GPT-6 Astra’s unmeasured claim#

OpenAI’s launch post for GPT-6 Astra (API name gpt-6-astra) makes a claim about a different lever. It says model capability, rather than prompt detail, can close the ambiguity gap on its own: “When instructions leave room for interpretation, GPT-6 Astra is better than previous models at making the right call. It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome” [10]. The post attaches no benchmark to that claim [10]. Treat it as a claim to watch. It neither backs up nor contradicts what NVIDIA measured.

Conclusion#

The code computed the right formula as soon as the prompt named a task at all. Every level NVIDIA tested confirmed it. More detail changed other things: cost, how much of the toolkit’s own machinery the code used, whether the physical settings matched the task, and whether the result could run again without a human rewriting it. The least reliable place to stop was one rung short of the top, where the prompt names internal pieces without a complete contract for how they fit.

So, against an API you don’t know well, name the system, the method, and the scale. When the answer depends on how the simulation or computation is run, spell out the protocol too. Reach for a full CLI contract only when the script has to run with nobody watching it. Budget for what that costs in tokens and in a reviewer’s time before you ask for it. None of these steps adds a premise check unless you ask for one. The final comparison against an independent reference still has to happen outside the script.


References#

[1] E. Tsai, F. Falcioni, J. S. Smith, P. Altoè, and W. J. Ong, “How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit,” NVIDIA Technical Blog, Aug 18, 2026. https://developer.nvidia.com/blog/how-ai-coding-agents-can-unlock-materials-simulation-with-nvidia-alchemi-toolkit/

[2] NVIDIA, “nvalchemi-toolkit,” GitHub, Apache License 2.0, tag v0.2.0. https://github.com/NVIDIA/nvalchemi-toolkit

[3] G. Dotzlaw, “Don’t Ship Skills Without Evals: A Reproducible Way to Prove Your Claude Code Skills Work,” Dotzlaw Consulting, Sep 9, 2026. /insights/claude-code-18-skill-eval-methodology/

[4] Anthropic, “Introducing Claude Opus 4.8,” Anthropic News, May 28, 2026. https://www.anthropic.com/news/claude-opus-4-8

[5] G. Dotzlaw, “Stop Treating Your Coding Agent Like a Chatbot. Treat It Like a Compiler.” Dotzlaw Consulting, Sep 17, 2026. /insights/ai-43-coding-agents-as-compilers/

[6] R. Hua, “How DiDi built intelligent contact center QA with Amazon Bedrock,” AWS Machine Learning Blog, Sep 8, 2026. https://aws.amazon.com/blogs/machine-learning/how-didi-built-intelligent-contact-center-qa-with-amazon-bedrock/

[7] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “The Context Engineering Stack: Compression, Retrieval, and Decision Memory,” Dotzlaw Consulting, Aug 12, 2026. /insights/ai-16-context-engineering-stack/

[8] S. Willison and C. Giordano, “How AI is changing software development,” Talking Postgres, 2026. https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison

[9] S. Willison, “Conceptual integrity and counting lines of code,” simonwillison.net, Aug 19, 2026. https://simonwillison.net/2026/Aug/19/conceptual-integrity-and-counting-lines-of-code/

[10] OpenAI, “GPT-6 Astra: A new generation of intelligence,” OpenAI News, Sep 3, 2026. https://openai.com/index/gpt-6-astra

The Prompt Specificity Ladder: How Much Detail Your Coding Agent Actually Needs
https://dotzlaw.com/insights/ai-48-prompt-specificity-ladder/
Author
Gary Dotzlaw
Published at
2026-09-28
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

Related reading

The Advisor Tool Is Real, and Anthropic Ranks It Last: The Cost Ladder and the One Number That Decides
Anthropic's advisor tool lets a cheap executor model consult a stronger advisor mid-task inside one API call, with no orchestration code. It is real, it is in beta, and Anthropic's own measured cost-lever guidance puts it dead last. A precise walk of the mechanism, the nine rungs you climb first, the consult rate that decides whether the pairing helps or hurts, and the four places a routing decision can live.
2026-09-03·AI & Modern Development
MCP Tool Design: The Two Ways Your Agent's Tools Fail (Bloat vs. Confusion)
Almost every MCP tool failure traces to one of two root causes, bloat or confusion, and the usual fix for one makes the other worse. A walk through AWS's six tool designs, Smartsheet's production token math, an independent eval where the arm with no tool catalog scored highest on correctness, and a checklist you can run against your own MCP server.
2026-09-02·AI & Modern Development
Prompt Architecture: Layer Your Prompts, Don't Bloat Them
One system prompt cannot be inviolable, situational, expressive, and self-checking at the same time. Split it into four stacked layers, make the last one code instead of text, and you get the only guarantee a prompt was never able to give you.
2026-08-24·AI & Modern Development
Stop Making the Model Watch the Whole Video
Most companies with a video archive priced AI analysis once, flinched, and shelved it. The number changed, and the reason is a change in who decides what gets looked at. Here is the arithmetic of the old approach, the mechanism that replaced it, what two outside benchmarks found when they went looking for the failure mode, and the general pattern underneath.
2026-09-18·AI & Modern Development
← Back to Insights