4369 words
22 minutes
The Recipe Card: How to Hand an Agent a Whole Job

Picture a back-office team that runs vendor onboarding the same way every time. Collect the tax form, confirm the business is real, set up banking details, route the file for a compliance check, activate the vendor in the accounting system. Every step is well understood. The job still takes days, because each step waits on a person to notice it, do it, and hand it to the next person. The team has a chat-based AI assistant and uses it the way almost everyone does: paste in one document, ask one question, get one answer, check it, paste in the next thing. Nobody hands the assistant the whole vendor file and asks it to run the process end to end. Doing that would mean deciding, in advance, what the assistant should ask before it starts. It would also mean deciding what it can do on its own, and which steps need a person’s sign-off first. That feels like writing a technical specification before the job can even begin. So nobody writes one, and the job stays a manual relay race.

Figure 1 - Diagram showing a one-line prompt and a full technical specification as two extremes, with a recipe card positioned as the middle layer between them

Figure 1 - A Recipe Card Sits Between a Prompt and a Spec: A one-line prompt is fast to write but gives an agent nothing to act on safely. A full specification gives an agent everything it needs, but almost nobody writes one before starting. A recipe card is the middle layer: structured enough to hand off, short enough that someone will actually write it.


Why a week of admin work never gets delegated#

We think the habit runs deeper than any one team’s workflow. Chat-based assistants trained a generation of users to work in a loop: ask, read, check, ask again. That loop is fine for a single question. It breaks down once a job spans a dozen steps that depend on each other. Checking every answer along the way can eat most of the time the delegation was supposed to save. Most people respond in one of two ways. They skip delegating the big job, or they hand it over piecemeal, one task-sized chunk at a time. The second option quietly hands the coordination work back to the human. How capable the model is doesn’t change that.

Practitioners like Nate B Jones now report agents staying on a multi-day, multi-app job without losing the thread [1]. The way we hand over work hasn’t caught up. A single prompt was never built to carry a job this size, and nobody has a template for what to write instead. Vendor onboarding shows the gap. The team knows the five steps by heart. It still has no format for telling an agent all five at once, in a way the agent could run safely without a person babysitting every step.

Fix 1: a one-line prompt#

We think the first thing most people try is the obvious move: type “onboard this vendor” into the assistant, attach the file, and see what happens. It takes ten seconds to write and asks nothing of the person typing it. That ease is the problem. The prompt says where the job should end up. It says nothing about how to get there, or which questions the assistant should ask before it starts. Does this vendor need a compliance review, or does it qualify for the fast lane? The prompt also leaves out what the assistant may do without checking in, and when it should stop and report back. It doesn’t name the actions that need a human to say yes first, like approving a bank account change. An assistant working from a one-line prompt has two choices when it hits an unclear step: guess, or stop and ask about everything. A team trying to hand off a job wants neither.

Figure 2 - Diagram showing a one-line prompt failing to specify starting questions, unsupervised scope, check-back points, or approval-gated actions

Figure 2 - A One-Line Prompt Answers None of the Handover’s Questions: “Onboard this vendor” names the destination. It says nothing about what to ask first, what the agent can do alone, when it should check back, or which actions need a person’s sign-off. An agent working from this has to guess or stop constantly, and neither is a real handover.

Fix 2: the full spec nobody writes#

The obvious fix for a vague prompt is a precise one. Before any work starts, you write out every sub-step and every branch. You add an approval matrix, a table naming exactly which actions need sign-off and from whom. If someone actually does this, it works. The agent has everything it needs, and there’s nothing left to guess at.

The problem is that almost nobody does it. Jones, who hosts the AI News & Strategy Daily YouTube channel, hit the same wall when he handed an agent a household move instead of a back-office job. He names the resistance directly: “Nobody wants to sit down before a household move and write an approval matrix and a 14-part spec for an AI” [1]. Offices feel the same resistance. A document that detailed, written before the first useful thing happens, is a barrier most teams won’t cross. So the spec never gets written. The job goes back to being done by hand or in disconnected pieces, which is exactly where it started.

Figure 3 - Diagram of a 14-part specification document with an approval matrix, shown as complete but never written, next to a job still being handled by hand

Figure 3 - A Complete Spec Works, But Almost Nobody Writes One: A full specification with an approval matrix gives an agent everything it needs. The barrier is the effort of writing it before any work starts. Most teams never cross that barrier, so the job stays manual.

The recipe card sits in the middle#

Jones’ answer to the gap between those two extremes is an artifact he calls a recipe card, and we think it’s worth taking seriously. It’s shorter than a full spec and more structured than a prompt. Its job is to give an agent enough to act on without asking anyone to write a specification. In Jones’ words, a recipe card “names a real job. It shows all of the jobs inside it, at least as a sketch. It tells the agent what to ask you for, what the agent can handle, when the agent ought to come back, and which actions need your approval” [1].

That’s six fields in one breath: the job, its sub-jobs, the starting questions, the unsupervised scope, the check-back points, and the approval-gated actions. Unsupervised scope is what the agent may do without asking. A check-back point is when it has to report in. An approval-gated action is a step that needs a person’s sign-off before it happens.

This is one practitioner’s name for the format, tested in one video against one kind of job. Read it as exactly that: one practitioner’s format for a loop that several vendors and teams already ship on their own. Nobody has made it an industry standard. We come back to the outside evidence, field by field, after filling in a card for a real job.

Figure 4 - Diagram of the recipe card's six fields: the job, sub-jobs, starting questions, unsupervised scope, check-back points, and approval-gated actions

Figure 4 - The Recipe Card’s Six Fields: The job named plainly, and the sub-jobs inside it sketched out. The questions the agent should ask first. What the agent may do without checking back, and when it has to check back. Finally, which specific actions need a person’s approval before they happen.

Who actually runs the card: a manager agent#

A recipe card only helps once something reads it and acts on it. That something is a manager agent, sometimes called a chief-of-staff agent. Its job is to coordinate the work, not to do it. The loop runs in three moves.

First, the manager takes the card’s starting questions and interviews the human before doing anything else, instead of guessing at missing details. For vendor onboarding, it might ask which compliance tier the vendor falls into, or whether banking details are already on file.

Second, once it has enough answers, the manager turns the card’s sub-jobs into tasks for execution agents (agents that do the actual work). Each one works a specific piece, such as verifying the business registration, drafting the compliance summary, or preparing the banking setup. Pieces that aren’t blocked keep moving in parallel instead of waiting in a single line.

Third, the human’s role shrinks to one point of contact: answer the manager’s questions, approve what the card marked as needing sign-off, and otherwise let the manager keep the execution agents working [1].

That narrowing is on purpose. Jones is direct about what it’s meant to protect: “I want to bring the person back to the choice, the risk, the responsibility that is ours as humans” [1]. An agent can do the research, draft the summary, and prepare the paperwork. It shouldn’t be the one deciding whether a vendor’s compliance risk clears the bar. That’s a judgment call, and getting it wrong has real consequences. The card’s approval-gated-actions field is where that line gets written down, once, in advance. Nobody has to improvise it mid-run, depending on who happens to be paying attention when the moment comes.

Figure 5 - Diagram of the manager loop: a manager agent interviews the human, dispatches sub-jobs to execution agents in parallel, and returns only approvals and open choices

Figure 5 - The Manager Loop: A manager agent interviews the human using the card’s starting questions. It dispatches the card’s sub-jobs to execution agents that run in parallel where possible. It brings back only the choices and approvals the card marked as needing a person. The human answers one point of contact instead of running every step.

Filling in a card for vendor onboarding#

The fields only mean something once they’re filled in for a real job. Here’s the card we would write for vendor onboarding. It’s an illustration, using the fields exactly as Jones states them [1].

The job: Onboard a new vendor into the accounts payable system before the first invoice is due. Done means the vendor is active, its banking details are verified, and a test payment file for it validates in the accounting system.

The sub-jobs inside it: Collect and validate the tax form. Confirm the business is a real, registered entity. Set up and verify banking details. Run a compliance check against the vendor’s industry and country. Activate the vendor record in the accounting system.

Questions the manager should ask first: Which compliance tier does this vendor fall into? Is there an existing relationship with this vendor under a different name? Are banking details already on file from a prior engagement? Is there a deadline tied to a specific invoice?

What the agent may do without checking back: Validate the tax form’s format and completeness, search public business registries to confirm the entity exists, draft the compliance summary, and prepare the banking setup for review. It’s the same unsupervised-scope idea covered in Give the Agent a Budget, Not a Token [2], applied here to one specific job instead of a general permission grant.

Check-back points: After the compliance summary is drafted and before it’s submitted for a human compliance review. After banking details are prepared and before they’re saved to the vendor record.

Approval-gated actions: Saving new banking details to the vendor record. Marking a vendor as cleared on a compliance flag. Activating the vendor in the accounting system.

Figure 6 - The vendor onboarding recipe card as six rows: the job named at the top, then sub-jobs, starting questions, unsupervised scope, check-back points, and the approval-gated actions highlighted in amber

Figure 6 - A Recipe Card Filled In for Vendor Onboarding: The six fields of the vendor onboarding card, with the job named at the top and the approval gate in amber. The full text above fills them in: five sub-jobs, four starting questions, four unsupervised actions, two check-back points, and three approval-gated actions, all written down once before the manager loop runs. This is our own worked example, built to show the fields in use, not a real client engagement.

Who else ships pieces of this#

Cursor, the coding tool, ships a version of the same structure in a product called Projects, in beta as of this writing [3]. “The coordinator doesn’t write code itself but directs other agents that do. Because it delegates rather than executes, it is never blocked and is always responsive to direction” [3]. That’s the card’s “one point of contact” idea, built into software. The coordinator has no execution role at all. That’s why the loop can keep moving on the unblocked parts of a job while a person is away. Cursor’s version researches the codebase before planning, instead of interviewing a human first. So it backs up the part where one coordinator hands work out to several agents, but not the card’s questions-first field.

OpenAI documents a two-step handover for Codex, its coding agent (the same docs also cover the ChatGPT desktop app, hence “ChatGPT” in the quote), covering the piece Cursor leaves out. You type /plan or /goal as commands in the chat box: “The goal text becomes both the first prompt and the completion criteria for the task. If the outcome is still unclear, start with /plan. Ask ChatGPT to interview you, identify constraints, and turn the result into a goal with measurable success criteria. Then start the refined goal with /goal” [4]. Plan mode is the card’s starting-questions field, run by the product itself rather than described in a video. The goal it produces states in advance what counts as finished, which is the card’s check-back field in a different shape: instead of a person deciding when the agent should return, the goal itself defines, up front, what “done” looks like.

Lyft, the rideshare company, described its approach in a post on the LangChain Blog. Every self-serve support agent, the kind its operations staff build without engineers, must fill in a five-part template before it ships: “identity (who is this agent, what user type, what topic area), primary objective (concrete verbs, not vague ‘help’ or ‘handle’), scope (both in-scope AND out-of-scope with explicit routing actions), phased workflow (numbered steps with entry conditions, branching for every if/else, and a terminal action for every phase), and content guidelines (concrete do/don’t rules with example phrases, not abstract principles)” [5]. Lyft pairs the template with “a review checklist that every prompt must pass before activation,” including “does every phase have an exit?” [5]. That template backs up two of the card’s fields. It names what an agent should refuse as well as what it should do, which is the scope field. It also requires an exit condition for every phase instead of leaving “done” implicit, which is the check-back field’s definition of done.

LangChain’s own GTM agent (GTM is go-to-market, the team that runs outbound sales outreach) adds a fourth piece: a card that persists instead of getting rewritten for every job. Before drafting an email, the agent “follows a defined outbound skill, a playbook it loads before drafting” [6]. That skill is enforced by “a consistent checklist (do-not-send checks, research, draft, rationale, follow-ups)” [6]. Its first move on any new lead is deliberately cautious: “The first thing it does is look for reasons not to send anything” [6]. Its check-back field also has a concrete default for when a person doesn’t answer. In LangChain’s system, a silver lead is the post’s middle-value tier of prospects, and a rep is the salesperson who owns the lead. The default is written as an SLA (a service-level agreement, a timed commitment for how fast a response is due): “we added a 48-hour SLA for silver leads: if a rep hasn’t approved or declined the draft within that window, it sends automatically” [6]. That default only covers one tier. It’s a deliberate, later loosening of the team’s own stated rule that “nothing is sent without an explicit rep review and approval” [6]. It’s the same kind of graduated trust covered in Risk-Graduated Agent Autonomy [7].

None of these four sources names the card or validates its full six-field list as a unit. Taken together, we think they give most of the fields independent, first-party support. The job and sub-jobs fields are the least tested, since they’re the easiest to write.

KEY INSIGHT: When a delegation format comes from one person’s video, check whether its pieces show up in products other teams have shipped. Look for the interview, the scope boundary, the definition of done, and the approval default. Independent shipping is stronger evidence than any single demo.

Figure 7 - Diagram mapping four vendor products (Cursor, Codex, Lyft, LangChain GTM) to the specific recipe card fields each one corroborates

Figure 7 - Four Products, Four Fields Corroborated: Cursor’s Projects coordinator corroborates the one-point-of-contact manager structure. Codex’s plan-then-goal mode corroborates the starting-questions field and the check-back field’s definition of done. Lyft’s five-part template corroborates scope and exit conditions. LangChain’s GTM agent corroborates a persisted card and a timeout default on approval. None of the four validates the card’s name or its full field set as a single unit.

Two fields need close watching once the card is actually running. The first is what “done” means at each check-back point. The card says when the agent comes back, and that only works if it also says what finished looks like. The second is how the approval field should change after the card has run cleanly a few times. Both lessons come from teams or testers who ran an agent and watched the results, not from a single description of the format.

What “done” means, and why the agent shouldn’t be the one checking#

A card can say “check back when the compliance summary is done” without saying what done actually means. That gap is where trouble hides. We’d argue this is the easiest field to get wrong, because skipping it costs nothing.

Hyperautomation Labs is a YouTube channel that handed ten small-business jobs to an agent. Seven ran on a plumbing company the creator made up for the test, and three ran on the creator’s own business data. The channel pairs a plain paragraph describing the finished result with an answer key the agent never sees: “one paragraph that describes what done looks like, not steps, just the finished result. Behind that folder sits an answer key” [8]. A script grades the run against that key after it ends. The creator’s advice for anyone repeating the test is to check by hand at first: “for the first month, check the output by hand every time. After that, you will know exactly which jobs you can stop checking” [8]. One job ran on the creator’s real sales export. There, the agent “noticed that the export ends with a totals row and left it out” [8]. The script’s answer key, not the agent’s own say-so, is what confirmed the resulting totals were right.

That practice matters more once you see what happens when the check doesn’t run against the real outcome. NVIDIA’s SWE-Serve benchmark (SWE is short for software engineering) gave coding agents real change requests against SGLang, an open-source server that runs AI models. Each patch was graded two ways. One used the full test suite. The other used the same suite minus the live-serving tests, which load a model and send it real requests through a running server. Across 627 patches, the pass rate was 45.9% under the complete verifier and 69.4% once the live-serving tests were removed [9]. Roughly one in three patches that looked finished under the lighter check actually failed once the real system ran them. The authors also note that the benchmark’s tests “don’t establish that an agent patch… is deployable, ready to merge, or endorsed by SGLang maintainers” [9].

The benchmark is a coding task, but the lesson carries over to back-office work. Name the finished result where it actually gets used. The vendor card’s job field does this by naming a payment file that validates in the accounting system. Don’t settle for a proxy check the agent can pass on the way there.

Figure 8 - Bar chart of the same 627 coding-agent patches, 45.9 percent passing with live-serving tests included and 69.4 percent passing with them excluded

Figure 8 - The Gap Between “Looks Done” and “Is Done”: On the same 627 coding-agent patches, the pass rate was 45.9% when live-serving tests were included in the check. Once those tests were excluded, it was 69.4%. About one in three patches that passed the lighter check still failed once the real system ran them.

KEY INSIGHT: Write the finished result as something checked where it will actually be used, not as a step the agent can complete on the way there. If the only thing confirming “done” is the agent that did the work, any mistake it can’t see also goes unchecked.

Approval starts tight and loosens with a track record#

The approval-gated-actions field is written once, at the start, but it doesn’t have to stay fixed. Cursor ran Projects on a design system (a shared library of reusable UI components), and its own account shows the gate relaxing as the coordinator earned trust: “At first, the engineer reviewed each fix and corrected the ones it got wrong. Now the coordinator scans every new PR [pull request, a proposed code change], extracts components that belong in the design system, and adds a lint rule [an automatic check that flags the mistake next time] whenever it sees the same mistake twice” [3]. Hyperautomation Labs gives its viewers the same advice from the other direction: check by hand for the first month, then stop checking the jobs that keep coming back clean [8].

Both point at the same practice. We recommend starting a new card with tighter approval gates than you expect to need for good. Then loosen a specific gate only once the agent has a track record on that specific job, not because the demo went well.

For vendor onboarding, the three approval-gated actions listed above stay in place for a while. They stay until the team has watched the manager loop run cleanly on enough vendors to know which ones really need a human eye. A compliance flag on a high-risk industry probably stays gated indefinitely. A routine domestic vendor with no flags might eventually move to a lighter check, once the pattern is established instead of assumed.

Figure 9 - Diagram showing an approval gate starting tight, with every action reviewed by hand, then loosening over time to only the actions that still show a track record of needing review

Figure 9 - Approval Starts Tight and Loosens With Evidence: Every gated action is checked by hand at the start. Only the specific actions that build a clean track record move to a lighter check later. The gate that never earns a loosening, like a high-risk compliance flag, is allowed to stay tight indefinitely.

Where this format still falls short#

The recipe card’s name and its full six-field layout still rest on one practitioner’s video. The outside evidence above is real and specific: a shipped manager loop, a shipped interview-then-goal handover, a shipped scope-and-exit template, and a shipped persisted card with a timeout default. None of the four sources uses the name “recipe card” or validates the complete six-field list as a single format. So call it what it is: a practitioner’s name for a loop several vendors and teams now ship. Nobody has agreed to it as a standard.

The approval field needs an ongoing decision about when the evidence is strong enough to loosen it. You don’t set it once and walk away. Get it wrong in either direction and the cost is real: an agent that never earns any autonomy, or one that earns it too fast. The job field’s definition of done needs a real outside check. Even a well-built one, like NVIDIA’s benchmark verifier, comes with its own caveat about what a pass actually proves [9]. Treat the format as a starting structure to adapt to your own job. Watch it closely the first several times it runs. Don’t treat it as a checklist to fill in once and trust forever.

Start with one card#

What kept vendor onboarding, and jobs like it, from being delegated was a missing format. A prompt was too thin to act on, and a spec was too heavy to write. The recipe card fills that gap with six fields any team can write for a job it already owns. A manager agent turns the card into a running loop. It interviews the human first, sends sub-jobs to execution agents, and brings back only the choices that matter. Once the loop is running, the check-back points and the approval gate deserve the most care. Check the real outcome instead of a proxy, and start conservative before loosening any gate.

The next step is small on purpose. Pick one recurring job already sitting on a team’s desk, such as vendor onboarding, month-end close, or new-hire paperwork, and fill in the six fields this week. A rough first card is fine. Your next run then starts from something more than a one-line prompt.


References#

[1] N. B. Jones, “There Are Jobs You Could Never Give AI. I Gave GPT-6 Astra 20 Hours Of Admin,” AI News & Strategy Daily, YouTube, Sep. 7, 2026. https://www.youtube.com/watch?v=ix8SsXjBc7M

[2] G. Dotzlaw, “Give the Agent a Budget, Not a Token: Four Dimensions That Replace a Yes-or-No Grant,” Dotzlaw Consulting, Sep. 30, 2026. /insights/ai-50-give-the-agent-a-budget-not-a-token/

[3] A. Robbins and F. Lindh, “Introducing Projects,” Cursor Blog, Sep. 10, 2026. https://cursor.com/blog/projects

[4] OpenAI, “Long-running work,” ChatGPT Learn Docs, accessed Sep. 29, 2026. https://learn.chatgpt.com/docs/long-running-work

[5] A. Sharma, “How Lyft Built a Self-Serve AI Agent Platform with LangGraph and LangSmith,” LangChain Blog, May 27, 2026. https://www.langchain.com/blog/lyft-built-a-self-serve-ai-agent-platform-for-customer-support-with-langgraph-and-langsmith

[6] V. Suresh and J. Ou, “How we built LangChain’s GTM Agent,” LangChain Blog, Mar. 9, 2026. https://www.langchain.com/blog/how-we-built-langchains-gtm-agent

[7] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “Risk-Graduated Agent Autonomy: Score Every Action by Blast Radius,” Dotzlaw Consulting, Sep. 4, 2026. /insights/ai-38-risk-graduated-agent-autonomy/

[8] Hyperautomation Labs, “What to Automate FIRST: 10 Boring Business Jobs, Tested in Claude Code (10/10 Passed),” YouTube, Sep. 12, 2026. https://www.youtube.com/watch?v=8U3zPU4cvaA

[9] J. Williams, D. Farris, J. Farris, and J. Jiao, “How SWE-Serve Exposes the Gap Between Local Tests and Live Serving,” NVIDIA Technical Blog, Sep. 23, 2026. https://developer.nvidia.com/blog/how-swe-serve-exposes-the-gap-between-local-tests-and-live-serving/

The Recipe Card: How to Hand an Agent a Whole Job
https://dotzlaw.com/insights/ai-58-recipe-card-delegating-multi-day-work/
Author
Gary Dotzlaw
Published at
2026-10-07
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

← Back to Insights