4493 words
22 minutes
Give the Agent a Budget, Not a Token: Four Dimensions That Replace a Yes-or-No Grant

An agent was cleaning up workloads (running jobs on a shared compute cluster) that it judged no longer useful. That’s a reasonable job to hand an agent. One stage in its filtering pipeline evaluated to nothing, so instead of matching the small set of stale jobs it was supposed to find, the selector matched everything. In 90 seconds the agent deleted about 200 workloads, affecting roughly 20 engineers’ worth of work, some of it long-running training runs that may not have been checkpointed [1]. Sachin Malhotra, an engineer on Anthropic’s CI (continuous integration, where code changes are tested and merged automatically) team, told this story at an AI Engineer conference talk [1]. That team builds the test machinery, merge automation, and autoscaling [2] that, in his words, a few thousand engineers rely on every day [1]. Nobody was being malicious. The agent, in his words, “genuinely thought it was tidying up after itself” [1].

The filter bug matters less than what the agent’s credential allowed while the bug ran. The agent held a token (a credential that proves to a system who is calling and what it may do), and that token was completely valid. Malhotra’s own diagnosis: “the agent technically hadn’t done anything that I couldn’t have done. It was using my token after all” [1]. A token only ever answers one question: can this caller do this kind of thing, yes or no? It has nothing to say about how much, how fast, or whether the damage can be undone.

Figure 1 - Diagram contrasting a token, a single yes-or-no gate, with a budget split into four dimensions: how much, how fast, what can undo, who is watching

Figure 1 - A Token Answers One Question, a Budget Answers Four: A token is a boolean gate: the caller either holds the scope or doesn’t. A budget replaces that single question with four dimensions: how much, how fast, what it can undo on its own, and who’s watching.


A credential that worked exactly as designed#

The first instinct is to fix the filter and move on. That fixes this bug and leaves the next one just as dangerous, because the real failure sat one layer up, in what the token allowed once something upstream went wrong.

A selector (the logic that picks which records a command applies to) matching everything is an ordinary bug. It shows up in code with no agent anywhere near it. What made it expensive was that the thing running it held a delete permission with no ceiling. We think a human might have caught the unusual count. An agent won’t unless something is watching for it, and here nothing was. Malhotra’s own read: “the failure wasn’t the model itself. The failure was that I was giving the agent unbounded amount of power to do something that I wasn’t watching super intently” [1].

Malhotra compares it to onboarding. Organizations already solved a version of this problem for new hires. Nobody watches a new hire’s every keystroke, yet a new hire can’t reach anything catastrophic on day one, because those actions are structurally out of reach. Agents, in his framing, get the onboarding problem without the onboarding structure: “they never get tired. They never sleep. And every so often they’re just like very confidently wrong” [1].

Figure 2 - Timeline diagram of a filter evaluating to nothing, a selector matching every workload, and 200 workloads deleted in 90 seconds under one valid token

Figure 2 - How 200 Workloads Disappeared: A filter stage evaluated to nothing, so the selector matched every workload instead of a handful of stale jobs. The agent’s token stayed valid the whole time, so nothing in the credential stopped it.

Why narrowing the token runs out of road#

The obvious response is to narrow the token. Take the delete permission away, or shrink the list of what the agent can touch until the blast radius (how much breaks when an action goes wrong) shrinks with it. It’s cheap and quick, and it does stop this exact incident from happening again, for a while.

Malhotra calls this “the standard fix,” and he’s also the one who explains why it doesn’t hold [1]. A token is a boolean: a static list of scopes (the named permissions a token carries) you either have or you don’t. Too tight, and the agent can’t do its job, so someone widens it again the first time it needs the permission for a legitimate reason. Too wide, and “you’re maybe writing a postmortem” [1]. Narrowing a scope after an incident buys maybe a week or two. Then the agent needs the verb it lost, and a human is back to approving every instance by hand. That human is now doing the agent’s work, with extra steps.

Go back to the new hire. You wouldn’t take a whole capability away from a junior engineer after one mistake. You’d limit how much of it they use unsupervised, set an escalation path for the rest, and let them keep working. A token can’t express “narrow how much,” only “have it or don’t.”

KEY INSIGHT: If your fix for an agent incident is “take the permission away,” check whether the agent has a legitimate, recurring reason to use that permission. If it does, the fix only schedules the next incident for whenever someone widens the scope back out.

Figure 3 - Diagram of a token as a single boolean gate, too tight blocking legitimate work and too wide enabling a large incident, with no setting between the two

Figure 3 - A Token Has Only Two Settings: A token is a yes-or-no gate. Narrow it and the agent can’t work. Widen it and one bug can do the cold open’s damage. There’s no setting in between.

Grading verbs by whether they fail loudly or quietly#

Malhotra names four dimensions, then teaches three primitives you enforce (asymmetric verbs, refilling rate limits, and trip wires) and one lens you use to size them (the undo test) [1]. The first primitive is about what an action does when it goes wrong. He calls it asymmetric verbs, where a “verb” is any operation the agent can perform: an API call, a CLI command, anything [1]. Two actions can look the same size on paper and carry very different consequences.

His worked example is skipping and unskipping a test in a CI pipeline. Unskip means the agent re-enables a disabled test. If that call is wrong, a bunch of builds turn red, a human notices right away, and the fix is to skip the test again. Skip is the reverse: the agent disables a test on its own judgment. If that call is wrong, nothing turns red, and a real bug can walk into production behind a green checkmark [1]. Same action, opposite failure visibility.

The rule follows from that pair: give the agent full autonomy on verbs that fail loudly, and keep a human on verbs that fail quietly [1]. In Malhotra’s system, the agent can re-enable a quarantined test (one temporarily skipped because it’s flaky or broken) on its own. Disabling one stays a “break glass” action that an on-call human reaches for under pressure, since a wrong call there is invisible [1]. Every call still gets logged, but the agent never writes the log entry. A separate layer stamps the real caller’s identity on it, which we come back to later.

Another team’s production agent, built independently, draws the same line. LangChain’s write-up of an autonomous SRE (site reliability engineering, the work of keeping production systems running) agent for Kubernetes (the most widely used system for running containerized services) grades actions by how easily a human can judge them at approval time: “Scaling a deployment to 3 is legible in a glance; a helm upgrade is one click that rewrites dozens of resources you can’t see from the approval prompt” [3]. A helm upgrade redeploys a whole packaged application at once. A verb that fails quietly and a verb a human can’t judge before approving are close cousins. Both need more than a rubber-stamp yes.

Grading verbs this way gets past the “too tight or too wide” trap of a single token. But it’s silent on volume. A verb’s grade says nothing about how many times the agent can fire it before something’s wrong.

Figure 4 - Grid comparing unskip, which fails loudly on a dashboard, and skip, which fails silently, showing which gets agent autonomy and which stays with a human

Figure 4 - Same Action, Opposite Failure Visibility: Unskip and skip look identical, but a wrong unskip turns builds red while a wrong skip lets a bug ship silently. The agent keeps the loud failure. A human keeps the quiet one.

A ceiling that refills#

The second primitive is a rate limit sized per caller, and Malhotra calls it “the most concrete form of the budget idea as a whole” [1]. Every caller gets a small allotment of disruptive actions per time window, to spend however it wants with no approval step. Cross the line and the request bounces, the caller waits, and the ceiling refills on its own. Nobody has to file a ticket to reset it [1].

A neighboring team built this after the cold-open incident [1]. An admission webhook (infrastructure that intercepts a request before it’s allowed to run) now caps deletes at a fixed number per hour, per resource kind, per namespace (a named partition of a cluster, usually owned by one team). A human bypass flag exists for real emergencies. Inside an agent session, though, the flag only tells the agent to ask a human to run the command [1]. The agent keeps the rate limit. The human keeps the override.

Figure 5 - Diagram of a rate limit ceiling that fills with actions, blocks a request once crossed, then refills automatically after a time window with no ticket required

Figure 5 - A Ceiling That Refills Itself: Every caller gets a small allotment of disruptive actions per time window. Cross the line and the request bounces back. The ceiling refills on its own, no ticket required.

Jamf, an enterprise device-management company, runs the same idea in production at a different layer, and there’s a public reference implementation you can inspect [4][5]. Jamf tracks each engineer’s daily spend on Amazon Bedrock (AWS’s managed API for large language models) and restricts access in tiers as spend approaches budget. Opus, the largest of the three Claude model families Jamf tiers, is denied at 80% of an engineer’s daily budget. Sonnet, a lighter model, is denied at 100%. Haiku, the lightest, is never denied, so a capped engineer always keeps a working fallback [4]. Every 15 minutes a scheduled process recomputes who’s restricted and rewrites the access policy, so the next call is checked against the fresh number with no re-authentication [4]. The reference implementation prices an unrecognized model at the highest tier rather than zero, so an unmapped model can’t slip past enforcement by accident [5]. That public AWS sample is a simplified version of what Jamf runs in production. Its documentation notes that spend-velocity anomaly detection and staged multi-account rollout were both cut [5].

Jamf’s write-up closes on a finding worth carrying forward: a hard per-engineer cap made leadership more comfortable expanding AI access. “Governance accelerates adoption rather than restricting it,” and “putting a hard per-user cap in place made leadership comfortable expanding access, not contracting it” [4]. The cap made spending observable, and Jamf says that visibility is what let leadership grow access [4].

Figure 6 - Chart showing Jamf's tiered Bedrock model access, Opus denied at 80 percent of daily budget, Sonnet denied at 100 percent, and Haiku never denied

Figure 6 - Jamf’s Tiered Spend Enforcement: As daily Bedrock spend climbs, Opus is denied first, at 80% of budget, then Sonnet at 100%. Haiku stays available, so a capped engineer keeps a working model.

The same shape shows up in a different stack. LiteLLM, an open-source gateway (a proxy that every model call passes through), supports a spend budget and rate limits on each API key it issues. In AWS’s walkthrough, four deployment settings feed those limits when a key is created. CODEX_KEY_MAX_BUDGET and CODEX_KEY_BUDGET_DURATION set the spend ceiling. CODEX_KEY_TPM_LIMIT and CODEX_KEY_RPM_LIMIT set the rate ceiling in tokens per minute and requests per minute [6]. Here a token means a chunk of text the model reads or writes, a different sense of the word from the credential in this article. The habit worth taking from that write-up is how to test the limits: “Validate those controls with a disposable identity by crossing a configured threshold and confirming that LiteLLM rejects the next request. This tests the policy itself rather than only confirming that the settings were accepted” [6].

KEY INSIGHT: Don’t trust a rate limit until you’ve deliberately crossed it once with a disposable credential and watched the system reject the next request. Seeing the setting saved only proves the configuration was accepted.

A rate limit stops a single loop from running too far. It only tells you that one request crossed a line, though. It says nothing about the pattern of an agent’s overall behavior.

Watching the aggregate instead of guessing up front#

The third primitive replaces a guess with data. An allow list (a fixed set of permitted actions, written before you have any evidence about how the agent behaves) is a guess made in advance. A trip wire gathers that evidence after the fact, without stopping the agent up front [1]. For cheap, reversible actions, let the agent act, stamp every action with an identity, and watch the total rather than any single call.

Malhotra’s example is a metric his team tracks: how many investigation threads an agent launches per hour for a given test-failure signature. One morning, as he tells it, the count sat well above baseline and the trip wire paged the on-call engineer. The page came after the threshold had already been crossed, which he describes as “the smoke detector, not the lock on the door” [1]. The agent had started a separate investigation for each of dozens of jobs that were all failing with the same underlying error. Each thread looked reasonable on its own. Together, they were one infrastructure failure disguised as dozens of unrelated bugs. The fix was about one line of guidance added to the agent’s context: correlate failures across jobs before launching a separate investigation for each one. The next time the pattern showed up, the agent did exactly that [1].

Allow lists stay fixed while the agent’s behavior changes, so they drift out of date. Trip wires get better with each trip, because each one points at a concrete fix [1].

None of the three primitives, alone or together, answers the question that decided how bad the cold-open incident got: once something goes wrong, can it be undone?

Figure 7 - Comparison diagram of a static allow list written before any data exists against a trip wire that watches aggregate agent behavior and pages when a threshold crosses

Figure 7 - A Guess Up Front vs. Evidence After the Fact: An allow list, written before you know how the agent behaves, goes stale. A trip wire watches the aggregate, pages once a threshold crosses, and improves with every trip.

The question that sizes the other three#

The undo test is a lens you apply while deciding how much of the other three an action needs. It’s two questions in one: can the agent put this back by itself, and how bad is it if it can’t [1]?

If the agent can roll back its own change and the worst case is acceptable, log it and move on. If either answer is no, the action needs a second key that the agent never holds, plus an audit record of who supplied that key and why [1]. Malhotra’s example outside CI is a feature flag (a switch that turns a feature on for some users). The agent has full authority over the flag on staging traffic (internal and early-access users), free to ramp it from 0 to 100 or roll it back. It has no scope to promote that flag to production, so the best it can do there is propose the promotion to a human [1].

Apply both to the cold-open incident. A rate limit would have capped the deletion at a few tens of workloads instead of 200. The undo test covers the rest: nobody can un-delete a running job in someone else’s namespace, so in Malhotra’s design anything past that cap needs a human holding a second key [1].

PayPal reached a closely related rule from its own production work and turned it into a reusable framework. Its three-tier matrix ties the evidence an action needs to the stakes and the counterparty (the party on the other side of the transaction). Inside a known ecosystem, system logs are enough. Between known partners, there’s shared trust. For autonomous transactions between strangers, the speakers say, the action needs cryptographic proof instead of a log entry [7]. They argue it reaches past payments: “this is a model that won’t just be used for payments, but we think it could be for any sort of high-stakes action that’s hard to reverse. So, medical orders, e-signatures, securities trading” [7]. PayPal describes an upcoming payments-signing mechanism as one concrete instance of the second-key idea for high-stakes cases, though it hasn’t shipped yet. What matters here is that two independent engineering teams landed on closely related rules from opposite starting points.

Figure 8 - Decision diagram asking whether an agent can undo its own action and whether the impact is acceptable, routing to log-and-go or a second key held by a human

Figure 8 - The Two Questions That Size Everything Else: Can the agent undo this by itself, and is the impact acceptable if it gets it wrong? Two yeses mean log it and move on. Either no means a human holds a second key.

The identity the agent can’t choose for itself#

Every primitive above assumes the identity attached to an action is the caller’s real identity. If that assumption breaks, they all collapse at once.

Malhotra’s closing point shows how. Say an agent can set its own identity in a request header. The moment it hits a rate limit, its cheapest fix is to relabel itself. It changes an ordinary header it was already allowed to set and gets a fresh budget under a new name. “In this case, you technically don’t have a rate limit. You just have a suggestion” [1]. His fix moves the stamping out of the agent’s reach. A proxy (a service that sits between the agent and the systems it calls) in the path of every outbound call holds the real credentials. It stamps each call with the identity it knows and ignores whatever the caller claims [1]. “That’s one rule. If you get that one rule right, everything else is just tuning” [1].

Figure 9 - Diagram of a proxy sitting between an agent and its outbound calls, stamping each call with the caller's real identity instead of trusting the identity the agent claims in the request

Figure 9 - The Proxy Stamps the Identity on Every Call: An agent that can set its own identity header can relabel itself for a fresh budget. A proxy holding the real credentials stamps every call, and the agent never gets to choose.

The same rule has a public, inspectable implementation one layer down, at the data layer. Google’s MCP Toolbox for Databases (an open-source tool server that connects agents to databases over MCP, the Model Context Protocol agents use to find and call tools) fills a tool’s identity parameter from a token the server validates itself, and ignores anything the model writes in its own tool call [8][9]. The value comes from that token’s verified claims (pieces of identity information, such as a user ID, inside a signed token whose signature the server has checked). A speaker on the project put it plainly: “it’s secured because we’re again extracting that user identity out of the agents control” [8]. In the talk’s end state, the model sees a tool like lookup_flights(date) with no identity argument at all. Where an implementation still lists the identity slot in the schema, whatever the model writes there is discarded [9]. A different team built the same rule at a different layer from Malhotra’s, which is a stronger signal than either instance alone.

It has two limits worth knowing before you rely on it. The parameter’s name can still appear in the tool list a model sees, even though its value is thrown away. That wastes a little context and can occasionally confuse which tool a model picks. It also only protects fixed, pre-written, parameterized tools. A tool that lets a model write its own free-form SQL has no fixed slot to bind identity into, so the model can still write a query that reaches another tenant’s row (a tenant is one customer’s isolated slice of a shared database). That gap needs its own enforcement below the model, which we cover in The Agent Can’t Guard Itself [10].

Figure 10 - Diagram of a tool server validating a signed token and filling an identity parameter from the token's verified claims, discarding any value the model tries to supply for that parameter

Figure 10 - The Model Never Fills the Identity Slot: The server validates the caller’s token and pulls the identity from its verified claims. Anything the model writes into that same field gets thrown away.

One creator, Hyperautomation Labs, documented a run on xAI’s Grok Bot in which a chief-of-staff agent hands work to single-job sub-agents. When one sub-agent tried to rewrite its own recurring schedule, the platform stopped the change and asked for explicit human sign-off, explaining that “the agent is persisting a new standing order with its own authored instructions and that requires explicit user authorization” [11]. That’s a platform enforcing the rule Malhotra argues for: an agent doesn’t get to rewrite its own permissions.

KEY INSIGHT: Ask where identity comes from for every write your agent performs. If the answer is “whatever the request says” or “whatever field the agent filled in,” none of your rate limits, verb grades, or approval gates are enforcing anything. They’re all reading a number the agent can change.

The dimension none of this catches#

Everything above assumes the enforcement itself is sound. A separate failure shows up when the number attached to a budget doesn’t mean what the person relying on it thinks it means.

Replit, a browser-based coding platform, walked through its new Routines feature (scheduled agent runs, described on air as beta) on a live stream. One host set a budget of “$2 per run”, and a co-host stepped in on air to clarify that the figure is the most a run may spend. A typical run, the co-host said, lands closer to 30 to 40 cents, and a separately demoed run cost 26 cents [12]. The stream never shows what happens when a run reaches that ceiling, so it’s not clear whether the platform enforces it or only displays it. What broke on air was legibility. The person setting the budget read a ceiling as a price, and a person who misreads a budget can’t use it to decide whether to keep going.

This failure belongs to the “who’s noticing” part of Malhotra’s budget. It’s easy to miss because the enforcement code may be working fine. The fix is to make sure the number a human sees means what they think it means before they rely on it.

Figure 11 - Diagram contrasting a displayed cost estimate that does not match what a system actually charges with the same budget after the number is corrected to be legible

Figure 11 - A Budget Only Works If the Number Is Read Correctly: The person setting a per-run ceiling read it as a per-run cost. A number the human misreads cannot guide the decision it exists for.

What this framework doesn’t cover#

This is one engineer’s framework, built from one company’s internal incident and recalled on a conference stage rather than published as an audited postmortem [1]. The proxy Malhotra describes is internal Anthropic tooling, with no public repository or documentation [1]. The public AWS sample built from Jamf’s design is a simplified stand-in for what the company runs internally, with two named production features left out [5]. PayPal’s second-key mechanism was described as upcoming and hasn’t shipped [7]. None of these sources gives a general benchmark for how many workloads is too many, or what a correct rate limit looks like for your system. Treat every number here as specific to the account it came from. This article also assumes an action’s risk tier is already decided. Scoring which actions deserve which tier is a separate question we cover in Risk-Graduated Agent Autonomy [13].

What the framework does give you is a small, testable set of questions for the moment you decide what an agent may do unsupervised.

What to ask before the next agent ships#

The filter bug started the cold-open incident, and bugs like it happen in ordinary code all the time. The damage came from a valid token being the only thing between that bug and about 200 workloads. A token has no way to say “a little, carefully, with someone watching.” A budget does.

For the next agent you give real write access to anything, four questions do most of the work before a line of enforcement gets written: how much can it do, how fast can it do it, what can it undo on its own, and who finds out when it can’t. One more question sits underneath all four. Is the identity on each action stamped by a layer the agent can’t edit, or is it something the agent reported about itself? If it’s the second, the rest of the budget is only a suggestion the agent happens to be following today.


References#

[1] S. Malhotra, “Give the Agent a Budget, Not a Token,” AI Engineer, YouTube, Aug. 22, 2026. https://www.youtube.com/watch?v=rbjWzZK2LU0

[2] “Sachin Malhotra,” AI Engineer speaker page. https://ai.engineer/speakers/sachin-malhotra

[3] E. Johanson, “How we built an autonomous SRE agent for Kubernetes,” LangChain Blog, Aug. 5, 2026. https://www.langchain.com/blog/how-we-build-an-autonomous-sre-agent-for-kubernetes-deployments

[4] Amazon Web Services and Jamf, “Tokenomics at Scale: How Jamf built real-time spend enforcement for Amazon Bedrock,” AWS Machine Learning Blog, Sep. 1, 2026. https://aws.amazon.com/blogs/machine-learning/tokenomics-at-scale-how-jamf-built-real-time-spend-enforcement-for-amazon-bedrock/

[5] aws-samples, “sample-bedrock-spend-enforcement,” GitHub repository. https://github.com/aws-samples/sample-bedrock-spend-enforcement

[6] Amazon Web Services, “Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock,” AWS Machine Learning Blog, Sep. 3, 2026. https://aws.amazon.com/blogs/machine-learning/set-up-openai-chatgpt-codex-with-litellm-on-amazon-ecs-and-amazon-bedrock/

[7] J. Mok and B. Coumes, “Your Agent Just Authorized What?!,” AI Engineer, YouTube, Sep. 1, 2026. https://www.youtube.com/watch?v=vGn6N4-bxBY

[8] A. Kitsch and P. Kakkar, “Build-Time vs. Run-Time: Why Dev Tools Fail in Production,” AI Engineer, YouTube, Sep. 9, 2026. https://www.youtube.com/watch?v=9R—1tg45Jg

[9] googleapis, “mcp-toolbox,” GitHub repository (Apache-2.0), v1.12.0, Sep. 17, 2026. https://github.com/googleapis/mcp-toolbox

[10] G. Dotzlaw, “The Agent Can’t Guard Itself,” Dotzlaw Consulting, Sep. 24, 2026. /insights/ai-47-the-agent-cant-guard-itself/

[11] Hyperautomation Labs, “My Grok Bot Chief of Staff Hired Its Own Team, I Metered Every Job,” YouTube, Sep. 1, 2026. https://www.youtube.com/watch?v=weaQjwLn6Q4

[12] Replit, “Free Mode + Routines in Replit (Live Deep Dive),” YouTube livestream, Aug. 26, 2026. https://www.youtube.com/watch?v=CaQ2FropN_I

[13] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “Risk-Graduated Agent Autonomy: Score Every Action by Blast Radius,” Dotzlaw Consulting, Sep. 4, 2026. /insights/ai-38-risk-graduated-agent-autonomy/

Give the Agent a Budget, Not a Token: Four Dimensions That Replace a Yes-or-No Grant
https://dotzlaw.com/insights/ai-50-give-the-agent-a-budget-not-a-token/
Author
Gary Dotzlaw
Published at
2026-09-30
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

Related reading

Beyond Code Completion: Building an AI Development Methodology with GitHub Copilot
GitHub Copilot suggests a line of code. Our enterprise codebase has 10,000+ functions across 22 modules. The gap between code completion and business context is where most AI adoption stalls. We closed it with 7 specialized agents, a code graph database, and a self-improving knowledge loop.
2026-03-22·GitHub Copilot
Self-Improving AI: How Code Reviews Feed a Knowledge Flywheel
Every code review harvests knowledge. Knowledge updates skills. Better skills produce better code. Eighteen domain skills and growing, each one making every Copilot agent smarter in that domain. Here is how we built a system that gets better every time someone uses it.
2026-03-25·GitHub Copilot
Self-Driving Products: From Observability Signals to Pull Requests
Two production teams at two different 2026 conferences described the same architecture: production signals in, pull requests out, and a human still clicking merge. Here is the loop, the grouping failure behind it, and the merge button nobody automated.
2026-08-10·AI & Modern Development
Multi-Tenant Agent Security: The LLM Is Not Your Security Boundary
A validator that inspects the SQL an LLM wrote is a probability, not a guarantee. Security in a multi-tenant analytics agent comes from the disallowed data being structurally absent from the model's context, proven here at three layers with our own re-runnable numbers, alongside five production architectures that land in the same place.
2026-08-25·AI Security
← Back to Insights