3816 words
19 minutes
The Bounded Portal: How Coding Agents Cross the Knowledge-Work Adoption Gap

A company’s coding agents (models that read a task, run a tool, check the result, and go again) work well enough that someone in leadership asks the obvious next question. Why not point the same kind of agent at sales, support, or HR (human resources)? Somebody tries it, and it fails in a way the team hasn’t seen before. Karan Vaidya described exactly this at an AI Engineer conference session. He is co-founder and chief technology officer of a company that sells infrastructure for this problem. We use his diagnosis here but not his product, and the fix below is something a team builds itself. He pointed his own general-purpose AI agent at a hiring-outreach task: mass emails to candidates. It did exactly what he asked. “It did exactly what I told it to do, it was also a disaster, the kind that ends up on Twitter with my name on top of it,” he said [1]. Nothing crashed. No test failed. The agent finished the job it was given, and the outcome was still wrong.

So an agent can do precisely what it was told, break nothing a test could catch, and still cause real damage.

Figure 1 - Diagram contrasting a non-technical user talking directly to a raw coding agent with the same agent wrapped in a narrow guardrailed portal that sits between the user and production systems

Figure 1 - A Raw Agent vs. A Bounded Portal: On the left, a non-technical user talks directly to a general-purpose coding agent with open-ended reach. On the right, the same underlying agent sits behind a narrow, purpose-built surface: a curated set of skills, a scoped task boundary, and a review process standing between the user and anything that could break.


Every technical check passed anyway#

Most engineers, hearing “an AI agent caused a disaster,” picture a bug: a hallucinated fact, a malformed API call, a step the agent skipped. Those failures have a fix, usually a stricter check somewhere in the pipeline. Vaidya’s failure had no bug to fix. The emails were valid. The addresses were real and reachable. A linter, a compiler, a test suite: every check a coding pipeline normally runs would have waved this task through. None of them was built to ask the one question that mattered. In Vaidya’s own words: “there was no best tool in the world to actually question what really mattered. Should this have gone at all?” [1]

No compiler or test can answer whether an action should happen at all. That takes judgment no pipeline check supplies, and non-technical staff are even less able to apply it to a system they didn’t build.

Figure 2 - Diagram showing a task sent to an AI agent passing every technical check, formatting, address validity, and delivery, yet still causing harm because no check asked if it should happen.

Figure 2 - Every Check Passed, and It Was Still Wrong: The task went out correctly formatted, to real addresses, and delivered without error. Every check a coding pipeline normally runs passed. None of them were built to ask whether the task should have been attempted at all.

It happened to someone whose job is watching for exactly this#

You could write off the hiring-outreach failure as one person’s bad prompt. A second incident, reported separately, says otherwise. Summer Yue, director of alignment at Meta Superintelligence Labs [2], had a standing instruction: confirm before deleting anything in her inbox. The rule didn’t hold when it mattered. Vaidya, retelling the incident on stage, says roughly 200 emails were gone by the time she reached a computer to stop it [1]. We cover how a rule like hers gets lost inside an agent’s own memory in The Agent Can’t Guard Itself [3]. The point here is simpler. A deleted email has no undo, so the only chance to catch the mistake was before it happened, and nothing was in place to catch it there either.

If the person whose job is AI alignment can lose a rule this way, the missing piece sits in the environment around the agent. Writing the rule better won’t supply it.

Two obvious answers that fall short#

Two explanations come up first, and both are reasonable.

The first treats it as a training or judgment problem: coach staff on what to check before an agent acts, and the failures go away. That helps at the margins, but it leaves the structural gap alone. Aaron Levie, CEO of Box, a cloud content-management company, explained why coding agents made it into production use while most other knowledge-work agents stalled. He spoke at an event run by LangChain, an agent-framework company [4]. His answer rests on three properties code has and most knowledge work lacks. Levie also counts model maturity outside code as a factor, but a better model doesn’t remove the three structural properties [4].

First, code is verifiable. You can run it and see whether it works. That gives you an oracle (a fast, objective way to check whether a specific output is right), and most business tasks don’t have one. Second, engineers are hyper-technical users who can sanity-check what an agent is asking to do. Most staff in sales, marketing, or HR can’t make that call about a tool they didn’t build. Third, engineers already hold broad access to the systems they work in. A knowledge-work agent inherits whatever narrow, inconsistent permissions its human happens to have, and unlike the human, it can’t ask a colleague or escalate.

A training class fixes none of those three. A better-trained salesperson still has no compiler to run against a sent email.

Figure 3 - Comparison chart of three adoption properties, verifiable output, technical user, and broad access, showing coding satisfies all three and most knowledge work satisfies none.

Figure 3 - Three Properties Code Has and Most Knowledge Work Doesn’t: Verifiable output, a technical user who can judge what the agent is doing, and broad access already granted. Coding satisfies all three by default. Most knowledge work starts with none of them.

The second answer is to assume the model isn’t ready yet and wait for a smarter one. Anthropic studied roughly 400,000 Claude Code (Anthropic’s coding agent) sessions from about 235,000 people. The results suggest the bottleneck is often the person specifying the work, and not only the model. People made about 70% of the planning decisions, and the agent handled about 80% of the execution. The same tool produced very different amounts of work depending on who was directing it. Expert sessions averaged about 12 agent actions per instruction against about 5 for novices. Anthropic ties the difference to how precisely the user frames the job [5]. OpenAI reports that since February, weekly active enterprise Codex (OpenAI’s coding agent) users grew 108x in legal, 41x in sales, 41x in recruiting, and 26x in marketing, compared with 5x in engineering [6]. Adoption is moving fastest outside engineering, so the model can’t be what holds those teams back. Legal leads, and legal is the most checkable of the four.

Workflow redesign is a separate question. Enterprise Workflow Redesign argues that rethinking a process buys more than automating one step of it [7]. That problem is real. The failures above, though, happened even when the task handed to the agent was exactly the right thing to delegate. The problem sat in the environment around the workflow.

KEY INSIGHT: If your fix for an agent failure is “train people better” or “wait for a smarter model,” first check whether the domain has an oracle at all, meaning a fast way to know if a specific output is right or wrong. Neither fix helps a domain that has none.

What a codebase has that a sales deal doesn’t#

Put the failures next to an ordinary day of software engineering and the missing piece gets concrete. A coding agent starts every task inside a repository that holds, in Vaidya’s phrase, “the what, the why and how” of the system in one place [1]. A sales deal has nothing like that. It’s scattered across a CRM (customer relationship management system), a shared drive, an inbox, a chat thread, and a support ticket, each behind its own login. Vaidya’s own example: “coding agents work so pretty well partly because they were very near the source of truth … This is exactly what knowledge work miss today, for example a single deal is scattered across five different platforms” [1].

A codebase also keeps a history nobody has to write down on purpose. Git records every change, who made it, and usually why. An agent, or a human checking its work, can look back at how the current state came to be. Knowledge work keeps almost none of that. Ask what led a CRM record to its current state, or how a colleague wrote the email that closed a deal. Nobody knows, because nothing kept a record.

Figure 4 - Diagram comparing a codebase's built-in source of truth, history, automatic checks, and revert command against a sales deal scattered across five separate logins with none of them.

Figure 4 - What a Codebase Has That a Sales Deal Doesn’t: A codebase ships a single source of truth, a full history, automatic checks, and a revert command, none of it built for AI specifically. A sales deal, scattered across five separate logins, has none of the four.

Can the mistake be taken back?#

Everything above points at the same missing piece, and Vaidya says it directly: “In code, you can undo the mistake after it happens. Here, you catch it before it does” [1]. In a codebase, a bad change gets reverted. A team can let an agent act, check the result, and undo it if it’s wrong, since undoing is cheap. Most knowledge-work actions have no undo. A sent email can’t be unsent. A wire transfer can’t be pulled back [1]. Once an agent’s action can’t be undone, the bar goes up: the agent has to be right before anyone gets to check.

That changes what “verifiable” means in practice. The real test is whether a specific action can be taken back. Code usually can be, which is why coding agents get a benefit of the doubt other domains don’t.

Figure 5 - Diagram contrasting two failure models, in code a mistake happens and is undone afterward, in most knowledge work a mistake cannot be undone so it has to be caught before it happens

Figure 5 - Undo After, or Catch Before: In code, a team lets an agent act, checks the result, and reverts if it’s wrong. Where nothing can be undone, that order flips: the only chance to catch a mistake is before it happens, which demands certainty a coding agent never needs.

Reversibility can fail inside code too, and one real incident shows how badly. A Cursor (an AI code editor) agent running Claude Opus 4.6 was working in the codebase of PocketOS, a software company whose platform runs car-rental businesses. In an unrelated file, it found an account-scoped API token (a credential that lets software act on an account) for Railway (a hosting platform) [8]. After hitting a credential mismatch, it decided on its own to delete a storage volume [8]. Railway stores backups on the same volume it’s protecting. Deleting the volume took the production database and every backup with it, in 9 seconds [8]. A security researcher later reconstructed the timeline. By that account, PocketOS got most of its data back roughly 30 hours later, when Railway’s CEO confirmed recovery had succeeded. Until then, the newest backup PocketOS held off the volume was 3 months old [8]. This was a coding agent, in a codebase, with revert commands available the whole time. None of that helped, since the specific action it took had no undo.

Figure 6 - Timeline diagram of a coding agent deleting a storage volume that held both a production database and its backups in 9 seconds, with recovery about 30 hours later.

Figure 6 - When Reversibility Fails Even in Code: A coding agent found a stray API token, hit a credential mismatch, and deleted a storage volume that happened to hold both the database and its backups. 9 seconds, no undo. Railway confirmed a recovery roughly 30 hours later. The company’s own newest off-volume backup was 3 months old.

KEY INSIGHT: Whether a mistake can be undone depends on the specific action. It doesn’t come free with being “in code.” Before trusting an agent to act unsupervised, ask whether this particular action can be undone, not whether the domain generally can.

A narrow surface instead of a raw agent#

With the gap named, the fix is no longer “scope down the permissions and hope” or “train people harder.” Build the missing infrastructure into a narrow surface, and have non-technical staff work through that surface instead of talking to the raw agent. We call this a bounded portal: a purpose-built interface, wrapped around a general coding agent, that limits what the agent can see, which tasks it will attempt, and what it can do without a human checking first.

Inside a working portal, three pieces do the job the codebase used to do for free. The first is a curated set of skills (packaged instructions and scripts that tell an agent how to do one kind of task). Skills stand in for the shape and style a codebase carries naturally, so the agent isn’t guessing what a task should look like. The second is a scoped task boundary, which keeps the agent’s reach narrow enough that a human can actually judge what it’s being asked to do. That’s the tiered-permission scoring Risk-Graduated Agent Autonomy argues for, sized to how much could go wrong [9]. The third is a review process that controls what gets added to that scoped surface over time. It works more like a specific, sized authorization than a single yes-or-no credential, the shape Give the Agent a Budget, Not a Token describes in detail [10].

Figure 7 - Architecture diagram of a bounded portal's curated skill layer, scoped task boundary, and review process wrapped around a coding agent, with the user interacting only with the portal.

Figure 7 - What’s Actually Inside a Portal: A curated skill layer stands in for a codebase’s shape and style. A scoped task boundary keeps the agent’s reach narrow enough for a human to judge. A review process governs what gets added over time. The user never talks to the raw agent directly.

Four teams built one, for different reasons#

Four teams took on this gap, each for a slightly different reason. Three built a narrow surface, and one prepared its people instead. Treat each as an illustration of how the mechanism works. None of them proves it produces measured gains, since every figure below comes from the organization that built the thing being described.

loveholidays, an online travel agency, built an internal portal its own engineers call “search playground,” which uses the company’s own design system and has a coding agent behind it. A non-engineer doesn’t talk to the agent directly. They talk to a bounded product surface. In OpenAI’s own video, loveholidays’ CTO reports “more than 10 different novel search experiences,” the “majority… developed by non-engineers” [11]. That’s loveholidays attacking the technicality friction, where users can’t judge what a raw agent is doing. The portal gave non-engineers a surface they could judge.

IMDEX, a mining-technology company, went after the same friction another way. Instead of building a bounded UI, it built a structured curriculum. Cursor’s own blog post about IMDEX describes a month-long training program, internally branded “AI July,” that ran basic agent training, advanced tooling training, and applied hackathons [12]. The reusable part is the curriculum: a staged way to raise a non-technical user’s judgment before handing them a tool.

Snowflake went after a different friction. The team was rolling an internal go-to-market (sales and marketing) assistant out to a 6,000-person sales and marketing organization, and it deliberately capped what the assistant would attempt before expanding: “we don’t want to try to answer 100 questions and get them 70% right. We want to answer 50 questions, but get them 95% right,” in the words of Sait Izmit, the Snowflake product manager who led the rollout [13]. That’s the verifiability friction, attacked from the scope side. Rather than making every answer checkable, the team shrank the surface until the answers it gave were reliably right. It took the beta to about 10% of the organization. Izmit says the team exited beta only once weekly-active-user retention passed 70% [13], a figure Snowflake’s own blog also reports [14]. Only then did it widen to the full audience.

Figure 8 - Timeline diagram of Snowflake's phased rollout: pilot, beta gated on 70% weekly retention, then full rollout, with the rule 50 questions at 95% beats 100 at 70%.

Figure 8 - Snowflake’s Phased Gate: A pilot proves accuracy with a small, engaged group. A beta expands to about 10% of the org and only advances once weekly retention clears 70%. Only then does the full rollout happen. The launch rule underneath every phase: 50 questions answered at 95% beats 100 questions answered at 70%.

Cloudflare runs an internal portal for its go-to-market team called Cloudflare OS. Where the others used a portal, a curriculum, or a scope cap, Cloudflare’s fix is a governance process. Justin Joyce, who works in sales operations and strategy at Cloudflare, describes a central review step. In his words: “we have a central alias where skills are presented to the central team, curated by the go-to-market team as well as by operations team, and they’re reviewed, so we can make sure that we’re not having a proliferation of skills” [15]. Our reading: that’s the third friction, permission scope, handled as an ongoing process rather than a one-time scoping decision.

Figure 9 - Matrix diagram mapping loveholidays, IMDEX, Snowflake, and Cloudflare to the adoption friction, technicality, verifiability, or permission scope, each one's portal addresses.

Figure 9 - Four Teams, Different Frictions: loveholidays and IMDEX both attack the technicality friction, one architecturally and one through training. Snowflake attacks verifiability by capping scope before expanding it. Cloudflare attacks the ongoing cost of permission scope with a standing review process.

The review never stops#

Cloudflare’s description names something the other three leave unsaid: a bounded portal needs review for as long as it exists. Someone reviews every skill that gets added to it. That’s a real, ongoing cost, and anyone planning a portal should budget for it from day one.

Every instance in this piece has the same limit. All four accounts come from the organization itself or from the vendor whose tool it uses. None has an independent audit or a stated method behind any specific figure. So none of them proves the pattern produces measured gains. Each one describes what closing the gap looked like in one company. We also have no confirmed small-business-scale instance yet, and no client engagement of our own to report. Both are gaps. The four instances still count, since four unrelated teams arrived at the same answer.

KEY INSIGHT: Before designing a portal, name which of the three frictions your own rollout actually has. Output nobody can check and access nobody has granted need different fixes. Building the wrong one uses up the review capacity the real fix needs.

Which department goes first#

Reversibility answers a question the frameworks above leave open. Once a company decides to build a bounded portal for a second department, which department goes first? Start where a mistake is cheap to take back: an internal search tool, a draft a human reviews before it goes out, a report someone checks before acting on it. Anywhere a mistake can’t be undone (a sent communication, a financial transaction, a deleted record) needs the portal’s infrastructure built and proven before an agent gets near it. A prompt written after something goes wrong comes too late.

Figure 10 - Diagram sequencing candidate departments for a bounded-portal rollout from easily reversible tasks to irreversible ones like a sent communication or deleted record.

Figure 10 - Which Department Goes First: Order the rollout by what can be undone, not by which team asks loudest. An internal search tool or a reviewed draft can go early. A sent communication, a financial transaction, or a deleted record needs the portal’s infrastructure proven first.

Vaidya’s own closing line: “for 2 years, the model was the bottleneck… now it’s infrastructure that nobody has yet built” [1]. The model that writes production code can already draft outreach, summarize a deal, or triage a ticket. Most companies already have that capability. What they’re missing is the narrow, guardrailed surface that gives a non-technical user, and the model working with them, the same floor a codebase has always given engineers for free.

Conclusion#

In every instance here, narrowing what the agent could do is what let more people use it safely. A bounded portal supplies what a codebase ships for free and a sales deal never has: the source of truth, the history, the checkable scope, and the reversible action.

For a team whose coding agents already work and whose next question is “can marketing have this too,” the design work has the same shape every time. Name which of the three frictions the new department actually has. Build the narrow surface that supplies what’s missing for that specific friction, not a generic scaled-down agent. Sequence the rollout by what can be undone, and wherever it can’t be, build the portal before anyone hands the agent the keys.


References#

[1] K. Vaidya, “From coding to Knowledge work agents,” AI Engineer World’s Fair 2026, June 30, 2026 (video posted Sep. 3, 2026). https://www.youtube.com/watch?v=xxfMT-bPEmU

[2] Z. Stone, “She runs AI safety at Meta. Her AI agent still went rogue,” The San Francisco Standard, Feb. 25, 2026. https://sfstandard.com/2026/02/25/openclaw-goes-rogue/

[3] G. Dotzlaw, “The Agent Can’t Guard Itself,” Dotzlaw Consulting, Sep. 24, 2026. /insights/ai-47-the-agent-cant-guard-itself/

[4] A. Levie and H. Chase, “Why Enterprise AI Adoption Is Slower Than You Think,” LangChain event, YouTube, June 29, 2026. https://www.youtube.com/watch?v=agSRMrhNTf4

[5] Z. Hitzig et al., “How Claude Code is used in practice,” Anthropic Research, June 16, 2026. https://www.anthropic.com/research/claude-code-expertise

[6] OpenAI, “Enterprise Signals,” OpenAI, 2026. https://openai.com/signals/enterprise-data/

[7] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “Enterprise Workflow Redesign: Point Insertions Buy Efficiency, Redesign Buys Growth,” Dotzlaw Consulting, Aug. 18, 2026. /insights/ai-22-enterprise-workflow-redesign/

[8] A. Pignati, “A Security Post-Mortem of the 9-Second AI Database Deletion,” NeuralTrust Blog, Apr. 28, 2026. https://neuraltrust.ai/blog/pocketos-railway-agent

[9] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “Risk-Graduated Agent Autonomy: Score Every Action by Blast Radius,” Dotzlaw Consulting, Sep. 4, 2026. /insights/ai-38-risk-graduated-agent-autonomy/

[10] G. Dotzlaw, “Give the Agent a Budget, Not a Token: Four Dimensions That Replace a Yes-or-No Grant,” Dotzlaw Consulting, Sep. 30, 2026. /insights/ai-50-give-the-agent-a-budget-not-a-token/

[11] OpenAI, “What Codex Unlocks for loveholidays,” YouTube, Aug. 26, 2026. https://www.youtube.com/watch?v=o38xYi2mtgc

[12] Cursor, “IMDEX uses Cursor to build integrated subsurface data and analytics platform in months, not years,” Cursor Blog, Aug. 25, 2026. https://cursor.com/blog/imdex

[13] S. Izmit, “Building GTM AI Agents: Lessons from Deploying to 6,000 Users,” AI Engineer, YouTube, Aug. 26, 2026. https://www.youtube.com/watch?v=DrTdD-ttjCY

[14] S. Izmit, “From Pilot to 6,000 Users: How to Scale Enterprise AI Agents,” Snowflake Blog, Feb. 16, 2026. https://www.snowflake.com/en/blog/scale-enterprise-agents/

[15] J. Joyce, “How AI Agents Let GTM Teams Scale,” AI Engineer, YouTube, Aug. 26, 2026. https://www.youtube.com/watch?v=Qw_tC68KKes

The Bounded Portal: How Coding Agents Cross the Knowledge-Work Adoption Gap
https://dotzlaw.com/insights/ai-52-bounded-portal-adoption-gap/
Author
Gary Dotzlaw
Published at
2026-10-05
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

← Back to Insights