A team builds a document chat feature. Each customer, or tenant (one customer’s isolated slice of the shared system), uploads its own files. That tenant’s users ask questions, and the chatbot answers from those files only. It works in the demo. Then, before it ships, a security reviewer asks one question. Can a user logged into Tenant A’s chat window ever see an answer built from Tenant B’s document? “We filter by user” is the honest first answer. It doesn’t survive the follow-up. Where does that filter come from, and when does it run? One aviation software vendor sells a per-airline AI agent (a model that can call tools and take steps on its own) over its customers’ operational data. Its write-up gives the buyer’s side of this same conversation in one sentence: “Every question an agent answers is tied to one airline’s identity and one airline’s data, which is exactly what airline security and IT teams evaluating this kind of tool want to see” [1]. That sale closes only if the team behind the demo can explain how the isolation actually works.

Figure 1 - Identity in, filter rebuilt every hop, results re-checked: A validated token supplies the tenant’s identity on the server. That identity builds an explicit filter on every retrieval call the system makes. A second, independent check re-verifies what actually comes back before it reaches the user.
Why “we filter by user” isn’t an answer
A document chat feature runs on RAG (retrieval-augmented generation), so the language model doesn’t answer purely from what it was trained on. First, the system searches the customer’s own documents for the passages that fit the question, then hands those passages to the model along with the question. The documents get broken into chunks, small pieces of text, usually a paragraph or a few hundred words. Each chunk becomes an embedding, a list of numbers that captures what the chunk means, so chunks about similar topics end up close together numerically. The chunks and embeddings live in a vector store, a database that searches by that numeric closeness instead of by keyword. A metadata filter is a rule attached to a search. In effect, it says: only return chunks tagged with this tenant’s ID.
The problem sits in that filter. A filter is a rule the application chooses to apply. On its own, it can’t prove it was applied correctly, applied at all, or applied to the right identity. A standards body that tracks security risk in AI systems names this exact failure, and it’s on that body’s top-ten list for large language model applications: “in multi-tenant environments where multiple classes of users or applications share the same vector database, there’s a risk of context leakage between users or queries” [2]. We made a deeper version of this argument in an earlier article on multi-tenant SQL agents, which answer questions by writing SQL database queries [3]. There, the point was that a language model itself is never a reliable security boundary. Document retrieval raises a narrower question. The filter decides what the model even gets to see. Was that filter built from something the server actually checked?
Fix 1: trust the client, or trust the wording of the question
The chat interface already knows which user is logged in, so the easiest thing to build is a filter that trusts that value. Either the client sends a tenant ID or user ID along with the question. Or the retrieval system guesses scope from the wording itself, treating “show me my invoices” as an implicit filter on the asker’s own account. Both versions demo cleanly. The demo user only ever asks about their own data. Nobody in a demo tries to break it.
A value the client sends has an obvious problem. Nothing stops a request from arriving with the wrong ID in it. A bug can do it. So can a stale session, or someone editing the request by hand before it goes out. A client-side dropdown that lets a user pick “which account am I asking about” is a convenience feature. It offers no protection, because the user can pick any value. AWS’s own documentation for its knowledge base’s document-level access feature makes the same point about a value the application passes in. An ACL (access control list) is a list of who is allowed to touch a given resource. Under a boxed warning about ACL awareness, AWS’s page says the service “filters results based on the identity you supply but does not constitute true authorization… You must not rely on this feature as a sole access control mechanism without upstream authentication” [4].

Figure 2 - A value the client sends is a value the client can change: When the tenant or user ID that scopes a search comes from the request itself, whether typed by the UI or edited by hand, the retrieval system has no way to know the value was tampered with.
Amazon Bedrock’s managed knowledge base product is a common choice for this kind of document chat. It ships the second version of this problem as a feature: it can infer a filter from the wording of a question. AWS’s own reference architecture for multi-tenant document chat says plainly that “that inference is a relevance feature, not an access control” [5]. It makes search results more useful when a query happens to name the scope it wants. It was never built to answer the reviewer’s question: is the person asking allowed to see this result at all?

Figure 3 - Wording guesses serve relevance, identity checks serve authorization: A retrieval system that guesses scope from how a question is phrased is optimizing for relevant results. Authorization has to come from something the server already checked, separate from what the question happens to say.
KEY INSIGHT: Build the tenant filter from an identity the server has already verified. Scope guessed from a question’s wording only helps relevance.
Fix 2: add a real filter, once, at the top
Once the client-supplied filter turns out to be untrustworthy, the next move looks obvious. Stop trusting anything in the request. Build a server-side filter from the logged-in user’s own session the moment the question comes in. That closes the forged-ID problem cleanly. A request that never carried the real user’s identity can no longer smuggle in someone else’s.
This fix misses one thing. Modern document chat rarely answers a question with a single search. An agentic retriever (a retrieval system that plans its own searches instead of running one fixed query) breaks a question into several sub-searches. It follows up on what the first search turns up before deciding what to search for next. This is called multi-hop retrieval. It’s why document chat answers feel more complete than a single keyword search. AWS’s own observability walkthrough for this product describes the ratio: for every question, about two retrieval calls, three model calls, and five tool operations [6]. Each of those retrieval calls is a fresh chance to leak. A filter attached once, at the top of the request, only protects code that remembers to pass it into every one of those calls. Picture a new tool added later, a retry path, or a sub-agent (a helper agent that the main agent hands part of the job to). Any of them can call the retriever directly and skip the filter entirely, and the top-level answer looks no different. The demo still passes, because the demo’s happy path runs exactly the code that carries the filter.

Figure 4 - One question, two retrieval calls, then a follow-on hop: An agentic retriever decomposes a single question into roughly two retrieval calls on average, plus additional model and tool calls. A filter attached once, at the first call, has to survive being carried into every later call too.
The fix that holds: derive identity from a validated token, on every hop
This pattern survives a security review. It does two things differently from Fix 2. First, the identity behind the filter comes from a token the server has already validated, which means checked cryptographically to confirm nobody forged it. Most often that token is a JWT (JSON Web Token), a signed credential the server can check without trusting the client’s word for what it contains. It never comes from a session variable that something upstream could have set wrong. Second, the system rebuilds the filter from that same validated identity on every retrieval call it makes, no matter how many hops the question takes.
This is the reference architecture AWS documents for this class of product. Each document gets a per-user metadata tag at upload time. The server extracts a validated identity from a JWT. On every retrieval query, it rebuilds an explicit equals-filter on that tag from the same verified identity, never from anything in the request body [5]. One vendor documenting a pattern doesn’t prove the pattern is sound. We trust this one because the same discipline shows up on stacks that share nothing with Bedrock, or with each other.
A second retrieval vendor’s own documentation draws the same distinction, between a filter built from a client-supplied string and one checked against an identity token. It calls the first a “security trimming workaround” using “simple string comparison,” and offers the second as a token-validating check, still in preview. That check “validates the caller’s… token on each request and trims result sets to only the documents the caller is authorized to read” [7]. An open-source knowledge-base project enforces the same rule at three separate layers in its own code [8]. First, a check in the request router (the code that decides which function handles a URL) works out which tenant owns the requested resource before any handler runs. Next, a check in the application confirms the caller has been granted access, and it refuses a cross-tenant match unless an explicit share exists. Finally, tenant scoping is built into the database query itself. At a developer conference, a startup founder describes his company-brain product, which he says is deployed at customers including large banks. He states the same rule in one sentence: “that is done every single time… the agent is always using the user’s credential to read the right part of the wiki” [9]. A creator who built his demo in partnership with Oracle shows the same pattern live. His database and agent framework have nothing to do with Bedrock, and he says it just as plainly: “you need the gate to sit in the database… we need to have the access be determined by the system that is running remotely.” He asks the identical question as two different logged-in users. He gets two different answers. One is a flat refusal, with no hint of what exists on the other side [10].

Figure 5 - The server rebuilds the filter every time: Identity comes from a token the server has already validated. The filter built from that identity gets rebuilt fresh on every call in a multi-hop chain rather than reused, so no hop can fall back to an unscoped search.
Unrelated teams built these four: a wiki, a document search service, a document knowledge base, and a general-purpose database, and the same rule shows up on all four stacks. That suggests a real engineering principle. It doesn’t depend on one vendor’s API.

Figure 6 - Three layers, each checking on its own: An open-source document knowledge base enforces tenant scoping at the router, in the application’s own access-check code, and again in the database query itself. Each layer re-derives the scope independently, so a mistake in one does not automatically become a mistake in the others.
It only holds if every code path touching the retriever rebuilds the filter. No exceptions. A call site added later, or an internal tool that talks to the vector store directly, breaks the guarantee while the rest of the system still looks correct. So the system needs a second layer that doesn’t depend on the first one having worked.
Checking twice: the returned result, and the call itself
The reference architecture doesn’t stop at the filter. Before a retrieved chunk reaches the model’s context, the system checks the chunk’s own owner tag against the caller’s identity a second time. It discards anything that doesn’t match, even though the filter that ran moments earlier should already have excluded it [5]. The second check exists for the day a code path runs without the filter, and it catches the leak at the last point before the data reaches a user.
A separate AWS pattern applies the same idea one layer earlier, at the moment a tool is about to run. Each tool is a Lambda function (a small piece of server code that runs on demand), and a separate interceptor (a gateway component that screens every tool call first) sits in front of it. In AWS’s own words: “each tool Lambda re-checks the caller’s permissions server-side, so authorization holds even if the interceptor is misconfigured or bypassed” [11]. It’s the same principle: a second check that doesn’t trust the first, and this time it decides whether a tool call may happen at all. One check protects a retrieval result after the fact, and the other protects a tool call before it starts. With both in place, one misconfiguration in the filter or the interceptor can’t leak data on its own.

Figure 7 - Two checks, two different points in the request: A chunk re-check runs after retrieval and discards any result whose owner does not match the caller. A tool re-check runs before a tool call is allowed to happen at all. Neither assumes the layer before it worked correctly.
KEY INSIGHT: The filter narrows what gets retrieved. The re-check verifies what came back. Build both, because each covers the other’s most likely failure.
The document that isn’t searchable yet
A document doesn’t become searchable the instant someone uploads it. Uploading, parsing the file into text, generating embeddings for each chunk, and adding those chunks to the search index all take time, and a question can arrive while a document sits somewhere in that pipeline. The reference architecture tracks each document through five states [5]:
- Accepted, but not yet processing.
- Queued.
- Actively being parsed and embedded.
- Text-searchable, while richer processing still runs (tables and images inside a PDF, for instance).
- Fully indexed.
The vendor measured the timings below on an idle system with small files, under 5 MB (megabytes), and it says plainly they’re no guarantee, only an order-of-magnitude reference. A plain text file typically reaches full indexing in a few seconds, while a PDF with tables and images typically becomes text-searchable in 5 to 30 seconds and fully indexed in about 90 seconds [5].
Which of those five states the product’s own screen calls “ready” is also an isolation decision. Say the system marks a document searchable the instant it’s accepted. A question that arrives during parsing can then go one of two ways. It returns nothing, which looks like a bug to the user. Or worse, it hits a retrieval path that skips the identity-based scoping the fully-indexed path uses. That code assumed indexing had already finished. Treating a half-processed document as already inside the isolation boundary is the same mistake as Fix 1: trusting a state nobody actually checked.

Figure 8 - Searchable only from the fourth state: A document becomes queryable only at the fourth state, text-searchable. Every path that can reach it from there on must apply the same identity-scoped filter.
Proving it: adversarial questions a reviewer can actually run
A reviewer needs something to run. One AWS walkthrough for a related multi-tenant product builds exactly this kind of proof. For each persona (a type of user) in the system, a table pairs a query that should succeed with a query that should be refused [12]. Each row names the structural reason the refusal holds. None of them leans on a probability. The same pattern works for a document chat system. For a customer support product with per-tenant knowledge bases, the table might read:
| Persona | Query | Expected result | Why the result is guaranteed |
|---|---|---|---|
| Tenant A support agent | ”What’s our refund policy for late deliveries?” | An answer citing Tenant A’s own policy document | In scope: the chunk carries Tenant A’s identity tag, which matches the validated token |
| Tenant A support agent | ”What’s Tenant B’s refund policy?” | No results, or a refusal | The filter excludes every chunk not tagged for Tenant A. No chunk exists in scope to answer from |
| Tenant A support agent, forged tenant ID in the request body | Same question, with a spoofed Tenant B ID in the payload | Still no results, or a refusal | The filter is built from the validated token, not the request body. The spoofed value is never read |
| Tenant A support agent | ”Summarize every refund policy in the system” | Only Tenant A’s policy | The filter excludes all non-Tenant-A chunks before the search ranks results by relevance, however broad the question |
Running a short table like this against a real system, before calling it production-ready, turns “we filter by user” into a claim someone can check. Some teams can’t produce a table like this. For others, the “why” column reads “the model was told not to” where a structural reason should be. Either way, they haven’t built the pattern this article describes. They’re still relying on the model to behave, the same weakness as Fix 1.

Figure 9 - A table a reviewer can actually run: Each row pairs an in-scope query with an adversarial one and states why the refusal is structurally guaranteed. A “why” column that names a mechanism turns a reassurance into a claim someone can check.
Where this stops working
We read the open-source project’s source code. We didn’t run a live test. That read confirmed its router and application-layer checks are present [8]. We didn’t audit the vector-store and search-retriever layers underneath. A gap there would be a real bypass, and our read wouldn’t have caught it. The same project runs code in a processing sandbox (an isolated container), and by default the sandbox’s network policy allows outbound traffic (network traffic leaving that container). Operators have to turn on the stricter setting themselves [8].
A separate platform-engineering write-up centralized identity checks behind a single identity gateway (one service every request passes through). It reports a 55% reduction in repeated login events, where users had to sign in again separately for each internal tool [13]. It also names a pitfall worth repeating on its own: “strip inbound identity headers before injecting trusted ones” [13]. Headers are the labels on a web request that can carry who the user is, so a client should never be able to claim an identity by setting one itself. That 55% figure is the vendor’s own number. It’s unaudited and self-reported, and we cite it as such.
We have shipped a server-side isolation guarantee ourselves, at a different layer than the one this article covers. Our own text-to-SQL engine turns a plain-English question into a database query, and it enforces tenant scoping by removing disallowed rows from a query’s context entirely. That’s the SQL-analytics half of this same family of problem [3]. We don’t have an equivalent first-party measurement for document retrieval yet, and this article doesn’t claim one. Everything here is a documented pattern, backed by several independent implementations from different vendors. We didn’t measure it ourselves.
Conclusion
The rule that survives a security review has three parts. Identity comes from something the server has already checked. The filter built from that identity is rebuilt on every retrieval call a question triggers. A second, independent layer re-checks what actually comes back, or what a tool is about to do.
Before calling a per-tenant document chat feature done, or before accepting a vendor’s or contractor’s word that isolation is handled, ask three questions. Where does the identity that scopes each query come from, and can it be traced to something the server validated? Is the filter rebuilt server-side on every retrieval call the system makes, including the later hops? What independently re-checks the results, or the action, if the filter is ever wrong? A team that answers all three with a named mechanism has built the thing its buyers are actually asking about. A sentence about filtering by user doesn’t count. The shape of the authorization grant itself is a separate design question, which we cover elsewhere [14].
If your team is putting a per-user document assistant in front of customer data, do this work before launch. A scoped design pass covers identity derivation, the per-hop filter contract, and the defense-in-depth check. It tests each one against your own retrieval calls instead of assuming a vendor’s reference architecture fits.

Figure 10 - Three questions that separate a design from a demo: Where the identity comes from, whether the filter is rebuilt on every retrieval call, and what independently re-checks the result. A team that can answer all three with a named mechanism has built the pattern this article describes.
References
[1] P. Lafond, M. Cardinaels, and N. Kheir, “How AvioBook builds turnaround insights from operational data with Amazon Bedrock AgentCore,” AWS Machine Learning Blog, Sep. 10, 2026. https://aws.amazon.com/blogs/machine-learning/how-aviobook-uses-generative-ai-to-drive-airline-turnaround-insights/
[2] OWASP Gen AI Security Project, “LLM08:2025 Vector and Embedding Weaknesses,” 2025. https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/
[3] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “Multi-Tenant Agent Security: The LLM Is Not Your Security Boundary,” 2026. /insights/ai-26-multi-tenant-agent-security/
[4] Amazon Web Services, “Document-level access controls,” Amazon Bedrock User Guide, accessed Sep. 2026. https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-ds-s3-acl.html
[5] G. Belsian, D. Mitchell, and O. Elkharbotly, “Build multi-tenant agentic chat applications on enterprise data with Amazon Bedrock Managed Knowledge Base,” AWS Machine Learning Blog, Aug. 31, 2026. https://aws.amazon.com/blogs/machine-learning/build-multi-tenant-agentic-chat-applications-on-enterprise-data-with-amazon-bedrock-managed-knowledge-base/
[6] L. F. Yepez Barrios, D. V. Batalov, and S. Singh, “Build observable enterprise agentic retrieval using managed Amazon Bedrock Knowledge Base with AWS CloudFormation,” AWS Machine Learning Blog, Aug. 31, 2026. https://aws.amazon.com/blogs/machine-learning/build-observable-enterprise-agentic-retrieval-using-managed-amazon-bedrock-knowledge-base-with-aws-cloudformation/
[7] Microsoft, “Document-level access control in Azure AI Search,” Microsoft Learn, updated Sep. 17, 2026. https://learn.microsoft.com/en-us/azure/search/search-document-level-access-overview
[8] Tencent, “WeKnora,” GitHub repository, v0.8.2, MIT License, accessed Sep. 2026. https://github.com/Tencent/WeKnora
[9] T. Gopal, “Your company brain will leak secrets: how we stopped it for big banks,” AI Engineer, YouTube, Sep. 2026. https://www.youtube.com/watch?v=0uC6u0lJJl4
[10] C. Medin, “You Built Your AI Second Brain. Now What? (Here’s How to Evolve It),” YouTube, Sep. 2026. https://www.youtube.com/watch?v=mjQlZrteMIY
[11] A. Sibanda, D. Perez Caparros, N. Goyal, J. DeMuth, and R. Diez Lejarazu, “Implementing defense-in-depth authorization for MCP tools on Amazon Quick,” AWS Machine Learning Blog, Sep. 17, 2026. https://aws.amazon.com/blogs/machine-learning/implementing-defense-in-depth-authorization-for-mcp-tools-on-amazon-quick/
[12] A. Ambavane, D. Paruchuri, V. Elangovan, and P. Sadhu, “Securing Amazon Quick from POC to production: Agents, Flows, and Spaces,” AWS Machine Learning Blog, Sep. 1, 2026. https://aws.amazon.com/blogs/machine-learning/securing-amazon-quick-from-poc-to-production-agents-flows-and-spaces/
[13] B. Khemchandani and R. Somvanshi, “How to Carry User Identity Across Federated Kubernetes and AI Platforms,” NVIDIA Technical Blog, Sep. 3, 2026. https://developer.nvidia.com/blog/how-to-carry-user-identity-across-federated-kubernetes-and-ai-platforms/
[14] G. Dotzlaw, “Give the Agent a Budget, Not a Token: Four Dimensions That Replace a Yes-or-No Grant,” 2026. /insights/ai-50-give-the-agent-a-budget-not-a-token/
Building production AI, or modernizing a legacy system?
That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.