A team ships a search feature over their own documents. It’s RAG (retrieval-augmented generation): find the relevant text, then hand it to a language model so it can answer a question. A user types a question a human would answer in ten seconds. The system comes back with the wrong page. The user says some version of “that’s literally the answer, why didn’t it find it.” So the team runs an eval, a scored test set of questions paired with the document that should answer each one, and the recall number comes back mediocre. Recall@100 is the fraction of questions where the correct document showed up anywhere in the top 100 results.
When that number is low, the blame lands on the retriever, the part of the system that searches the document index and hands back candidates. So the team goes shopping. First comes a newer embedding model, a model that turns text into a list of numbers so a computer can measure how close two pieces of text are in meaning. Maybe a reranker follows. That’s a second model that takes the retriever’s top candidates and re-sorts them by relevance. Both cost money. Both nudge the eval number. The complaints keep coming. Nobody opens the code that decided what the retriever could find in the first place: the code that took each raw document and turned it into the text sitting in the index. It was written once, early in the project, filed under “preprocessing,” and never opened again.

Figure 1 - Diagnose in English, fix in code: An analysis agent compares a failing query against the document that should have matched it and writes down why in a sentence. A code agent reads that sentence and edits the program that builds the search index. Only changes that measurably help get kept, and the loop repeats.
The two moves everyone reaches for first
The first instinct is to blame the matching function itself. A dense retriever searches by embedding. It turns the query and every document into vectors (lists of numbers) and returns the documents whose vectors sit closest to the query’s. A lexical retriever matches on the words themselves. The standard one is BM25 (a scoring formula that counts matching words and gives rare words more weight). Dense retrieval usually wins on raw accuracy. The paper that popularized it reported a dense retriever beating a classic lexical baseline by 9 to 19 percentage points on general-knowledge questions, measured by whether the right passage appeared in the top 20 results [1]. So the first fix is a better embedding model. A smarter matching function should find more of the right documents.
It can help, a little. What it can’t do is explain a specific failure. If a query still misses after the upgrade, the new model hands you a different wrong ranking with a different score attached. Nobody on the team can point at a line and say why this document lost to that one.
The second fix is a reranker. It re-sorts the top candidates the first-stage retriever already found before anything reaches the user, which cleans up the order among documents that made the cut. It does nothing for a document that didn’t. If the answer never entered the top 100 candidates, the reranker never sees it.
Both fixes target the matching function. Neither touches the step before matching starts: the code that decided what each document would look like once it landed in the index.

Figure 2 - The step nobody revisits: A document moves through representation code before it ever reaches the retriever or the reranker. Teams tune the retriever and the reranker often. The representation code that decided what the retriever could see gets written once and left alone.
Four teams patch the same problem
Once a team sees that the real gap sits upstream of the retriever, the usual next move is to enrich the documents before they get embedded, so the indexed text looks more like what a user would type.
Weaviate, a vector-database company, published a build log for an internal demo called Foundry that shows this for a single file. Weaviate also runs the podcast where AutoIndex comes up later in this article, so Foundry and that podcast count as one company’s view, not two independent ones. The test archive was full of the filenames a real creative team leaves behind: final_FINAL_v7.svg, BROLL_NEW2.svg, scene_14_USE_THIS.svg. A file named rain-floor.jpg stays invisible to a search for “rain striking a reflective floor” until something attaches meaning to it. Foundry’s preparation stage adds a description, tags, extracted text, a relationship role, and a source link (a URI, the address that points back to the original file) to every file before anything gets embedded [2]. Afterward, a query for a rabbit in a grassy landscape returns a rabbit reference image and a related trailer, and a filename search would never surface either one [2]. In the post’s example, it works. By the post’s own account, it’s also driven by a manifest, a small file a person wrote by hand listing each asset’s description and tags, so every demo run gives the same result [2]. Nothing about it updates itself when the corpus changes or a new kind of query starts failing.

Figure 3 - What “representation” adds to one file: The raw file is named rain-floor.jpg, invisible to a natural-language search. The enriched record adds a description, tags, a relationship role, and a source link before embedding, which is what makes the same file findable by a query like “rain striking a reflective floor.”
Reducto, a document-parsing vendor, found the same problem one layer down, inside a single table. A model reasoning over a table does well when it gets the table as HTML, because the structure (which cells are merged, which header applies to which column) carries real meaning. Adit Abraham, Reducto’s co-founder and CEO, described the catch in an AI Engineer conference talk: “embedding models really struggle to correlate that natural language human prompt with this messy blob of HTML tags and numbers” [3]. Nobody asking “how did revenue change over time” types table markup. Reducto’s fix makes two representations of the same table from one extraction pass: HTML for the model that reasons over it, and a plain-language summary built for the embedding index [3]. “A big thing that you can do that actually takes very little effort is creating a representation that’s more so designed for the embedding model itself,” Abraham said [3].

Figure 4 - One table, two jobs: Reducto extracts a table once and produces two representations from it. The HTML version keeps the structure a reasoning model needs. The plain-language summary is what actually gets embedded, because a user’s question never looks like table markup.
LangChain rebuilt its own public documentation chatbot after deciding that chunking, the standard practice of splitting a document into fixed-size pieces before embedding, was working against it. In their words: “Chunking breaks structure. When you chop documentation into 500-token fragments, you lose headers, subsections, and context” [4]. Their fix gave the agent direct access to full pages and to the existing structure of their docs, so it “didn’t need smarter retrieval” [4]. That’s a different kind of move. Instead of representing documents better for a retriever, they dropped the retrieval step for content already organized well enough to search directly. As of November 2025, LangChain says embeddings still suit unstructured content. This fix is for organized material like documentation and code [4].
Anthropic’s Contextual Retrieval is a fourth version of the idea, applied more generally. Before a chunk gets embedded, a model prepends a short paragraph saying what the chunk is and where it sits in the larger document. Anthropic reports it cut failed retrievals by 49%, and by 67% when combined with reranking [5]. A model writes the added context here. A person still decided, once, what that context should be.
Four companies, working at different layers of the pipeline, reached the same diagnosis: the retriever was fine, and the representation feeding it was the weak spot. In every case a person decided the enrichment once. Nothing watches for new failures and changes the code in response.
KEY INSIGHT: When a search feature can’t find an obvious document, check what the indexing code actually kept from that document before blaming the retriever or the ranking model. A better matching function can’t recover information that never made it into the index.
Letting an agent rewrite the indexing code
A piece of research called AutoIndex sets out to close that gap. Instead of a person noticing a failure and hand-editing the enrichment code once, a loop does the noticing and the editing, and keeps doing it. AutoIndex is a paper from a team at the University of Massachusetts Amherst, with a co-author from Databricks Mosaic Research, posted in July 2026 to arXiv, the public server where researchers share papers before formal review [6]. It’s a preprint. It hasn’t gone through peer review at a conference or journal yet, so read it as the researchers’ own account of their own work.
The paper’s core object is what it calls a representation program. That’s a small piece of code (in the experiments, a single Python file) that takes a raw document and turns it into the units a retriever actually searches: chunks, enriched text, reweighted sections, or something else entirely [6]. AutoIndex holds the retriever fixed and optimizes the representation program instead. You stop tuning the search engine. You tune the code that decides what it gets to see.
The system runs two AI agents in a loop. Each one is a large language model (LLM), the kind of model behind a chatbot, wired up with specific tools and a specific job. The analysis agent gets read-only tools: it can run a test search against the current index, read the source document a query should have matched, and search for specific terms across documents. The paper calls that should-have-matched document the “gold document,” an information-retrieval term for the correct answer in the test set. The agent investigates a batch of failing and succeeding queries. Then it writes a plain-language summary that names a cause in words. In one case from the paper, repeated LaTeX markup in the gold documents was drowning out the words the queries actually used [6].
The code agent reads that summary and a log of every representation program tried so far. Then it proposes edits. Each edit gets scored on a set of validation questions. Only edits that clear a minimum improvement are kept. When several pass, the code agent tries combining them. The final program is scored once on separate held-out questions it never saw during the search [6].

Figure 5 - What each agent actually does: The analysis agent has three tools: search the current index, read a specific document, and search for a term across documents. It uses them to write a plain-language diagnosis of a batch of failing and succeeding queries. The code agent edits the representation program in response, and the edit is tested before it’s kept.
What the written diagnosis adds
The paper tests this with an ablation, an experiment that removes one piece of the system at a time to see what that piece contributes. AutoIndex ran under four conditions: the full two-agent loop, a single-iteration version, a version without the search history, and a version without the analysis agent. In that last one, the code agent had nothing but the raw metric delta, whether recall went up or down. The full loop improved 8 of 8 tasks. Cutting it to one iteration dropped that to 3 of 8. Without the search history, 5 of 8 improved. Without the analysis agent and its written diagnosis, 6 of 8 still improved.
That’s the smallest drop of the three. So the case for the diagnosis rests on how big the gains were. The paper reports that “removing the Analysis Agent shrinks effect magnitudes… indicating that grounded failure analysis is a substantial source of signal” [6]. Its explanation of why two separate agents exist at all makes the same point from the other side (nDCG@10, in the quote, is a ranking score defined later). In the paper’s words, “a single coding agent given only aggregate Recall@100/nDCG@10 deltas had to guess what was going wrong in the corpus, producing low-signal hypotheses biased toward generic preprocessing tricks” [6].

Figure 6 - What each part of the loop contributes: Four versions of AutoIndex, scored by how many of 8 benchmark tasks each one improved. The full loop improved all 8. Cutting it to one pass helped only 3. Without the analysis agent’s written diagnosis, 6 still improved, but the paper reports the gains were much smaller.
A different research team reached a similar conclusion a year earlier, on a different problem: tuning the instructions given to AI systems rather than how documents get indexed. Their system, GEPA (a prompt optimizer), was accepted at ICLR (the International Conference on Learning Representations) 2026, a leading machine-learning research conference, and the AutoIndex paper cites it as related work [6][7]. GEPA improves prompts, the written instructions given to a model, by having a model reflect in plain language on why a specific attempt failed. The usual alternative is reinforcement learning. It trains a model by rewarding it with a score. GEPA’s own stated reason: “the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards” [7]. In both systems, a sentence explaining a failure gave the optimizer more to work with than a score alone.
A Google Research team hit the mirror image of this problem. They were training a model to break a complex search question into simpler sub-searches, a costly step, so it could run cheaply when a user asks. That’s a different technique with a different goal. The model was trained against a single narrow score, and it learned to game the score instead of solving the task: “without a diversity term, the model quickly collapses into generating degenerate, nonsensical strings… to mathematically exploit the vector coordinates of the database” [8]. It says nothing about whether AutoIndex’s method works. It comes from a different team on a different problem. What it shows is the failure a narrow single-number signal invites.
Why the loop runs on an old-fashioned retriever
One design choice surprises most readers. The paper runs almost all of its experiments against BM25, a lexical scoring method from the 1990s [9], rather than a modern dense embedding retriever. Most teams treat lexical search as something you graduate from once you can afford embeddings.
The paper’s stated reason is legibility: “We use BM25 as the primary retriever because it provides a strong, widely used, and transparent testbed for studying document representation. Its sensitivity to segmentation, normalization, and term reweighting makes it well suited for isolating the effect of learned representation programs” [6]. When a BM25 search fails, an agent can trace the failure to an actual missing or diluted word and write a sentence a person could check against the source text.
Sam O’Nuallain, the first-listed of the paper’s four equal-contribution authors, went further on the Weaviate Podcast. Dense retrievers, he said, are “a little blackboxy,” and he suspected it’s harder for an agent to get useful feedback from them about why one document ranked above another [10]. That stronger claim, that legibility helps an agent act on its own diagnosis, is O’Nuallain’s own extension in the interview. The paper states the narrower version. BM25 is a transparent testbed for isolating what a representation change actually does.

Figure 7 - Why a plain-word retriever is easier to debug: A BM25 miss traces to specific words an agent can point at and check against the source. A dense retriever’s ranking comes from a neural network’s learned geometry, with no comparably legible explanation for why one document scored closer than another.
KEY INSIGHT: If you want an automated system to explain its own failures, pick a retriever whose misses you can trace to specific words, even if another retriever scores a little higher on accuracy.
What the loop found
AutoIndex was tested on CRUMB (Complex Retrieval Unified Multi-task Benchmark), an existing set of 8 quite different retrieval tasks from another research group, built to stress multi-part queries and long documents [6]. Against a plain BM25 baseline, the learned representation programs improved recall on all 8 of 8 tasks. The average gain was 8.4% in Recall@100 and 8.3% in nDCG@10, a ranking-quality measure that rewards putting the right document near the very top rather than somewhere in the list [6]. The biggest single gain came on a task called SetOpEntity, one of the tasks where the paper says plain BM25 suffered most from queries and documents using different words: 30.5% in Recall@100 and 43.6% in nDCG@10 [6]. Those numbers use one of the two code-writing models the paper tested. With the other, Claude Sonnet 4.6, 7 of the 8 tasks improved [6].

Figure 8 - Eight tasks, eight improvements: AutoIndex’s Recall@100 gain over a plain BM25 baseline on each of CRUMB’s 8 tasks. Every task moved up, though PaperRetrieval’s +0.1% is effectively flat. The largest gain, on the SetOpEntity task, was 30.5%. The dashed line is the paper’s +8.4% gain in the average Recall@100 across tasks, which is not the same as averaging the eight percentages.
Two case studies show the code agent working out fixes nobody told it to look for. On the StackExchange task, built from questions on the Stack Exchange Q&A sites, the analysis agent noticed that documents heavy in LaTeX (mathematical notation markup) were being pushed down the rankings. The repeated markup was diluting the content words BM25 needed to match on. The code agent’s fix was narrow. It stripped LaTeX only from documents over a heavy-markup threshold and left prose-only documents alone [6]. It chose narrow on purpose. The search history showed that an earlier, broader cleanup had removed too much and caused regressions [6].
The second case is a “tip of the tongue” task, where users describe a movie or show they can’t quite name and the correct answer is a plain plot summary. Here the code agent learned something else. It repeated a document’s plot and cast sections several times within the same chunk, which raises how heavily BM25 weights those terms without throwing away the rest of the article’s context [6].
The paper reports that both fixes came from the analysis agent reading the actual failing documents, not from a rule written in advance [6]. A person debugging a search index by hand works the same way.
Where this stops working
The paper is explicit about its own limits. AutoIndex’s reported results optimize mainly for one metric, Recall@100, against one fixed retriever, BM25, with a limited number of search iterations and a small number of repeat runs [6]. The paper does report one dense-retrieval result. A representation program already learned against BM25 was applied to a dense retriever called Qwen3-Embedding-0.6B on the StackExchange task. Held-out Recall@100 rose from 0.7391 to 0.8741, an 18.3% relative gain [6]. That’s a real, positive result. It’s also narrower than it sounds. It reused an already-learned program on a new retriever. The paper did not run the full optimization loop against a dense retriever from scratch. The paper says broader dense, hybrid (keyword plus embedding), and reranking evaluations “remain future work” [6].

Figure 9 - A transfer, not a new training run: Reusing a representation program the loop had already learned against BM25, applied instead to a dense embedding retriever on the StackExchange task, raised Recall@100 from 0.7391 to 0.8741. The paper is explicit this is a transfer test, not an AutoIndex loop run from scratch on a dense retriever.
O’Nuallain floated a further idea in that same interview. Could the same analysis-and-edit loop learn to predict a database’s table layout (its schema) from unstructured source documents, for a text-to-SQL system that turns a plain-English question into a database query [10]? That idea appears nowhere in the paper. It’s O’Nuallain’s own speculation about where the framework might go next, not something the paper built or tested.
What this changes about the next retrieval complaint
Most teams don’t need a two-agent optimization loop to use this finding. What changes is where you look first. The next time a search feature misses an obvious document, pull up the failing query and the document that should have matched it before you reach for a new embedding model or a reranker. Write down, in a full sentence, what went wrong. That’s the move AutoIndex’s analysis agent makes with tools. A person can do a rough version by hand on a handful of failures. Foundry, Reducto, and LangChain all show that even a one-time, hand-written diagnosis, turned into a specific change to enrichment or chunking code, closes real gaps a better retriever alone never would. Once retrieval works, keeping what it finds current and attributable is its own problem, and The Production RAG Stack Nobody Ships Complete covers it [11].
The same habit tells you whether you have a representation problem at all. Some retrieval complaints come from querying structured data, the rows and columns in a database, through natural language, which is an architecturally different problem from finding the right passage of unstructured text. If your diagnosis keeps landing on “the answer is really a row in a table, not a paragraph in a document,” that’s your signal to get a retrieval-architecture review, not another embedding-model swap. It’s also the exact territory our txtToSql bolt-on engine was built for: giving a team’s actual structured data a natural-language front door, instead of pretending the problem is a chunking strategy.
Conclusion
The default answer to a retrieval complaint is the retriever. A newer embedding model, a reranker, a bigger index. Four vendors and one research paper point somewhere fewer people look: the code that decided how each document was represented before any matching started. AutoIndex’s controlled comparison shows why a written diagnosis moves that code further than a bare metric does. Its choice of a decades-old lexical retriever shows something else. The easiest system to debug and the most accurate system aren’t always the same one. The next time search can’t find the obvious answer, check your own representation code first.

Figure 10 - What’s proven, what’s speculation, what to do next: AutoIndex’s results are demonstrated against a fixed BM25 retriever on one primary metric, with one narrow, positive dense-retrieval transfer test. The text-to-SQL extension is the researcher’s own speculation, not a built system. The concrete next step for a reader: write a plain-language diagnosis of your own next retrieval failure before reaching for a new model.
References
[1] V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. Yih, “Dense Passage Retrieval for Open-Domain Question Answering,” Proceedings of EMNLP 2020, Nov. 2020. https://arxiv.org/abs/2004.04906
[2] Weaviate Blog, “Building Foundry Part 3: From archive to creative search,” Weaviate, Sep. 8, 2026. https://weaviate.io/blog/building-foundry-creative-search
[3] A. Abraham, “From Ingestion to Agents: How AI Teams Build on Document Intelligence,” AI Engineer conference, Sep. 23, 2026. https://www.youtube.com/watch?v=0I07YAuF8xA
[4] L. Bush, “Why We Rebuilt LangChain’s Chatbot and What We Learned,” LangChain Blog, Nov. 5, 2025. https://www.langchain.com/blog/rebuilding-chat-langchain
[5] Anthropic, “Introducing Contextual Retrieval,” Sep. 19, 2024. https://www.anthropic.com/news/contextual-retrieval
[6] S. O’Nuallain, N. Rajkumar, R. Narayanasamy, H. Jiang, S. Chaudhari, and A. Drozdov, “AutoIndex: Learning Representation Programs for Retrieval,” arXiv:2607.18603, Jul. 21, 2026. https://arxiv.org/abs/2607.18603
[7] L. A. Agrawal, S. Tan, D. Soylu, et al., “GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning,” arXiv:2507.19457, accepted ICLR 2026 (Oral), Jul. 25, 2025 (rev. Feb. 14, 2026). https://arxiv.org/abs/2507.19457
[8] P. Jiang and J. Y. Li, “Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train,” Google Research Blog, Sep. 15, 2026. https://research.google/blog/bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train/
[9] S. Robertson and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends in Information Retrieval, Vol. 3, No. 4, pp. 333-389, 2009. https://doi.org/10.1561/1500000019
[10] “AutoIndex with Sam O’Nuallain: Weaviate Podcast #143,” Weaviate, Sep. 7, 2026. https://www.youtube.com/watch?v=mAj92SoEhjc
[11] G. Dotzlaw, “The Production RAG Stack Nobody Ships Complete: Freshness, Evidence, Versioning, and Where It Stops Working,” Dotzlaw Consulting, Aug. 26, 2026. https://dotzlaw.com/insights/ai-27-production-rag-stack/
Building production AI, or modernizing a legacy system?
That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.