3547 words
18 minutes
There Is No Right Chunk Size: Multi-Scale Indexing for RAG

Picture a team building document search on top of RAG (retrieval-augmented generation: find the relevant text first, then hand it to a language model so it can answer a question). At index time (when documents are processed and stored so they can be searched), they split every document into chunks: small pieces of text an embedding model can actually look at. An embedding model turns text into a list of numbers so a computer can measure how close two pieces of text are in meaning. The team uses a common default, 512 tokens (a token is a word or word-piece, roughly three-quarters of a word), with a little overlap so nothing gets cut off mid-thought. It ships.

A user asks a narrow, specific question and gets the right document back immediately. Another user asks a question whose answer is spread across a longer passage, and the system comes back with nothing useful. The answer is sitting right there in the same corpus (the full collection of documents). The team tries a newer embedding model. The number barely moves. They add a reranker (a second model that re-sorts the first model’s top candidates). That doesn’t help either.

Nobody goes back to the chunk size. It feels like a decision you make once, at the start. It doesn’t feel like something that could be right for one question and wrong for the next.

Figure 1 - Illustrative bar chart with a cluster of six bars, one per chunk size, for each of three datasets, QMSum, NarrativeQA and Seinfeld, with an oracle marker sitting above every cluster and a labeled 20 to 30 percent recall gap

Figure 1 - The chart nobody looks at after launch: For each of three datasets, QMSum, NarrativeQA, and Seinfeld, six bars show recall@1 (how often the right document comes back as the first result) for six fixed chunk sizes. No single size wins across all three clusters. An oracle marker, a hypothetical system that always picked the best size for each question, sits above every cluster. It illustrates gaps AI21 Labs (an AI research company) measured at 20-30% in recall@1, reaching 40%+ on some datasets [1].


Two questions the same chunk size can’t both answer#

AI21 built and published a small trivia dataset out of Seinfeld episode transcripts to test this directly [1][2]. Two example queries from their published code show the problem. The first: “What’s the name for Jerry’s favorite shirt?” That’s a focused, verbatim fact sitting in one sentence somewhere in the transcripts. Small chunks, 100 and 200 tokens, retrieve the right document at rank one. Chunks of 500 tokens fail to find it anywhere in the top 20 results [2].

The second query: “What is Kramer’s first name?” That one needs more surrounding context before the search finds it reliably. Small 100-token chunks rank the correct document fifth. Medium, 200-token chunks improve that to third. Larger, 500-token chunks reach second [2].

Run both queries against one fixed chunk size and one wins while the other loses. Pick 100 tokens and the shirt question lands at rank 1 while Kramer slips to rank 5. Pick 500 and Kramer climbs to rank 2 while the shirt drops out of the top 20 entirely. The middle size, 200, does reasonably on both, but it isn’t the best choice for Kramer. Across a bigger set of questions, compromises like that add up. That’s where the oracle gap hides: the distance between any one fixed size and the best size for each question.

Figure 2 - Diagram of two Seinfeld trivia queries, one about Jerry's favorite shirt answered correctly by small 100 to 200 token chunks and one about Kramer's first name that improves as chunk size grows from 100 to 500 tokens, showing opposite winners

Figure 2 - The same corpus, two opposite winners: “Jerry’s favorite shirt” is answered correctly by small chunks and missed entirely by large ones. “Kramer’s first name” does the opposite, improving steadily as the chunk grows from 100 to 500 tokens. No single size is best for both questions.

Fix one: pick the default and move on#

The usual approach, in AI21’s words, is to find a “sweet spot” chunk size, typically 500-800 tokens, with a small overlap between neighbors so nothing important gets cut at a boundary [1]. It’s a reasonable place to start. It gets a working index shipped in an afternoon instead of a design meeting, and for a lot of corpora it’s good enough to launch with.

The trouble is that the number was arbitrary from the start. Nobody measured whether 512 tokens is right for this corpus, let alone for any single question a user might ask. The team picked it because everyone else uses it. Changing it later means admitting the first choice was a guess. Once the system is live, re-chunking also means re-indexing everything, so the number tends to stay.

Figure 3 - Diagram of a document being split into fixed 512 token chunks with overlap and indexed, labeled as an arbitrary choice never revisited after launch

Figure 3 - The default nobody re-checks: A document gets split into fixed 512-token windows and indexed exactly once. The size was picked because it’s the common default, not because anyone measured it against this corpus or these questions.

Fix two: tune the number on an eval set#

Once the first complaint comes in, most teams start measuring. They build a small eval set: a list of questions, each paired with the document that should answer it. They score every chunk size they tried against it and keep whichever size scores best on average. That’s a real improvement over a guess. It replaces “512 because everyone uses 512” with a number backed by a test.

The underlying problem is still there. AI21 ran this kind of test at a larger scale. They indexed the same corpus at six chunk sizes, 50, 100, 200, 500, 1,000, and 2,000 tokens, across three benchmark datasets: QMSum (meeting-transcript summarization), NarrativeQA (questions over full-length stories), and their own Seinfeld set. For each size on each query, they measured recall@1 (whether the correct document lands as the very first search result) [1]. Plotted together, the six lines cross each other again and again. No single size wins.

Then they built what they call an oracle aggregator. It’s a hypothetical system that, for every query, picks whichever chunk size actually worked best for that question, as if it already had the answer key. The oracle line sits well above every fixed size. In AI21’s words: “The oracle substantially outperforms all fixed choices. Gaps of 20-30% in recall@1 are common, reaching 40%+ in some datasets” [1].

That’s the ceiling on fix two. An eval-tuned chunk size gets the best average score across the eval set. It tells you nothing about the individual questions that still lose under that “best” size. The oracle gap measures exactly that blind spot.

Figure 4 - Diagram of an eval-set-tuned chunk size selection process, showing an average recall score chosen as best, next to the same six crossing lines from the oracle chart with an arrow noting that some queries still lose under the chosen size

Figure 4 - Tuning optimizes the average, not the question: An eval set picks the chunk size with the best average score. The same crossing-lines pattern from the oracle experiment still applies underneath that average: some individual queries lose no matter which single size gets chosen.

KEY INSIGHT: An eval-tuned chunk size is only as good as the average it was tuned against. Before you trust one number, look at the per-query results underneath it as well as the mean.

Why tuning keeps missing the same questions#

Neither the default nor the tuned number can close this gap, and the reason is timing. Yuval Belfer is a senior developer advocate at AI21 and co-author of the blog post [1]. He named the problem in a later conference talk on the same research: “If I’m looking at the indexing part where I do have control over the chunk size, I don’t know what the queries will be… And the retrieval part where I do have my queries, I cannot control the chunk size, right? It’s already fixed” [3].

Chunk size gets decided once, at indexing time, before a single user question exists. The query shows up later, at retrieval time, when the chunk boundaries are already baked into the index. The stage that knows the question can’t change how the document was cut up. The stage that can change it has no idea what’s coming. So tuning the number at indexing time can’t close the gap. The stage making the choice simply doesn’t have the information it needs.

Figure 5 - Two-stage diagram showing an indexing stage that controls chunk size but does not know the future queries, separated by a gap from a retrieval stage that knows the query but cannot change the chunk size already fixed

Figure 5 - The information trap: Indexing time controls chunk size but doesn’t know what will be asked. Retrieval time knows the question but the chunk size is already locked in. Neither stage alone has both the information and the power to fix it.

KEY INSIGHT: Any fix has to move the size decision to query time. It can predict the right size for each question, or it can index every size ahead of time.

Fix three: a better embedding model or a reranker#

Before most teams spot the timing problem, they try two other fixes: a better embedding model and a reranker. A dense retriever searches by comparing the query’s embedding to every document’s embedding and returning the closest matches. A better embedding model can make closeness in vector space (the numbers an embedding model produces) line up better with actual relevance, and it can nudge recall upward. A reranker takes the retriever’s top candidates and re-sorts them before anything reaches the user.

Neither one touches the chunk boundary. A reranker can only reorder documents that already made the candidate list. If the right document never got retrieved, because it lived in a chunk of the wrong size for that question, the reranker never sees it. Our companion piece on AutoIndex covers this gap in more depth, including a research system that rewrites the code deciding how documents get represented before indexing [4]. We won’t repeat that heavier fix here. For the chunk-size problem, a better matching function still picks from candidates drawn from one fixed-size index. It can’t find a document that was never chunked at a size that lets it be found.

Fix four: let a router choose per query#

If the right size depends on the question, the direct response is to build something that picks the size per question. Academic work outside AI21 supports this. Mix-of-Granularity trains a small router that scores every granularity’s (chunk size’s) available chunks for an incoming query, weights those scores, and returns the chunk from whichever granularity scored highest [5]. The work is peer-reviewed, and it confirms the same claim AI21’s oracle experiment measured: the best chunk size varies by query, as well as by corpus or document type.

It also brings a new problem. A router is a model. It has to be trained on labeled data, evaluated, and maintained, and its weights decide which granularity’s chunk reaches the language model. If it learns the wrong preference for one kind of question, every question of that kind inherits the error.

Figure 6 - Diagram of a query flowing through a trained router that weights every chunk size's score and returns only the chunk from the highest-weighted size

Figure 6 - A trained router still has to be trained: Mix-of-Granularity scores every chunk size for an incoming query, weights those scores, and returns the chunk from the highest-weighted size. The router itself is a model that has to be trained, evaluated, and maintained, and a wrong learned preference for a class of question propagates to every question in that class.

The fix that stops choosing: multi-scale indexing#

AI21’s own answer skips the guess entirely. Pick several chunk sizes instead of one, and stop trying to decide which is right before the question arrives. The method has three parts: index the same corpus at several sizes, query all of them at once, then merge the results with reciprocal rank fusion (RRF, a formula for merging several ranked lists into one) [1][2].

  1. Index the same corpus at several chunk sizes. Build separate indices, each chunked at its own fixed size (AI21’s own examples use sizes drawn from 50 up to 2,000 tokens [1]). No new chunking algorithm is needed. It’s the same fixed-size chunking from fix one, built several times.
  2. Query every index in parallel. Each incoming query runs against all of the indices at once, producing one ranked list per chunk size instead of a single list.
  3. Map every retrieved chunk back to its parent document, then fuse the rankings with RRF. The formula scores each candidate as the sum of 1/(rank + k) across every list it appears in. Here k is a constant that controls how steeply rank position is weighted. RRF uses only each list’s rank position, not its raw similarity score. AI21’s demo code uses k=1 [2], while the original RRF paper uses 60 [6]. A small k rewards a document for one very high rank. A large k rewards it for showing up often. Check which one an implementation uses before comparing its results with AI21’s.

Figure 7 - Architecture diagram of a corpus indexed at several chunk sizes in parallel, a query hitting all indices at once, chunks mapped back to parent documents, and the results fused by reciprocal rank fusion into a single ranked list

Figure 7 - Index at several sizes, decide at query time: The corpus is duplicated across several chunk-size indices. A query hits all of them in parallel. Retrieved chunks are mapped back to their parent documents, and reciprocal rank fusion merges the per-size rankings into one list.

Why voting across scales works when ranking chunks doesn’t#

Two separate choices make this work. First, ranks replace scores. A similarity score from a 200-token index and one from a 2,000-token index aren’t on the same scale, so RRF ignores scores and uses only each result’s position in its list [1]. Second, chunks become documents. Each retrieved chunk counts as a vote for the document it came from, so lists of different-sized fragments end up voting on the same set of candidates. Once every index ranks the same set of documents, the lists can be fused directly.

This is what fixes the timing problem from earlier. Retrieval time still can’t change the chunk size after the fact. It doesn’t need to, since indexing time already built every size it might need. The decision that used to be locked in months before any query existed now gets made per question. A document that several scales, or several chunks within one scale, keep ranking high rises to the top.

KEY INSIGHT: When you merge rankings from different chunk sizes, merge on rank position instead of raw score, and count each chunk as a vote for its parent document. Those two steps are what make lists from different scales comparable.

What multi-scale indexing costs#

Building several indices instead of one multiplies storage. AI21 reports that the number of chunks to embed and store grows by a factor of 2x to 5x across their experiments, depending on how many chunk sizes get indexed [1]. Every query also becomes several retrieval calls, one per chunk size. AI21 states these run in parallel [1][3], and we found no measured latency figure published beyond that. Nobody has published how much slower a live query gets, if at all. The honest version of the claim is narrower than “free”: there are no extra model calls and no new vendor, but there is real added storage and query volume.

Each extra index is also an extra bulk load. In our own Qdrant (an open-source vector database) measurements, search ran hundreds of times slower at p95 (the latency that 95% of searches beat) for several minutes after a bulk load. The median looked normal the entire time [7]. Budget the rebuild window for every scale, along with the disk.

The gains are also smaller than the oracle gap makes them sound, and the two numbers measure different things. The oracle line is an upper bound nothing can fully reach, since it assumes it knows the correct answer before searching. The method you can actually deploy, multi-scale plus RRF, gained “1-37% across benchmarks” [1].

AI21 ran it on MTEB (Massive Text Embedding Benchmark, a standard suite of tasks used to score embedding and retrieval quality), using 2 embedding models across 8 configurations. The method beat the single-size baseline in 7 of 8. The gains were mostly a modest 1-3%, and one configuration didn’t improve at all. The standout was TRECCOVID (a COVID-19 scientific-literature search benchmark), where E5-small (the smaller of the two embedding models tested) gained 36.7% [1]. On the other four benchmarks AI21 gives no numbers in its text, only that the method matched or beat the best single size and often came close to the oracle [1]. So expect a small, steady improvement. The occasional large win shows up on a dataset where a particular scale mismatch happened to be severe.

Fusion doesn’t always match the best single size, either. On AI21’s shirt example, RRF returned the correct document at rank 2, one spot behind the rank 1 that the 100- and 200-token indices reached alone [2].

Figure 8 - Stat-tile chart of AI21's reported numbers: storage growing 2x to 5x, 7 of 8 MTEB configurations improved, a typical 1 to 3 percent gain, and one 36.7 percent outlier on TRECCOVID

Figure 8 - The price and the payoff: Storage grows by 2x to 5x. Every query becomes several parallel retrieval calls instead of one, with no extra model calls. On AI21’s MTEB runs, 7 of 8 configurations improved, typically by a modest 1-3%, with one dataset (TRECCOVID) reaching a 36.7% outlier.

KEY INSIGHT: Multi-scale indexing trades storage for retrieval quality. Measure the gap on your own data before you pay for it, and don’t expect to get the oracle number back.

Measuring the gap on your own corpus#

AI21’s public repository ships the retrieval demo, but not the measurement code behind the oracle chart [2], and we found no ready-made script elsewhere. The measurement can be rebuilt from the same method in four steps:

  1. Build a labelled query set. Write a list of realistic questions, each paired with the document that should answer it.
  2. Index the corpus at a few different fixed chunk sizes. 3 or 4 sizes spanning a meaningful range are enough to see whether the problem exists at all.
  3. For every query, compute recall@K at each chunk size. Check whether the correct document appears in the top-K results for each size on its own.
  4. Compare the per-query best against the single best fixed size. For each query, take whichever size worked best, and average those scores across the query set. Average the best single fixed size the same way. The difference between the two is the oracle gap for that corpus.

If the gap is small, a single well-tuned chunk size is probably good enough, and the extra storage for multi-scale indexing isn’t worth paying for. If it’s large, the gap itself is the evidence needed to justify the storage cost before spending it.

Figure 9 - Four step flow diagram showing building a labeled query set, indexing at several chunk sizes, computing per query recall at each size, and comparing the per query best against the single best fixed size to get the oracle gap

Figure 9 - Measuring the gap before paying for the fix: Build a labelled query set, index at a few sizes, compute per-query recall at each size, then compare the per-query best against the single best fixed size. The difference is the oracle gap for that corpus.

Conclusion#

The right chunk size depends on the question. The two stages that could fix it, indexing and retrieval, never have the right information at the same time. Of the fixes we walked through, multi-scale indexing is the only one that stops guessing altogether, at a real and measurable price in storage.

Start by measuring. Run the four-step oracle-gap check against a real corpus and a real query set. It shows whether this problem is even present, and how large it is, before any storage gets spent closing it. That’s the shape of a retrieval-quality audit: measure the oracle gap on the documents and questions that matter to a given system, then decide, with a number in hand, whether multi-scale indexing earns its keep.

References#

[1] N. Granot and Y. Belfer, “Chunk size is query-dependent: a simple multi-scale approach to RAG retrieval,” AI21 Blog, Jan. 29, 2026. https://www.ai21.com/blog/query-dependent-chunking/

[2] AI21 Labs, “multi-window-chunk-size,” GitHub repository. https://github.com/AI21Labs/multi-window-chunk-size

[3] Y. Belfer, “Stop Chunking Like It’s 2022,” AI Engineer World’s Fair 2026, AI Engineer (YouTube), Sep. 16, 2026. https://www.youtube.com/watch?v=r9OwPx_HoV0

[4] G. Dotzlaw, “Your Retriever Isn’t the Problem, Your Representation Is: What Happens When an Agent Diagnoses Retrieval Failures in Plain English,” Dotzlaw Consulting, Oct. 2, 2026. /insights/ai-57-autoindex-code-optimization-retrieval/

[5] Z. Zhong, H. Liu, X. Cui, X. Zhang, and Z. Qin, “Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation,” COLING 2025, arXiv:2406.00456. https://arxiv.org/abs/2406.00456

[6] G. V. Cormack, C. L. A. Clarke, and S. Büttcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods,” SIGIR 2009. https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf

[7] G. Dotzlaw, “Qdrant in Production: The Slow Window After a Bulk Load, and Why Your Median Hides It,” Dotzlaw Consulting, Oct. 1, 2026. /insights/ai-51-qdrant-optimizer-numbers/

There Is No Right Chunk Size: Multi-Scale Indexing for RAG
https://dotzlaw.com/insights/ai-59-multi-scale-indexing-chunking-fix/
Author
Gary Dotzlaw
Published at
2026-10-08
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

← Back to Insights