4427 words
22 minutes
Qdrant in Production: The Slow Window After a Bulk Load, and Why Your Median Hides It

We loaded 1,759,180 embedding vectors (lists of numbers that capture the meaning of a piece of text) into a fresh collection in Qdrant, an open-source vector database, using default settings. The upload finished in 162 to 165 seconds, and every batch came back acknowledged. We started searching the moment the last batch returned. Once this database settles, a search takes about 1.9 ms. Right after the load, a typical search took 600 to 900 ms, and it stayed that way for the first 3.5 to 4 minutes. At the 95th percentile (the latency that 95% of searches beat, which describes the slow tail), search in that window was 252x and 325x slower than it would be a few minutes later, across our 2 runs.

Then we looked at the median, the latency that half of all searches beat. The plain median of every search we ran during that window came out at 2.14 ms and 2.03 ms. That’s about 1.1x the settled speed. The number most latency dashboards lead with said nothing had happened.

Figure 1 - Timeline chart of search latency after a Qdrant bulk load: about 600 to 900 ms for several minutes, then a drop to about 2 ms once the optimizer goes idle

Figure 1 - The load finished, the database wasn’t ready: Our run on 1,759,180 vectors. The upload returned, then search sat at 600 to 900 ms for minutes before falling to about 2 ms. The plain median of every search in that window was 2.14 ms, which is why a median-only dashboard misses it.


The load finished, but the database was still working#

A vector database stores embeddings and finds the ones closest to a query. Checking every vector one by one is slow, so Qdrant builds an index called HNSW (Hierarchical Navigable Small World, a graph that lets a search hop toward the nearest neighbors instead of checking everything). Qdrant also splits a collection into segments (smaller chunks of the data, each with its own index).

An upload drops points into segments fast. The heavy work comes later. Background jobs called optimizers build the HNSW graph for each segment (indexing), combine small segments into bigger ones (merging), and clean out deleted points (vacuum). The upload call returns long before they finish. Qdrant’s study calls the time between “upload returned” and “optimizers idle” the draining window, and the time after it the steady state [1]. We use the same names.

During the draining window, a search fights the optimizers for the same CPU and disk. Some segments aren’t indexed yet, so the search scans them slowly. That’s why ours took hundreds of milliseconds instead of 2.

Qdrant measured this first. Clelia Bertelli’s Qdrant blog article loaded about 1.76 million 1,024-dimension Cohere embeddings of MS MARCO passages (a public set of web-search text from Microsoft) onto “a single Qdrant node running Ubuntu 26.04 x86_64, with 32GB of RAM and 14 Intel CPU cores” [1]. With default continuous indexing, “median latency was 780 ms, p95 reached 2.0 s” during draining, and once the optimizers went idle, “median search latency dropped 180x, from 780 ms to 4.3 ms” [1]. The window lasted about 11 minutes [1].

Short names: the median is p50, and the 95th percentile is p95.

Picture who sees this. A nightly job refreshes your product catalog. It finishes at 2:00 a.m. and reports success. A user in another time zone searches at 2:01 and waits almost a second for results that normally take 2 ms. The load job is green, the error rate is zero, and the latency panel shows a median that barely moved.

We reproduced the effect on our hardware with the same floating qdrant/qdrant:v1.19 image tag the study pins [2]. Pulled in September 2026, that tag gave us v1.19.1 [3]. The study ran in August, before that patch shipped, so we were one patch release ahead of Qdrant. Our container had 14 CPUs and 24 gigabytes (GB) of RAM, and we used the study’s own data file: the full 1,759,180 points, plus 1,000 held-out queries. We wrote our own harness against Qdrant’s REST API (its standard HTTP interface).

Figure 2 - Bar chart of median search latency in 30-second bins across one draining window: about 600 ms for 4 minutes, falling to 208 ms, then 2.0 ms in the final bin

Figure 2 - What the draining window looks like, 30 seconds at a time: Our first full-scale run with defaults. Every 30-second bin for the first 4 minutes had a median between 593.9 and 625.0 ms. Latency then fell to 414.6, 286.8 and 208.2 ms, and only the final bin dropped to 2.0 ms.

The obvious fix: wait until it’s done#

The obvious move is to wait. Load the data, pause, then send traffic. That works if you wait long enough. The hard part is knowing how long.

Our 2 full-scale windows were 321.3 and 349.8 seconds. Qdrant’s published run took about 11 minutes [1]. Qdrant’s own checked-in results [2] hold 4 runs that loaded the same 1.76 million points with the same effective default settings. Those 4 windows ranged from 474 to 1,027 seconds. Upload speed matters too. When we sent a 500,000-point set over 1 connection, indexing kept pace with the upload, and the window was only 24.5 seconds with a 2.4x p95 penalty. Over 8 parallel connections, the same data built a real backlog, and the p95 penalty jumped to 79x to 90x.

So a fixed sleep is either too long or too short. It’s better to check whether the database is done and watch latency. If you watch the median, though, it already says everything is fine.

Tag every search with the phase it ran in#

Qdrant’s study answers the “how do we know it’s done” part [1]. The harness polls the /collections/{name}/optimizations endpoint, which reports queued and running optimizations and has been available since v1.17.0 [4]. It searches nonstop while that endpoint shows work and tags every sample as draining or steady. Once the optimizer reports idle, it runs 5 passes over 1,000 queries as the steady baseline [1].

The tags give you two groups to compare instead of one latency number for the whole day. Qdrant published the harness, scripts and results on GitHub [1], [2], which let us check their numbers. The repository has no license file [2], so we rebuilt the method ourselves instead of reusing the code. That takes about an afternoon.

Figure 3 - Diagram of the three benchmark stages: upload with no searches, draining with searches tagged draining while the optimizer endpoint is polled, and steady with 5,000 searches

Figure 3 - Phase tagging, the method Qdrant published: Upload with no search traffic, then search continuously while polling the optimizations endpoint, tagging each sample. After the optimizer confirms idle, 5 passes over 1,000 queries form the steady baseline.

We did the same, polling every second and calling the window closed after 3 idle polls in a row. The tags split our data cleanly.

Then we took the median of the draining group, and it said 2.14 ms. The tags were right. The median misled us.

Why the median of the tagged samples still lies#

The benchmark is closed-loop, which means it keeps exactly one query in flight and sends the next one only when the last one returns. Qdrant’s article flags this for sample counts: “When a query takes 800 ms, only about 1.25 queries fit into a second of wall-clock time” [1]. So a slow minute produces about 75 samples. A fast minute at 2 ms produces thousands.

In our runs, 71% and 78% of all draining samples came from the final 30 seconds of the window, after the last big segment finished indexing. Those fast samples outvote 4 minutes of slow ones, so the plain median lands on the fast side.

There are two honest ways to read the window. One is the p95, which our runs put at 633.85 and 811.93 ms against a steady p95 of 2.52 and 2.50 ms. The other is a time-weighted median (each sample weighted by how long it took, which gives the latency that a user arriving at a random moment would see). Ours came out at 588.89 and 668.49 ms. A chart of 30-second bins, like Figure 2, shows the same thing at a glance.

Figure 4 - Comparison of three statistics for the same draining window: plain median 2.14 ms, time-weighted median 588.89 ms, p95 633.85 ms, with 71% of samples from the last 30 seconds

Figure 4 - Same samples, three answers: Our first full-scale run. The plain median of the draining samples says 2.14 ms because 71% of them came from the last 30 seconds. The time-weighted median (588.89 ms) and the p95 (633.85 ms) describe what users actually waited.

We ran the same analysis on Qdrant’s checked-in result files [2]. The headline numbers held up. We recomputed the 780 ms median as 779.66 ms and the 180x as 180.5x. The surprise was how much the runs differed. The 4 runs with the same effective default settings had draining sample medians of 779.66, 5.19, 44.00 and 4.49 ms. The published 780 ms is real. It’s the highest of the 4, the one run where the slow period happened to dominate the sample count. The time-weighted medians of those same 4 runs were 1,048, 1,785, 525 and 1,405 ms, between 129x and 439x their steady state.

Credit to Qdrant: they published every sample, and their article warned about closed-loop sample counts [1]. One median from one run hides how noisy this window is. The p95 and the time-weighted median move far less between runs.

Figure 5 - Chart of four identical Qdrant runs: sample medians 779.66, 5.19, 44.00 and 4.49 ms against time-weighted medians of 1,048, 1,785, 525 and 1,405 ms

Figure 5 - Four runs, one configuration: Our re-analysis of Qdrant’s checked-in data for the 4 runs that loaded 1.76 million points with default settings. The sample median swings from 4.49 to 779.66 ms. The time-weighted median stays between 525 and 1,785 ms in every run.

KEY INSIGHT: When you measure latency during any background catch-up period, don’t report the plain median of a closed-loop test. Report the p95 or a time-weighted median, and chart it in time bins, because slow periods produce few samples.

Seeing the window doesn’t make it shorter. Qdrant’s optimizer settings do that.

prevent_unoptimized: fast search, hidden writes#

The setting with the biggest effect is prevent_unoptimized. Qdrant’s docs describe it this way: “points written to an unindexed segment that is larger than indexing_threshold are accepted and durably stored but are not visible in search results” [4]. indexing_threshold is the segment size, in kilobytes (KB), above which Qdrant builds an HNSW index instead of scanning. Qdrant’s docs use 10,000 KB in their example configuration [4], and our default collections reported the same value. Those held-back points are called deferred points, and they “only become visible after the optimizer has indexed the segment” [4]. The docs also warn: “prevent_unoptimized is an experimental feature; its behavior may change slightly in future releases and it must be used with care” [4]. It has been available since v1.17.1 [4].

The speed gain is large. Qdrant reports that it “dropped draining-phase p50 latency from 780 ms to 10.2 ms, a 76x improvement, with p95 at 81.3 ms” [1]. In our full-scale runs, the draining p95 fell from 633.85 and 811.93 ms to 20.98 and 17.14 ms, which is 30x and 47x better. Our time-weighted median fell from 588.89 and 668.49 ms to 14.26 and 9.95 ms. The plain median made it look worse: 4.45 and 8.65 ms against the defaults’ 2.14 and 2.03 ms, because of the sampling effect above. Our re-analysis of Qdrant’s own 2 runs, using the time-weighted median, puts the gap at 28x rather than 76x.

Qdrant’s article names the price: “prevent_unoptimized trades write visibility for query latency” [1]. It says freshly written points can be “durable, but invisible to search” until their segment is optimized, and that for large collections you should “evaluate how long that delay gets before enabling it” [1]. We measured that delay, and it hit more than search.

When our upload returned, 1,136,680 and 1,318,180 of the 1,759,180 points were deferred. That’s 64.6% and 74.9% of the collection. We checked the 200 most recently written points every 5 seconds in 4 ways:

  1. Retrieve by id found 0 of 200.
  2. An exact count filtered to those ids returned 0.
  3. A vector search filtered to those ids returned 0.
  4. Searching 10 of them with their own vector found 0 of 10 in the top 10.

In one run, those points stayed hidden for the whole 275.6-second window. In the other, they appeared 244.3 seconds in. For about 4.5 minutes, a client can write a point, get an acknowledgment, and then fail to read it back by id. The newest points weren’t the only ones. In every run with the flag on, 4 to 6 of 10 older control points, written in the first half of the load, also failed to come back when searched with their own vector as draining began.

Nothing in the collection info warns you. points_count reported the full 1,759,180 the entire time. The field that shows the deferral is update_queue.deferred_points, inside the update_queue section Qdrant’s docs point to for monitoring [4]. The docs describe the effect on search results. They don’t say what happens to retrieve or count, which is what we measured on 1.19.1.

Figure 6 - Diagram of write visibility with prevent_unoptimized: points_count reports 1,759,180 while retrieve, count and search each find 0 of the 200 newest points

Figure 6 - Acknowledged, counted, and invisible: Our full-scale runs with prevent_unoptimized on. Up to 74.9% of points were deferred when the upload returned. The collection’s points_count stayed at 1,759,180 while retrieve, count and filtered search found none of the 200 newest points. Only update_queue.deferred_points showed the gap.

Search results shifted too. Hundreds of thousands of points stayed visible, so searches never came back short, but the answers changed. During draining, the top 10 results matched the final top 10 for the same query only 94.3% and 92.6% of the time at full scale, and 58.9% and 61.3% at 500,000 points. With the flag off, that overlap was at least 99.8%.

Watch out if you use the Python, TypeScript, .NET or Java clients, which default to wait=true (a write mode that returns only once the write is searchable) [4]. With this flag on, “the response is held until every deferred point, including the current update, has been indexed” [4]. We used wait=false in every run, as the docs recommend [4].

KEY INSIGHT: Before you turn on a setting that trades freshness for speed, test reads by id and counts as well as search, and alert on the field that exposes the backlog. On Qdrant 1.19.1, that field is update_queue.deferred_points, and points_count won’t tell you.

So prevent_unoptimized fixes slow search and breaks read-your-own-writes (the guarantee that a client can read back what it just wrote) for minutes. If your app can’t live with that, you need a different setting.

Smaller segments: a faster recovery, a slower forever#

Next is segment size. Fewer, bigger segments mean more merge work after a load, but faster search once it’s done. Qdrant tested this. A single segment gave the best steady search, 3.2 ms against 4.1 ms for the default, but took “just over one hour” to clear its backlog [1]. Capping segments at 100,000 KB (about 100 MB each) cleared the backlog in 283.9 seconds, “12.7x faster” than the single segment [1], with a steady median of 17.3 ms.

One detail from our re-analysis: the single-segment run stopped at the harness’s 3,600-second timeout with indexing still running [2]. It never finished, so 12.7x is a lower bound on the speedup.

We ran a single 500,000-point test with a smaller segment cap. The backlog cleared 1.6x faster than the mean of our 3 default runs at 500,000 points, and steady search was 2.6x slower (4.82 ms against 1.80 to 1.85 ms). That’s 1 run, so treat it as a direction rather than a measurement.

So smaller segments shorten a window you pay once per load, and you pay for that on every search afterward.

Figure 7 - Chart of the segment-size trade: a smaller segment cap clears the backlog faster but raises steady-state median latency, from Qdrant's 4.1 ms to 17.3 ms

Figure 7 - Recovery time against everyday speed: Qdrant’s results for default segments against a 100,000 KB cap, with the 12.7x speedup over a single segment marked as a lower bound. Our 1 run at 500,000 points showed the same shape, 1.6x faster recovery and 2.6x slower steady search.

Turning indexing off during the load#

Another common move is to switch indexing off while loading and turn it back on afterward. Qdrant measured the “off” half. With no index, search falls back to brute-force scans (checking every vector), and steady latency was “256.6 ms, about 60 times higher than the continuously indexed collection after optimization” [1]. Turning indexing back on without prevent_unoptimized was worse still. Search ran at a “median of 2.7 s, with the tail reaching 12.1 s” [1].

Our single 500,000-point run of the full off-then-on sequence agreed. The upload was the fastest of any setting, at 23.8 seconds. After we switched indexing back on, the backlog took 193.8 seconds to clear, 2.7x longer than our 500,000-point default runs (mean 71.9 s). Delaying the index moved all the work to one moment after the load.

Figure 8 - Timeline comparing continuous indexing with indexing turned off then on: the off-then-on run uploads faster but its backlog takes 2.7x longer to clear

Figure 8 - Deferring the index defers the cost: Our single 500,000-point run. With indexing off, the upload was fastest, but switching it back on created a 193.8-second backlog, 2.7x longer than continuous indexing at the same size. Qdrant’s never-indexed collection searched at 256.6 ms.

Fewer optimizer threads and a lazier vacuum#

The last two settings come from Qdrant’s study. We didn’t test them.

Capping the optimization and index-building threads at 1 smooths search but takes longer. Qdrant set max_optimization_threads and max_indexing_threads to 1, which “stretched the draining window to 3,244.1 seconds, 6.8 times longer,” while draining p95 “was capped at 373.7 ms, less than half of the default configuration’s 820.4 ms” [1].

Vacuum is the cleanup that physically removes deleted points, and deleted_threshold sets how much deleted data triggers it. After deleting about 25% of the collection, a 20% threshold triggered vacuum, and “p95 search latency rose from 5.0 ms in steady state to 22.7 ms” [1]. At 50%, vacuum didn’t run, and “p95 latency fell from 5.4 ms to 4.3 ms” [1]. You trade disk space for a quiet search path.

Figure 9 - Two small panels: one optimizer thread gives a 6.8x longer window with draining p95 373.7 ms against 820.4 ms, and a 20% vacuum threshold raises p95 from 5.0 to 22.7 ms

Figure 9 - Two settings that trade time for smoothness: Qdrant’s figures. Capping the threads at 1 cut draining p95 from 820.4 to 373.7 ms but made the window 6.8x longer. A 20% vacuum threshold raised p95 from 5.0 ms in steady state to 22.7 ms while vacuum ran, after a large delete. At 50%, p95 went from 5.4 to 4.3 ms.

The bulk-load settings trade the same three costs#

Each bulk-load setting moves cost between three places: how long recovery takes, how fast search is once everything settles, and how soon new writes can be read. None of them removes the cost.

  • Defaults: minutes of slow search, then the fastest steady state.
  • prevent_unoptimized: fast search during recovery, but minutes of hidden writes, including reads by id.
  • Smaller segments: a shorter recovery, but slower search forever.
  • Indexing off during load: a fast upload, then a longer backlog later.
  • Fewer optimizer threads: smoother search, but a much longer window.

The vacuum threshold is separate. It matters after large deletes, where it trades disk space for a quieter search path.

So decide which cost your users can bear, then check that you got it. For a catalog refreshed overnight, a few minutes of slow search at 2:00 a.m. may be fine. For a chat product that writes a user’s message and immediately searches for it, hidden writes are a bug. You can only choose once you’ve measured your own window with phase tags and a time-aware statistic. That measurement works on other vector databases too, since it tracks what users experience instead of what the load job reports.

Figure 10 - Triangle diagram of three costs, recovery time, steady-state speed and write visibility, with each Qdrant setting placed near the cost it pays

Figure 10 - One trade, five settings: Each bulk-load setting shifts cost among recovery time, steady-state search speed and write visibility. Measuring your own draining window is how you see which cost you picked.

KEY INSIGHT: Treat “the load returned” and “the database is ready” as two separate events, and gate traffic or alerts on the optimizer status, not on the load job finishing.

The same lesson on different machinery#

Other systems hide costs the same way, through different machinery. At the AI Engineer conference, Jacob Lauritzen of Legora (a legal AI platform) and Simon Eskildsen, CEO and co-founder of Turbopuffer, described Legora’s project search on Postgres [5]. Legora had split one table into about 4,000 partitions and bin-packed projects into them, so hot and cold projects landed in the same partitions [5]. Queries pulled whole oversized partitions into memory, thrashing the cache (repeatedly evicting and reloading data from memory). Lauritzen said what that did to the slowest 1% of requests, the p99 (99th percentile): “we went from like search and ingestion P99 of 100 milliseconds into 20 seconds” [5]. The cause was partition layout and cache behavior. Legora moved project search to Turbopuffer and showed a latency chart on stage, saying median latency improved by roughly an order of magnitude and p99 by more [5]. The chart had no published numbers, so treat it as a loose, self-reported result.

AWS found a third kind of cost, one you see only if you measure quality as well as size. Deepak Dalakoti, Rhys Lewis and Petar Avramovic benchmarked binary quantization (storing each vector’s numbers as single bits to shrink the index) on Amazon OpenSearch Serverless [6]. They scored quality with NDCG (normalized discounted cumulative gain), a standard measure of how well a search ranks the right results near the top. The smaller index was easy to see: “At 1024 dimensions, binary embeddings reduced index size by 13.4x (0.40 vs 5.34 GiB) with a 5.2 percent NDCG loss …” [6]. The quality cost depended on the embedding size: “At 256 dimensions the quality cost rose to 28.3 percent …” [6]. That’s AWS testing its own backend in one internal sweep, and the post says its workloads are “separate workloads tailored to each database rather than a head-to-head benchmark …” [6], so it isn’t a comparison across databases.

Each of the three costs sat behind a number that looked like success: the load returned, the partitions held up until Legora scaled, and the index shrank 13.4x. Each cost showed up only when someone measured what users get, in search latency or ranking quality.

Figure 11 - Three side-by-side cards: Qdrant optimizer backlog, Legora Postgres partitions with p99 from 100 ms to 20 s, and AWS quantization losing 5.2% NDCG at 1,024 dimensions and 28.3% at 256

Figure 11 - Three mechanisms, one lesson: A background optimizer backlog in Qdrant, cache thrashing from partition layout in Legora’s Postgres setup, and a quantization penalty that grows as AWS shrank the embedding size. Different causes, each visible only once someone measured what users actually get.

Where our numbers stop#

Our setup differs from Qdrant’s, and some of our results are thin. We ran on Windows 10 with Docker Desktop on WSL2 (Windows Subsystem for Linux), with 14 logical CPUs on a 10-core desktop i9, and the client ran on the same machine. We uploaded over REST JSON with no payloads, and the study used gRPC (a binary protocol) and stored a small payload of document fields with each vector [2]. A visibility probe ran during every run, defaults included, which adds a little load. We did 2 full-scale runs per setting, so treat our results as ranges, not point values. The segment-size and indexing-off tests were single runs at 500,000 points. Our windows were about half as long as Qdrant’s, and we didn’t dig into why.

Our test was closed-loop, like Qdrant’s. An open-loop test (one that sends queries on a fixed schedule whatever the response time) might show a different shape. We didn’t test wait=true writes with prevent_unoptimized.

Qdrant’s numbers are Qdrant benchmarking its own product. Our runs confirm the direction, and our re-analysis confirms most of their published figures, but neither is an independent audit. The flag names are version-fragile too. prevent_unoptimized is experimental [4], and the retrieve and count behavior we saw is for v1.19.1 only. Tuning advice for one Qdrant version won’t carry over to another database. The phase-tag measurement will.

Figure 12 - Checklist diagram of the measurement: poll optimizer status, tag each search draining or steady, report window length, draining p95 and time-weighted median against steady p95

Figure 12 - The measurement to run on your own system: Poll the optimizer status, tag every search, and report the window length, the draining p95 and the time-weighted median next to the steady p95. Test reads by id and counts too if you enable a setting that defers writes.

Conclusion#

On our hardware, the gap between a finished bulk load and a ready database was 5 to 6 minutes, and search in that gap ran hundreds of times slower at the tail. A plain median hid it in our data and made it look small in 3 of Qdrant’s 4 identical runs. Qdrant’s settings can shorten or smooth the window, but each one moves the cost somewhere else, and the most effective one hides acknowledged writes from reads by id for minutes.

Before your next bulk load, measure your own window. Poll the optimizer status, tag each search as draining or steady, and report the window length and the draining p95 next to the steady p95. Then pick the setting whose cost your users can live with. If you run regular performance audits, this is one line item in the loop we described in Agentic Performance Audits [7].


References#

[1] C. Bertelli, “Configure Qdrant’s Optimizer for Predictable Search Latency,” Qdrant, Aug. 25, 2026. https://qdrant.tech/articles/tuning-qdrant-optimizer/

[2] Qdrant, “optimizers-in-action,” GitHub repository, 2026. https://github.com/qdrant-labs/optimizers-in-action

[3] Qdrant, “Releases,” qdrant/qdrant, GitHub, accessed Sep. 24, 2026. https://github.com/qdrant/qdrant/releases

[4] Qdrant, “Optimizer,” Qdrant Documentation, accessed Sep. 24, 2026. https://qdrant.tech/documentation/ops-optimization/optimizer/

[5] S. Eskildsen and J. Lauritzen, “Connect AI to Billions of Legal Documents,” AI Engineer, Sep. 2026. https://www.youtube.com/watch?v=V-isu4eTHgw

[6] D. Dalakoti, R. Lewis, and P. V. Avramovic, “Selecting a vector store for Amazon Bedrock Knowledge Bases,” AWS Machine Learning Blog, Sep. 17, 2026. https://aws.amazon.com/blogs/machine-learning/selecting-a-vector-store-for-amazon-bedrock-knowledge-bases/

[7] G. Dotzlaw, “Agentic Performance Audits: From Production Blind Spots to ROI-Scored Fixes,” Dotzlaw Consulting, Sep. 14, 2026. /insights/ai-40-agentic-performance-audits/

Qdrant in Production: The Slow Window After a Bulk Load, and Why Your Median Hides It
https://dotzlaw.com/insights/ai-51-qdrant-optimizer-numbers/
Author
Gary Dotzlaw
Published at
2026-10-01
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

← Back to Insights