The Failure That Returns Results
Here’s the thing about filtered vector search that took me too long to internalize: when it breaks, it doesn’t break. No error, no timeout, no empty response. You get a handful of results, a bit weak, a bit off-topic. And if an agent sits on top of that, it reads “weak results” as “nothing relevant exists” and moves on with total confidence.
That’s why I think it’s the most important engineering fact in retrieval, more than which embedding model you pick or whether you rerank.
So let me go through the mechanism, the numbers, what the engines do about it, and then the part I care about most: what it means for multi-tenant systems and for the order of operations on the read path.
Why a Filter Shreds the Graph
A graph index like HNSW works because it’s navigable. Each node has enough links that you can reach any neighbourhood in a few hops. With default parameters on a 1M-point set you get roughly 21 links per node.
Now filter out 96% of the points (so the filter matches 4%). The removed nodes are gone from traversal. Per the vendor’s own write-up, that leaves “fewer than one link per node on average”. The graph is shredded.
The query still traverses. It just traverses a graph with holes in it and returns whatever it happened to reach. Nothing in the API tells you.
The measurements, from the same vendor self-evaluation (so take it as directional, one vendor measuring its own engine):

Plain HNSW recall as the filter gets more selective. Collapse is reported all the way down to 0.012% selectivity.
- 62.9% recall at 20% selectivity
- 0.1% recall at 1% selectivity
- collapsed at 0.012%
Two more data points that point the same way. The pgvector docs say it plainly: if a condition matches 10% of rows, with HNSW and the default hnsw.ef_search of 40, only 4 rows will match on average. And it’s not just graphs. A peer-reviewed ICML 2025 paper from Pinecone says simple filtering in IVF “usually does not work”, and at 90% filtered out the results are “unusable”.
Why does 1% matter so much? Because that’s exactly where a tenant filter lives. 100 tenants on one shared index means one tenant’s slice is about 1% of the corpus. You’re in the 0.1% recall regime on the most common filter in a multi-tenant system.
What the Engines Do About It
Everyone has a different answer, which tells you the problem is real and not solved:
- Qdrant filters in place during traversal and runs a per-query planner. Two repairs: filterable HNSW (extra edges at index time between points sharing an indexed value) and ACORN-1 (neighbours-of-neighbours at query time).
- Weaviate has explicit pre-filtering with an allow-list, ACORN as the default strategy in recent versions, and brute force below a configurable fraction of the dataset. I couldn’t verify a quantitative claim for it, only the qualitative one.
- Pinecone compiles the filter to a compressed bitmap over immutable slabs that each carry their own index. It reports mean recall@10 of 0.989 flat across selectivity ranges at about 20 ms. The same paper concedes that graph plus filter remains an open problem.
- Vespa hands the pre/post/exact decision to the operator, with thresholds, one of which is documented as improving response time at potential recall cost. I like the honesty.
- pgvector has iterative index scans since 0.8.0, with a bounded scan-tuple limit.
The most useful result is that no single repair dominates. Vendor-published numbers, same engine:
- 1% selectivity, single filter: filterable HNSW gets 99.8% recall at 1.0 ms; ACORN-1 plateaus at 67.7% at 4.7 ms.
- 4% selectivity, AND of filters: filterable HNSW drops to 63.7%; ACORN-1 gets 95.2%.
- Cost side: filterable HNSW builds 4.4 to 5.6x slower (507-652 s vs 116 s); ACORN-1 has no index cost but 2.1 to 2.9x query latency.
So the design rule I take from this: a stable, single-valued dimension like tenant wants an index-time repair, or better a physical partition. Ad-hoc dimensions that combine arbitrarily (service, severity, time window, release) want a query-time repair, because you can’t pre-build for every combination.
One caveat I want to be loud about. There is no neutral public benchmark for filtered ANN at tenant-scale selectivity. Every number above comes from a vendor or a paper using its own harness. That’s a reason to measure on your own corpus, not a reason to assume the numbers are wrong. I haven’t measured mine yet, and honestly that’s the first thing to do.
A Predicate Is Not Isolation
Now the security side. If you do tenant separation as WHERE tenant = ?, you have two problems at once: the recall collapse above, and the fact that one forgotten or injected clause returns another tenant’s content with no error. Same silent-failure shape.
I rank the options in four rungs:
- Physical partition per tenant (namespace, tenant shard, partition key with isolation). Isolation is a property of the data layout, the only class where a forgotten predicate cannot leak. Foot-gun: one engine documents that deleting a tenant deletes its shard and all its objects. Fast right-to-erasure, also a sharp edge.
- Payload partitioning with a tenant hint at the index level. Good for the long tail of small tenants, with a documented path to promote a big one to rung 1. Documented cost: global requests without the tenant filter are slower because they scan all groups.
- Pure metadata filtering in one namespace. I’d disqualify this. Three independent reasons: cost, recall (it’s exactly the section above), and the leak risk. On cost, one vendor’s own number is 1 read unit for a namespace query versus 100 for the same query as a metadata filter on a 100 GB namespace, because the filter scans the entire namespace regardless.
- Row-level security. The most leak-prone, and the docs say so themselves.
On rung 4 the Postgres docs list the bypasses: superusers and BYPASSRLS roles always bypass; table owners normally bypass unless forced; security-definer functions can access data unavailable to the caller; referential integrity checks always bypass, with care needed to avoid covert-channel leaks. Think about what that means for a search function written with definer rights: it silently returns everything.
It’s defensible only with forced enforcement, no definer-rights functions near the search path, invoker-rights views, no bypass role in the app path, and explicit acceptance of the covert channel. That’s a checklist you can fail silently. I still think “isolation below the application” is the right instinct. I just think relational RLS is the weakest place to put it.
The Fusion Window Is a Correctness Knob
Small detour, same theme. Hybrid search usually means reciprocal rank fusion (a two-page SIGIR 2009 short paper originally; I verified the bibliographic record, not the full text). Implementations expose a rank constant and a window size, and the window is treated as a latency knob: 50 is faster and slightly worse, 500 is slower and slightly better.
A recent preprint argues that’s wrong. Its claim is that truncated fusion is not generally equivalent to complete-list fusion: unread cross-list ranks can change top-K membership or order. Making channel depth per-request reproduced complete-list top-20 across 150 query-snapshot combinations, with a large latency improvement over exhaustive fusion.
It’s one preprint, so I’d hold it loosely. But the framing is the useful part: if you tune the window for latency, your result set is not the result set. It’s another parameter that changes answers without changing anything you’d see in a log.
And when do you need hybrid at all? When documents contain exact tokens like service names and error codes that dense retrieval blurs. That’s the trigger. Before that corpus exists, I wouldn’t bother. Also BEIR’s headline is that BM25 is a robust baseline (NeurIPS 2021 Datasets and Benchmarks), so whatever you build gets measured against keyword search, not against nothing.
The Whole Read Path, in Order
Here’s the ordering I’d use. The principle is filter before rank, and make the filters deterministic, not search-based:

The read path: deterministic filters, then search and rank over what survived, then a last check before the model call.
A few notes on the non-obvious parts:
- Steps 1 to 6 happen before any search. Authority, applicability, validity and purpose are hard predicates, and unknown means withhold, not include.
- Retrieval only ever sees what survived. That’s what makes the recall problem a design question (partition or repair) and not a surprise.
- De-duplicate by event identity, not text similarity. Four paraphrases of the loudest signal are not four pieces of evidence. For selection, start with maximal marginal relevance (1998, free, understood) and only escalate if the pairwise penalty proves too crude.
- Drop whole units when fitting to budget and log every drop with a reason. A silently truncated item is the same class of bug as everything above.
- The final revocation gate is the one I rarely see. Steps 8 to 11 take time. A revocation that lands in that window would otherwise slip through, so you re-read each item’s state immediately before the model call.
Make the Silent Failure Loud
Back to the original problem. From the agent’s side, “nothing matched” and “the index could not be traversed” look identical. You can’t fix that by being careful with parameters.
The cheapest fix I know: plant one canary item per filter scope that must always be retrievable under that filter. If retrieval returns the canary, the index is healthy and “nothing matched” is true. If it misses the canary, the traversal is broken, so alarm and don’t conclude anything.
That’s one item per scope, and it turns the failure from silent to loud. Of everything in this post it has the best ratio of effort to value.
What I’m still unsure about: whether per-tenant indexing changes recall at all. As far as I could find, nobody has published multi-tenant retrieval quality. All the multi-tenancy documentation talks about isolation and cost, none reports recall. So the partition-first advice rests on isolation and on the mechanism, not on a measured recall comparison. Worth testing before you take my word for it.
References
- Qdrant, filtered vector search and ACORN article (vendor self-evaluation)
- pgvector README (vendor docs)
- Pinecone, ICML 2025 (PMLR 267)
- ACORN, arXiv 2403.04871
- Curator, arXiv 2601.01291
- Fusion window preprint, arXiv 2608.07152
- BEIR, arXiv 2104.08663