How to Make Multitenant Vector Indices Scalable

12 min read
Hunter Zhao
Engineering

Short answer

Multitenant vector indices scale reliably when you match the isolation model to your tenant profile before you write any indexing code. Shared indices with metadata filters work for hundreds of uniform tenants; namespace-per-tenant handles thousands; physical separation, which is how Powabase projects are structured by default, is the right call when tenants have regulatory exposure or wildly different sizes.

A single HNSW graph holding every tenant's vectors is the fastest way to build a multitenant RAG app, and the fastest way to wreck it in production. One shared index means one tenant's bulk ingest stalls everyone else's queries, metadata filters degrade recall, and a GDPR deletion request forces a graph rebuild. Scaling a multitenant vector index is really a question of where you draw the isolation boundary, and that choice has to survive a tenant distribution where your top 1% hold more vectors than the bottom 90% combined.

This guide walks through the five choices that matter: isolation strategy, ANN algorithm tradeoffs, sharding, platform-specific patterns, and the compliance surface. Our bias is for physical isolation per project, because that's how Powabase ships every workspace: a dedicated Postgres with pgvector, so there's no shared logical database and no cross-tenant graph to escape from.

Why multitenant vector indices are hard to scale

A multitenant vector index is simultaneously a storage problem, a query-routing problem, and an isolation problem. Each tenant has its own corpus, its own query cadence, its own compliance expectations, and, crucially, its own size. Build for the median tenant and one whale breaks you; build for the whale and you're renting idle capacity for the 95% who'll never use it.

The tenant size discrepancy and power-law distribution problem

Tenant corpora almost never cluster around an average. Weaviate's team, after rebuilding their multitenancy stack, now supports 50,000+ active tenants per node and a 20-node cluster for a million active tenants holding billions of vectors in aggregate, a design that only makes sense once you accept that most tenants are tiny and a handful are enormous. A shared HNSW graph sized for a 10M-vector tenant wastes memory on tenants with 2,000 vectors; a flat index tuned for small tenants collapses when a whale onboards.

The noisy-neighbor problem and how it degrades query latency

Shared infrastructure means shared contention. When one tenant triggers a bulk re-embed, HNSW insertions grab locks, caches evict, and p99 latency spikes for every other tenant on the shard. Research on multi-tenant RAG explicitly flags noisy neighbor effects in a vector database alongside tenant isolation as the two defining constraints of enterprise RAG. The only durable fixes are hard resource quotas per tenant or physical separation of the index itself.

Tenant isolation strategies: the core design choice

Isolation sits on a spectrum. On one end, every tenant shares a table with a tenant_id column. On the other, every tenant gets a dedicated database on dedicated compute. Everything else is a hybrid.

Shared index with metadata filtering (tenant_id pre-filter)

The simplest pattern: one collection, a tenant_id field on every vector, a filter clause on every query. Nile describes this as the tenant column approach, where developers add a filter to each query that limits the response to the current tenant. Minimal overhead, maximum operational simplicity, weakest isolation guarantees.

It works at small scale. It stops working when a shared HNSW graph has to prune millions of candidates to find the few hundred belonging to a given tenant. Pinecone explicitly warns against the degenerate version of this pattern, stuffing every user into one namespace and filtering by ID. Their docs call out that large $in filters increase payload size and latency, and that each filter operator caps at 10,000 values, after which requests fail outright.

Namespace or collection per tenant

The middle ground is one logical container per tenant inside a shared cluster: Pinecone namespaces, Weaviate tenants, Qdrant payload groups, Milvus partitions. Pinecone's own guidance is to target each tenant's queries at the namespace for the tenant so one tenant's query load never affects another's, a form of one index per tenant that still shares a control plane. Running one index per tenant also makes tenant offboarding and data deletion a single destructive call rather than a graph rebuild.

Database or silo per tenant (physical isolation)

At the strict end, each tenant gets a dedicated database on dedicated compute. No shared logical state, no shared query planner, no shared buffer pool. Nile characterizes this as the "database per tenant" approach that maximizes isolation and suits a smaller number of high-paying customers. It's how we provision Powabase projects: every project gets its own Postgres instance with pgvector preloaded. The tradeoff is operational: more databases to patch, back up, and monitor. Noisy neighbors stop being a design concern — because there are no neighbors.

Choosing between physical and logical isolation

Pick physical when tenants are large, regulated, or paying enterprise rates. Pick logical when tenants are many, small, and homogeneous. Most SaaS products end up with both: a tiered scheme where free and starter tenants share a namespace per tenant, and enterprise tenants get their own database.

Per-tenant indexing vs. shared index: the ANN algorithm tradeoffs

Isolation decides where tenant data lives. Indexing decides what structure you build over it. The two choices interact: some algorithms degrade badly under one tenant per index, others degrade badly under shared indexing.

HNSW per-tenant vs. shared HNSW performance and memory cost

HNSW is the default for a reason: fast queries, high recall, predictable behavior up to tens of millions of vectors per graph. The catch is memory. Every graph carries overhead for its upper layers and entry points, so 10,000 tenants with 1,000 vectors each in 10,000 separate HNSW graphs waste a huge fraction of RAM on scaffolding. HNSW multitenant designs only pay off when each tenant has enough vectors to amortize the per-graph overhead.

Flat and IVFFlat indexes for many small tenants

For the long tail of small tenants, flat (brute-force) indexes often beat HNSW. There's no graph to build, no memory overhead beyond the vectors themselves, and at a few thousand vectors a sequential scan with SIMD is milliseconds. MongoDB's multi-tenant architecture guide recommends flat indexes for many small tenants, optionally with scalar or binary quantization to shrink the memory footprint further.

Curator and clustering-tree approaches to efficient indexing

Qdrant's team recommends a different trick for high-throughput ingest: bypass the construction of a global vector index and build smaller per-group indexes instead, so indexation throughput isn't bottlenecked by a monolithic graph. Research prototypes like Curator generalize this with per-tenant clustering trees over a shared quantized backbone.

Sharding and partitioning for scale

Once you've picked an isolation level and an index type, the next question is how to distribute data across physical machines.

User-defined and custom sharding by tenant_id

Default sharding spreads data by hash, which is terrible for multitenancy. One tenant's vectors scatter across every shard, so every query fans out to every node. Custom sharding for vector search fixes it: Qdrant exposes a shard_key_selector that pins each tenant's vectors to a specific shard, so queries hit one node. In Postgres, a tenant_id btree index plus partitioning by tenant_id gets you the same locality.

Tiered multitenancy and promoting tenants to dedicated shards

Tiered multitenancy treats tenant size as a lifecycle. Small tenants start on a shared shard with other small tenants. When a tenant crosses a threshold (vector count, query rate, revenue) they're promoted to a dedicated shard with a dedicated index. This matches how Pinecone, Qdrant, and Weaviate production users actually operate, and it's the pattern our per-project isolation defaults to from day one.

Active vs. inactive tenant offloading

Most SaaS tenants are inactive most of the time. Keeping their HNSW graphs in RAM is pure waste. Weaviate's native multitenancy distinguishes active from inactive tenants so you aren't paying compute for idle users, trading a cold-start penalty for a large reduction in resident memory. On Postgres, the equivalent is leaving a tenant's partition on cheaper storage and letting the OS page it in when queried.

Platform-specific implementation patterns

The abstractions above map differently onto each vector platform. Here's how the major options actually implement multitenancy.

Pinecone serverless namespaces

Pinecone's recommended pattern is one namespace per tenant, with queries targeted to a specific namespace at request time so one tenant's read and write load can't affect another's, and offboarding reduced to deleting the namespace. Namespaces scale independently and are the right default for Pinecone-based designs.

Weaviate native multitenancy with lightweight shards

Weaviate rebuilt its multitenancy stack around the premise that a tenant should be a lightweight shard, not a filter. Weaviate reports 50,000+ active tenants per node and millions of tenants per cluster, though the practical ceiling is the per-process open-file limit rather than a fixed number: their own nine-node test cluster held roughly 170,000 active tenants, nearer 19,000 per node. Inactive tenants can be offloaded, which takes them out of that budget entirely.

Qdrant payload partitioning and tiered multitenancy

Qdrant combines a group_id payload field with custom shard keys. The group field scopes queries logically; the shard key scopes them physically. Where many tenants share a collection, Qdrant's advice is to set HNSW m to 0, which disables the collection-wide index so vectors get indexed per tenant instead.

Milvus partition key isolation

Milvus exposes four strategies: database, collection, partition, and partition key. Partition-key isolation is the only one their comparison table marks as physical and logical: a physical partition can hold several tenants while keeping their data logically separate, which lets one collection scale to large tenant counts without per-partition overhead.

pgvector and Aurora PostgreSQL with Row-Level Security

In Postgres, isolation lives at the row level. A tenant context is set per session (via SET LOCAL or a GUC bound from a JWT claim), and pgvector row-level security policies gate every query. The exact policy, session binding, and index arrangement matter. A poorly placed policy predicate causes the planner to fall back from HNSW to a sequential scan, which is catastrophic. The tenant-scoped ANN index pattern uses a partial HNSW with a WHERE clause on tenant_id so the index survives the RLS filter. The verification pattern is also non-negotiable: set the context to tenant A, query rows belonging to tenant B, assert empty.

Powabase takes this one level further. Rather than relying on RLS alone, every project gets its own dedicated Postgres with pgvector. RLS still protects end-users within a project, but there is no shared logical database to escape from in the first place.

MongoDB Vector Search and Azure Cosmos DB sharded DiskANN

MongoDB recommends flat indexes with quantization for tenants with small corpora. Azure Cosmos DB's sharded DiskANN takes the physical-partition route, confining each search to the relevant shard's graph rather than fanning out across the whole index.

Security, access control, and compliance

Isolation isn't just a performance concern. A cross-tenant leak is a reportable incident, and in regulated verticals it's a contract-ending one.

Enforcing isolation with RLS, JWT tenant context, and RBAC

The defense-in-depth pattern is: tenant ID in a signed JWT claim, extracted by the API gateway, bound to a Postgres GUC, enforced by RLS policies on every table. Powabase's PostgREST layer respects RLS policies by default, so browser-side queries stay inside the tenant's rows. The same discipline applies to vector columns: a SELECT ... ORDER BY embedding <=> $1 against a shared table must carry the tenant predicate the HNSW planner can push down.

For a deeper treatment of agent-driven RLS, see our write-up on whether AI agents should run as the user or as a service role.

Preventing cross-tenant embedding leakage

Embeddings themselves are sensitive. A nearest-neighbor result that leaks across tenants doesn't just reveal a document ID, it reveals semantic content. Three failure modes to design against: RLS policies missing on a secondary table joined into the retrieval query; caches keyed on query embedding without a tenant prefix; and retrieval pipelines that fall back to a global corpus when the tenant corpus returns zero results. The last is particularly insidious because it looks like a UX improvement.

GDPR-compliant tenant offboarding and data deletion

A tenant invoking Article 17 has to result in cryptographic certainty that their vectors are gone. In a shared HNSW graph, that means rebuilding the graph; tombstones aren't deletion. In a per-tenant collection or database, it's a DROP. This is one of the strongest arguments for physical isolation at the project level: tenant offboarding and data deletion is a single destructive operation, not a background job that leaves residual edges in a graph for weeks.

Building a scalable multitenant RAG application

RAG amplifies every multitenancy mistake. Retrieval errors become hallucinations grounded in another tenant's data, much worse than a stale cache.

Tenant-scoped retrieval and LLM context isolation

Our walkthrough of tenant isolation in multi-tenant RAG covers the retrieval side of this in more depth. Scope at every stage. Embed with a tenant-aware pipeline (so model fine-tunes don't cross boundaries), retrieve from a tenant-scoped index, pass only tenant-owned chunks into the LLM context, and log the retrieval set with the tenant ID so audits can reconstruct what was shown. Powabase's standalone context handlers make this explicit: you call retrieval against a specific knowledge base, get the chunks back, and compose your own LLM call. The knowledge base IDs are the tenant boundary, and they live inside a project that is already physically isolated.

For teams building custom hybrid retrieval over the ai.chunks table directly, the AI schema recipes show how to use PostgREST filters to scope across multiple knowledge bases in a single query without application-side fan-out, useful when a single tenant has several corpora.

Capacity planning and choosing the right strategy

The right strategy depends on how many tenants you have, how big they are, and how much they vary.

Scalability limits by platform

Rough ceilings worth knowing: Milvus defaults to 64 databases per cluster (configurable), making database-per-tenant unworkable past a few dozen enterprise accounts. Pinecone namespaces scale into the tens of thousands per index. Weaviate supports 50,000+ active tenants per node. Postgres can host thousands of schemas per database, but connection pooling becomes the bottleneck long before storage does, hence our per-project Postgres model.

A decision framework for your tenant profile

A pragmatic rubric:

Tenant profileRecommended strategy
<100 tenants, large and regulatedDatabase or project per tenant
100–10,000 tenants, mixed sizesNamespace/collection per tenant, tiered promotion for whales
10,000+ tenants, mostly smallShared index with tenant partition key, flat indexes per group
Hybrid enterprise + self-servePer-project isolation for paid tiers, shared namespace for free

The temptation at the start is to pick the cheapest option and migrate later. Migrating vector indices across isolation boundaries is painful: you're re-embedding, re-indexing, re-validating recall, and coordinating cutover without dropping queries. Start on a platform where isolation is enforced at the infrastructure level from day one and the hardest architectural decision has already been made in your favor.

FAQ

It depends on tenant count and risk tolerance. A shared index with a tenant_id filter is cheapest for small, uniform tenants. Namespace-per-tenant (Pinecone, Weaviate, Qdrant) handles thousands with reasonable isolation. Physical isolation, one database per tenant, is the right choice when tenants have compliance requirements or data that can never mix.

Every Powabase project gets its own Postgres StatefulSet with pgvector, its own API gateway, auth service, and storage in a dedicated Kubernetes namespace. That gives physical separation at the project level, with row-level security available inside each project for finer-grained scoping. A leaked API key cannot reach another project's vectors.

A monolithic HNSW graph reflects the data distribution of every tenant, so a query from a tenant with 500 vectors still traverses a graph built across billions. Recall drops, latency rises, and one tenant running a bulk re-embed degrades everyone else. Per-tenant graphs or flat indexes for small tenants avoid this.

With physical or namespace-per-tenant isolation, deletion is a DROP or namespace delete that covers chunks, embeddings, metadata, and the ANN graph entries in one operation. A shared index forces a DELETE plus reindex plus cache invalidation across every table that stored tenant data, which is error-prone and slow for regulated workloads.

For tenants holding fewer than a few thousand vectors, a brute-force flat scan is faster than HNSW graph traversal setup. MongoDB's multi-tenant guidance recommends flat indexes for small tenants with scalar or binary quantization for memory efficiency. A tiered design, flat for the long tail and HNSW for large accounts, is the practical default.

You enable RLS on the shared chunks table, bind the tenant identifier to the session, and every query including similarity search is automatically scoped to that tenant's rows. The policy fires before the ANN scan returns results. Switching to another tenant's session context and querying rows belonging to the first tenant returns nothing.

Keep reading

multitenant vector indices

Share this article