Back to RAG & Vector Databases
RAG & Vector Databases

How do metadata filtering and HNSW index parameters affect query latency in multi-tenant RAG systems?

Practical answer and configuration guide for How do metadata filtering and HNSW index parameters affect query latency in multi-tenant RAG systems?.

G
Gaurav Bhasin 👑 Tier 3 Elite
Aug 9, 2026 · 1 min read

Most RAG quality issues come from bad document chunking rather than the LLM model itself. Here is how to optimize retrieval:

1. Use Semantic Structure Splitting: Split Markdown and PDFs on header boundaries (`
## `) rather than arbitrary character counts to keep tables and code blocks intact.

```python
from langchain.text_splitter import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
separators=["
## ", "
### ", "

", "
", " "]
)
```

2. Combine Vector + Keyword Search: Pair vector embeddings with BM25 keyword search, then pass top results to a Cohere Reranker model. This catches both semantic context and exact product/code matches.

3. Parent-Child Indexing: Store 128-token chunks for vector retrieval, but return the surrounding 1024-token parent section to the LLM.

Read the evidence

Sources used in this thread

Open the original material, compare the claims, and form your own view.

Community notes

Add context, not noise (1)

Corrections, lived experience, useful examples, and better sources belong here.

A
2 hours ago
👍 0 Upvotes

Audit your permission set policies regularly using AWS IAM Access Analyzer to catch any wildcard `*` permissions that creep in over time.

Click here to write a reply...
🔒

Authentication Required

Join Trendzza to begin your journey. Submit tasks, complete batches, help peers, and earn your way to Tier 3.