Metadata Filtering in RAG Pipelines

RAG
AI
PostgreSQL
pgvector
How to use metadata filters to scope vector similarity search — and why filtering before retrieval beats filtering after.
Author

Aakriti Dhakal

Published

November 1, 2025

The naive RAG setup retrieves the top-k most semantically similar chunks and passes them to the LLM. This works until your vector store contains documents from multiple contexts — different courses, different users, different domains — and you start getting cross-context contamination in responses.

Metadata filtering fixes this.

Pre-filter vs. post-filter

Post-filter: Retrieve top-k, then discard chunks that don’t match your metadata criteria. Simple but wasteful — if most of your data is in the wrong category, you’ll discard most of your results and end up with low effective k.

Pre-filter: Apply metadata filters before similarity search, then retrieve top-k from the filtered subset. This is almost always what you want.

With pgvector, pre-filtering is a WHERE clause:

SELECT id, content, 1 - (embedding <=> $1) AS similarity
FROM document_chunks
WHERE metadata->>'course_id' = $2
  AND metadata->>'content_type' = 'lecture'
ORDER BY embedding <=> $1
LIMIT 10;

The index interaction problem

If your filter selects a very small fraction of rows (< 5%), PostgreSQL may fall back to a sequential scan rather than the index. You can check with EXPLAIN:

For highly selective filters, consider a partial index:

CREATE INDEX ON document_chunks
USING hnsw (embedding vector_cosine_ops)
WHERE metadata->>'content_type' = 'lecture';

Metadata schema design

Metadata that you’ll filter on should be promoted to top-level jsonb keys (or dedicated columns).

Good:

{ "course_id": "CS4320", "content_type": "lecture", "week": 3 }

Multi-tenant isolation

For multi-tenant RAG (e.g., per-user or per-course knowledge bases), metadata filtering is your access control layer. Every query must include a tenant filter, and that filter should be enforced at the application layer:

def retrieve(query: str, course_id: str, k: int = 10) -> list[Chunk]:
    embedding = embed(query)
    return db.query(
        "SELECT ... WHERE metadata->>'course_id' = $1 ORDER BY embedding <=> $2 LIMIT $3",
        course_id, embedding, k
    )

Never let course_id come directly from user input without validation against the authenticated user’s allowed courses.