Building Production-Grade RAG Pipelines with Vector Embeddings and Hybrid Search
Architecting high-accuracy Retrieval-Augmented Generation systems using chunking strategies, dense vector embeddings, BM25 sparse keyword ranking, and reciprocal rank fusion.
System Design: From Zero to Production
A comprehensive engineering series guiding backend developers from single-node instances to highly resilient distributed architectures.

Why Pure Vector Search Fails in Production
Semantic vector search is powerful for retrieving documents based on meaning, context, and conceptual similarity. However, in production environments, semantic similarity alone is not sufficient for reliable information retrieval.
Consider a technical support query:
Why does error
P2025occur when using Prisma with PostgreSQL?
A vector search engine may retrieve documents about Prisma database errors in general, even when a document explicitly explaining P2025 exists but receives a lower semantic similarity score. Similar problems occur with API endpoints, product SKUs, version numbers, configuration keys, exception messages, and other exact-match identifiers.
Hybrid search addresses this limitation by combining dense vector retrieval with sparse keyword retrieval. Dense retrieval captures semantic relationships, while sparse retrieval identifies documents containing relevant keywords and technical identifiers.
When combined with Reciprocal Rank Fusion (RRF) and an optional reranking stage, hybrid retrieval can improve the relevance of the context supplied to a Retrieval-Augmented Generation (RAG) pipeline.
Key takeaway: Vector search helps find documents that mean the same thing. Keyword search helps find documents that contain the exact terms. Hybrid search combines both signals.
1. Why Pure Vector Search Is Not Enough
Vector search converts queries and documents into numerical embeddings. A similarity function, such as cosine similarity, is then used to retrieve documents whose embeddings are close to the query embedding.
This works well for natural-language questions and conceptually related content, but several limitations become apparent in production.
1.1 Exact-Match Retrieval Problems
Embedding models do not guarantee that rare identifiers or exact technical strings receive sufficient weight in the similarity calculation.
For example, these identifiers may have very different operational meanings:
ERR_CONNECTION_RESETERR_CONNECTION_REFUSEDHTTP 401HTTP 403PostgreSQL 16PostgreSQL 17
A semantically similar document is not necessarily the document that contains the exact error code, version, or configuration value the user needs.
1.2 Semantic Drift
A query about a specific implementation detail may retrieve conceptually related but operationally irrelevant documents.
For example, a search for a Prisma connection-pool timeout could retrieve general database performance guidance instead of the configuration documentation that explains the actual timeout parameter.
1.3 Rare Terms and Technical Identifiers
Technical documentation frequently contains tokens that appear infrequently in a corpus:
- API names and endpoint paths
- Exception messages and error codes
- Package names and dependency versions
- Database table and column names
- Environment variables
- Configuration keys and command-line flags
Sparse retrieval methods such as BM25 are particularly useful when these exact terms determine relevance.
1.4 Similarity Scores Are Not Relevance Guarantees
A high vector similarity score indicates that two embeddings are close under the chosen similarity metric. It does not guarantee that the retrieved document contains the exact evidence required to answer the question.
This distinction is important in production RAG systems, where irrelevant context can reduce answer quality and increase the risk of unsupported responses.
2. What Is Hybrid Search?
Hybrid search combines two complementary retrieval approaches:
| Retrieval method | How it works | Primary strength |
|---|---|---|
| Dense vector search | Compares query and document embeddings | Semantic similarity and paraphrases |
| Sparse keyword search | Matches query terms against indexed document terms | Exact keywords, identifiers, and rare terms |
| Hybrid retrieval | Combines rankings from both methods | Broader retrieval coverage |
Dense Retrieval
A dense embedding model converts text into a fixed-dimensional vector. For example, text-embedding-3-small supports 1,536-dimensional embeddings by default.
A vector database can use an approximate nearest-neighbor index, such as HNSW, to retrieve candidate documents efficiently.
Dense retrieval is useful when the query and document use different wording but express similar concepts.
Sparse Retrieval with BM25
BM25 is a probabilistic information-retrieval ranking function that scores documents using factors such as term frequency, inverse document frequency, and document length normalization.
It is useful when a query contains specific terms that should match the indexed content.
Implementation note: PostgreSQL provides built-in full-text search through tsvector, tsquery, and ranking functions such as ts_rank_cd. These provide lexical retrieval, but they are not automatically equivalent to a full BM25 implementation. If BM25 is a requirement, use an implementation that explicitly supports it, such as a suitable search engine or a PostgreSQL extension.
3. Hybrid RAG Architecture
A hybrid RAG pipeline retrieves candidate documents using both dense and sparse search, merges their rankings, optionally reranks the candidates, and passes the most relevant evidence to the language model.
Architecture DiagramMermaid Flow
Retrieval Pipeline Explained
- Query processing: Normalize the query where appropriate and extract important technical terms, identifiers, and filters.
- Dense retrieval: Generate an embedding for the query and retrieve semantically similar document chunks.
- Sparse retrieval: Search the lexical index for relevant keywords and identifiers using BM25 or another suitable ranking method.
- Candidate fusion: Merge the ranked lists using Reciprocal Rank Fusion.
- Reranking: Optionally use a cross-encoder to evaluate query-document relevance more directly.
- Context assembly: Select relevant, sufficiently diverse chunks while respecting the model's context budget.
- Generation: Provide the selected evidence to the language model and require it to distinguish supported facts from missing information.
The candidate counts and final context size should be tuned to the corpus, latency budget, and evaluation results rather than treated as universal defaults.
4. Reciprocal Rank Fusion (RRF)
Reciprocal Rank Fusion combines ranked search results without requiring the scores from different retrieval systems to be directly comparable.
This is useful because vector similarity scores and BM25 scores generally have different scales and distributions. Adding them directly can cause one retrieval system to dominate unless their scores are calibrated or normalized.
RRF instead assigns a contribution based on each document's rank in each result list.
RRF Formula
For a document (d), its fused score is:
Where:
- (R) is the set of ranked retrieval result lists.
- (\operatorname{rank}_r(d)) is the one-based rank of document (d) in result list (r).
- (k) is a smoothing constant, commonly set to 60.
- Documents absent from a result list contribute zero for that list.
A document ranked highly by both dense and sparse retrieval receives contributions from both lists, increasing its fused score.
Example
Assume a document appears in both retrieval lists:
| Retrieval source | Document rank | RRF contribution with (k=60) |
|---|---|---|
| Dense vector search | 1 | (1/61 \approx 0.01639) |
| Sparse keyword search | 3 | (1/63 \approx 0.01587) |
| Combined score | — | 0.03226 |
A document appearing only in the first position of the dense list receives a score of approximately (0.01639). The document appearing in both lists receives a higher combined score because both retrieval systems support its relevance.
RRF does not prove that a document is relevant. It combines ranking evidence and should be evaluated against representative queries and labeled relevance judgments.
5. TypeScript Implementation of RRF
The following implementation accepts already-ranked dense and sparse result lists. Each list must be ordered from most relevant to least relevant.
typescriptexport type RankedDocument = { id: string; score: number; }; export type FusedDocument = { id: string; rrfScore: number; }; export function reciprocalRankFusion( denseResults: RankedDocument[], sparseResults: RankedDocument[], k = 60, ): FusedDocument[] { if (!Number.isFinite(k) || k <= 0) { throw new Error("k must be a positive finite number"); } const scoreMap = new Map<string, number>(); const addRankedResults = (results: RankedDocument[]) => { results.forEach((doc, index) => { const rank = index + 1; scoreMap.set( doc.id, (scoreMap.get(doc.id) ?? 0) + 1 / (k + rank), ); }); }; addRankedResults(denseResults); addRankedResults(sparseResults); return Array.from(scoreMap, ([id, rrfScore]) => ({ id, rrfScore, })).sort((a, b) => b.rrfScore - a.rrfScore); }
Implementation Considerations
- Ranks must be one-based: The first result receives rank 1, not rank 0.
- Input ordering matters: RRF uses the position of each result, not its supplied
scorefield. - Scores are intentionally unused: The
scorefield is retained for compatibility and debugging, but RRF uses rankings only. - Duplicate document IDs: Repeated IDs within a single result list contribute multiple times in this simple implementation. Deduplicate each list before fusion if duplicate entries are possible.
- Stable tie-breaking: If deterministic output is required, add an explicit secondary sort key.
- Metadata preservation: In a production implementation, maintain a document lookup map so fused results can be enriched with document text, source URLs, access permissions, and metadata.
For systems with more than two retrieval sources, apply the same rank-based accumulation to each list.
6. Reranking After Hybrid Retrieval
RRF is effective for merging candidate lists, but it does not deeply evaluate the semantic relationship between the original query and each candidate document.
A cross-encoder reranker evaluates a query-document pair together and produces a relevance score. Because it performs more computation per candidate, it is generally applied after the initial retrieval stage.
A typical pipeline is:
- Retrieve 20–100 candidates from each search system, depending on corpus size and latency requirements.
- Merge and deduplicate the candidates with RRF.
- Select a manageable candidate pool for reranking.
- Rerank the candidates using a cross-encoder.
- Select the final context chunks based on relevance, diversity, permissions, and context limits.
Reranking can improve the ordering of retrieved candidates, but it cannot recover a relevant document that neither retrieval system returned.
7. Benchmarking Hybrid Search Against Vector Search
Hybrid retrieval should be evaluated against a representative test set rather than assumed to be better in every workload.
For a technical documentation corpus, build a labeled dataset containing natural-language questions, exact-identifier queries, version-specific questions, and paraphrased queries. Compare the systems using identical query sets and relevance judgments.
Recommended Evaluation Metrics
| Metric | What it measures |
|---|---|
| Recall@5 | Fraction of relevant documents retrieved within the first five results |
| Recall@20 | Fraction of relevant documents retrieved within the first twenty results |
| MRR | How highly the first relevant result is ranked |
| nDCG@k | Ranking quality when relevance has multiple grades |
| Retrieval latency | Time spent retrieving and ranking candidate documents |
| Answer faithfulness | Whether generated claims are supported by retrieved evidence |
| Citation accuracy | Whether cited sources actually support the associated claims |
Interpreting Benchmark Results
A benchmark might report the following hypothetical results:
| Metric | Pure vector search | Hybrid search |
|---|---|---|
| Recall@5 | 64.2% | 89.4% |
| Recall@20 | Measure experimentally | Measure experimentally |
| Retrieval latency | Measure experimentally | Measure experimentally |
| Answer faithfulness | Measure experimentally | Measure experimentally |
Important: The Recall@5 figures above are illustrative values, not independently verified benchmark results. Replace them with measured results from your own evaluation dataset before publishing them as empirical findings.
A claim such as an 82% reduction in hallucination rate requires a separate controlled evaluation of generated answers. Retrieval recall alone cannot establish that result.
8. Common Production Challenges
Query and Document Preprocessing
Use consistent normalization where appropriate, but preserve meaningful distinctions in identifiers, versions, and case-sensitive strings. Aggressive normalization can make technically different identifiers appear equivalent.
Chunking Strategy
Chunk size affects both retrieval precision and the amount of evidence available to the language model. Overly large chunks can introduce irrelevant content, while overly small chunks may separate important facts from their surrounding context.
Duplicate Results
The same source passage may appear in both dense and sparse results. Deduplicate candidates using stable document or chunk IDs before reranking, and consider diversity across source documents when assembling context.
Metadata Filtering and Access Control
Apply tenant, document-permission, language, and version filters consistently across retrieval paths. A hybrid pipeline must not expose documents simply because one retrieval branch failed to enforce access restrictions.
Latency and Cost
Hybrid retrieval adds work because multiple retrieval branches may run in parallel. Reranking adds another stage. Measure end-to-end latency, database load, embedding cost, and reranking throughput before choosing candidate counts.
Low-Confidence Queries
If the retrieved evidence is weak, conflicting, or incomplete, the generation layer should ask for clarification, qualify the response, or state that the available documents do not establish the answer.
9. Production Implementation Checklist
- Generate embeddings for document chunks and store them in a vector index.
- Build a lexical index using BM25 or a suitable full-text search implementation.
- Run dense and sparse retrieval in parallel where appropriate.
- Apply consistent access-control and metadata filters.
- Deduplicate and merge candidates using RRF.
- Add a cross-encoder reranker when its quality and latency trade-offs are justified.
- Preserve source metadata and document-level citations.
- Evaluate Recall@k, MRR, nDCG, latency, and answer faithfulness.
- Test exact-match queries separately from semantic and paraphrased queries.
- Monitor retrieval failures, stale indexes, and changes in document distribution.
- Verify that the final generated answer is supported by retrieved evidence.
Conclusion
Pure vector search is effective for semantic retrieval, but it can miss exact identifiers and other lexical signals that are critical in technical documentation. Hybrid search combines dense semantic retrieval with sparse keyword retrieval, while Reciprocal Rank Fusion merges their rankings without requiring directly comparable relevance scores.
For production RAG systems, the complete retrieval strategy should be evaluated as a pipeline: dense retrieval + sparse retrieval + rank fusion + optional reranking + evidence-grounded generation.
The objective is not simply to retrieve more documents. It is to retrieve the right evidence, rank it appropriately, and ensure that the generated answer remains faithful to the available sources.
Editorial Transparency & Verification Standards
Provenance, research methodology & primary citations
Staff-written distributed systems architectural explainer adhering to NexusBlog rigorous verification and reproducibility standards.
All architectural diagrams, code snippets, and distributed protocol assertions are technically reviewed prior to release.
Microservices Observability: OpenTelemetry, Metrics, Logs & Distributed Tracing
Dev
@krish
Core technical contributor to NexusBlog.
Discussion & Technical Notes0
Peer architectural reviews, benchmark insights, and implementation Q&A
Sign in to ask questions, share benchmark findings, or participate in architecture reviews.
Loading discussions...