AI / EngineeringStaff Architecture ExplainerADVANCEDPart 10 of 10 in Series

Building Production-Grade RAG Pipelines with Vector Embeddings and Hybrid Search

Architecting high-accuracy Retrieval-Augmented Generation systems using chunking strategies, dense vector embeddings, BM25 sparse keyword ranking, and reciprocal rank fusion.

D
Dev
@krish
October 01, 2026 15 min read
0
Part 10 of 10Technical Learning Track

System Design: From Zero to Production

A comprehensive engineering series guiding backend developers from single-node instances to highly resilient distributed architectures.

Building Production-Grade RAG Pipelines with Vector Embeddings and Hybrid Search

Why Pure Vector Search Fails in Production

Semantic vector search is powerful for retrieving documents based on meaning, context, and conceptual similarity. However, in production environments, semantic similarity alone is not sufficient for reliable information retrieval.

Consider a technical support query:

Why does error P2025 occur when using Prisma with PostgreSQL?

A vector search engine may retrieve documents about Prisma database errors in general, even when a document explicitly explaining P2025 exists but receives a lower semantic similarity score. Similar problems occur with API endpoints, product SKUs, version numbers, configuration keys, exception messages, and other exact-match identifiers.

Hybrid search addresses this limitation by combining dense vector retrieval with sparse keyword retrieval. Dense retrieval captures semantic relationships, while sparse retrieval identifies documents containing relevant keywords and technical identifiers.

When combined with Reciprocal Rank Fusion (RRF) and an optional reranking stage, hybrid retrieval can improve the relevance of the context supplied to a Retrieval-Augmented Generation (RAG) pipeline.

Key takeaway: Vector search helps find documents that mean the same thing. Keyword search helps find documents that contain the exact terms. Hybrid search combines both signals.

1. Why Pure Vector Search Is Not Enough

Vector search converts queries and documents into numerical embeddings. A similarity function, such as cosine similarity, is then used to retrieve documents whose embeddings are close to the query embedding.

This works well for natural-language questions and conceptually related content, but several limitations become apparent in production.

1.1 Exact-Match Retrieval Problems

Embedding models do not guarantee that rare identifiers or exact technical strings receive sufficient weight in the similarity calculation.

For example, these identifiers may have very different operational meanings:

  • ERR_CONNECTION_RESET
  • ERR_CONNECTION_REFUSED
  • HTTP 401
  • HTTP 403
  • PostgreSQL 16
  • PostgreSQL 17

A semantically similar document is not necessarily the document that contains the exact error code, version, or configuration value the user needs.

1.2 Semantic Drift

A query about a specific implementation detail may retrieve conceptually related but operationally irrelevant documents.

For example, a search for a Prisma connection-pool timeout could retrieve general database performance guidance instead of the configuration documentation that explains the actual timeout parameter.

1.3 Rare Terms and Technical Identifiers

Technical documentation frequently contains tokens that appear infrequently in a corpus:

  • API names and endpoint paths
  • Exception messages and error codes
  • Package names and dependency versions
  • Database table and column names
  • Environment variables
  • Configuration keys and command-line flags

Sparse retrieval methods such as BM25 are particularly useful when these exact terms determine relevance.

1.4 Similarity Scores Are Not Relevance Guarantees

A high vector similarity score indicates that two embeddings are close under the chosen similarity metric. It does not guarantee that the retrieved document contains the exact evidence required to answer the question.

This distinction is important in production RAG systems, where irrelevant context can reduce answer quality and increase the risk of unsupported responses.

Hybrid search combines two complementary retrieval approaches:

Retrieval methodHow it worksPrimary strength
Dense vector searchCompares query and document embeddingsSemantic similarity and paraphrases
Sparse keyword searchMatches query terms against indexed document termsExact keywords, identifiers, and rare terms
Hybrid retrievalCombines rankings from both methodsBroader retrieval coverage

Dense Retrieval

A dense embedding model converts text into a fixed-dimensional vector. For example, text-embedding-3-small supports 1,536-dimensional embeddings by default.

A vector database can use an approximate nearest-neighbor index, such as HNSW, to retrieve candidate documents efficiently.

Dense retrieval is useful when the query and document use different wording but express similar concepts.

Sparse Retrieval with BM25

BM25 is a probabilistic information-retrieval ranking function that scores documents using factors such as term frequency, inverse document frequency, and document length normalization.

It is useful when a query contains specific terms that should match the indexed content.

Implementation note: PostgreSQL provides built-in full-text search through tsvector, tsquery, and ranking functions such as ts_rank_cd. These provide lexical retrieval, but they are not automatically equivalent to a full BM25 implementation. If BM25 is a requirement, use an implementation that explicitly supports it, such as a suitable search engine or a PostgreSQL extension.

3. Hybrid RAG Architecture

A hybrid RAG pipeline retrieves candidate documents using both dense and sparse search, merges their rankings, optionally reranks the candidates, and passes the most relevant evidence to the language model.

Architecture Diagram
Mermaid Flow

Retrieval Pipeline Explained

  1. Query processing: Normalize the query where appropriate and extract important technical terms, identifiers, and filters.
  2. Dense retrieval: Generate an embedding for the query and retrieve semantically similar document chunks.
  3. Sparse retrieval: Search the lexical index for relevant keywords and identifiers using BM25 or another suitable ranking method.
  4. Candidate fusion: Merge the ranked lists using Reciprocal Rank Fusion.
  5. Reranking: Optionally use a cross-encoder to evaluate query-document relevance more directly.
  6. Context assembly: Select relevant, sufficiently diverse chunks while respecting the model's context budget.
  7. Generation: Provide the selected evidence to the language model and require it to distinguish supported facts from missing information.

The candidate counts and final context size should be tuned to the corpus, latency budget, and evaluation results rather than treated as universal defaults.

4. Reciprocal Rank Fusion (RRF)

Reciprocal Rank Fusion combines ranked search results without requiring the scores from different retrieval systems to be directly comparable.

This is useful because vector similarity scores and BM25 scores generally have different scales and distributions. Adding them directly can cause one retrieval system to dominate unless their scores are calibrated or normalized.

RRF instead assigns a contribution based on each document's rank in each result list.

RRF Formula

For a document (d), its fused score is:

RRF⁡(d)=∑r∈R1k+rank⁡r(d)\operatorname{RRF}(d) = \sum_{r \in R} \frac{1}{k+\operatorname{rank}_r(d)}

Where:

  • (R) is the set of ranked retrieval result lists.
  • (\operatorname{rank}_r(d)) is the one-based rank of document (d) in result list (r).
  • (k) is a smoothing constant, commonly set to 60.
  • Documents absent from a result list contribute zero for that list.

A document ranked highly by both dense and sparse retrieval receives contributions from both lists, increasing its fused score.

Example

Assume a document appears in both retrieval lists:

Retrieval sourceDocument rankRRF contribution with (k=60)
Dense vector search1(1/61 \approx 0.01639)
Sparse keyword search3(1/63 \approx 0.01587)
Combined score—0.03226

A document appearing only in the first position of the dense list receives a score of approximately (0.01639). The document appearing in both lists receives a higher combined score because both retrieval systems support its relevance.

RRF does not prove that a document is relevant. It combines ranking evidence and should be evaluated against representative queries and labeled relevance judgments.

5. TypeScript Implementation of RRF

The following implementation accepts already-ranked dense and sparse result lists. Each list must be ordered from most relevant to least relevant.

typescript
export type RankedDocument = {
  id: string;
  score: number;
};

export type FusedDocument = {
  id: string;
  rrfScore: number;
};

export function reciprocalRankFusion(
  denseResults: RankedDocument[],
  sparseResults: RankedDocument[],
  k = 60,
): FusedDocument[] {
  if (!Number.isFinite(k) || k <= 0) {
    throw new Error("k must be a positive finite number");
  }

  const scoreMap = new Map<string, number>();

  const addRankedResults = (results: RankedDocument[]) => {
    results.forEach((doc, index) => {
      const rank = index + 1;

      scoreMap.set(
        doc.id,
        (scoreMap.get(doc.id) ?? 0) + 1 / (k + rank),
      );
    });
  };

  addRankedResults(denseResults);
  addRankedResults(sparseResults);

  return Array.from(scoreMap, ([id, rrfScore]) => ({
    id,
    rrfScore,
  })).sort((a, b) => b.rrfScore - a.rrfScore);
}

Implementation Considerations

  • Ranks must be one-based: The first result receives rank 1, not rank 0.
  • Input ordering matters: RRF uses the position of each result, not its supplied score field.
  • Scores are intentionally unused: The score field is retained for compatibility and debugging, but RRF uses rankings only.
  • Duplicate document IDs: Repeated IDs within a single result list contribute multiple times in this simple implementation. Deduplicate each list before fusion if duplicate entries are possible.
  • Stable tie-breaking: If deterministic output is required, add an explicit secondary sort key.
  • Metadata preservation: In a production implementation, maintain a document lookup map so fused results can be enriched with document text, source URLs, access permissions, and metadata.

For systems with more than two retrieval sources, apply the same rank-based accumulation to each list.

6. Reranking After Hybrid Retrieval

RRF is effective for merging candidate lists, but it does not deeply evaluate the semantic relationship between the original query and each candidate document.

A cross-encoder reranker evaluates a query-document pair together and produces a relevance score. Because it performs more computation per candidate, it is generally applied after the initial retrieval stage.

A typical pipeline is:

  1. Retrieve 20–100 candidates from each search system, depending on corpus size and latency requirements.
  2. Merge and deduplicate the candidates with RRF.
  3. Select a manageable candidate pool for reranking.
  4. Rerank the candidates using a cross-encoder.
  5. Select the final context chunks based on relevance, diversity, permissions, and context limits.

Reranking can improve the ordering of retrieved candidates, but it cannot recover a relevant document that neither retrieval system returned.

Hybrid retrieval should be evaluated against a representative test set rather than assumed to be better in every workload.

For a technical documentation corpus, build a labeled dataset containing natural-language questions, exact-identifier queries, version-specific questions, and paraphrased queries. Compare the systems using identical query sets and relevance judgments.

MetricWhat it measures
Recall@5Fraction of relevant documents retrieved within the first five results
Recall@20Fraction of relevant documents retrieved within the first twenty results
MRRHow highly the first relevant result is ranked
nDCG@kRanking quality when relevance has multiple grades
Retrieval latencyTime spent retrieving and ranking candidate documents
Answer faithfulnessWhether generated claims are supported by retrieved evidence
Citation accuracyWhether cited sources actually support the associated claims

Interpreting Benchmark Results

A benchmark might report the following hypothetical results:

MetricPure vector searchHybrid search
Recall@564.2%89.4%
Recall@20Measure experimentallyMeasure experimentally
Retrieval latencyMeasure experimentallyMeasure experimentally
Answer faithfulnessMeasure experimentallyMeasure experimentally

Important: The Recall@5 figures above are illustrative values, not independently verified benchmark results. Replace them with measured results from your own evaluation dataset before publishing them as empirical findings.

A claim such as an 82% reduction in hallucination rate requires a separate controlled evaluation of generated answers. Retrieval recall alone cannot establish that result.

8. Common Production Challenges

Query and Document Preprocessing

Use consistent normalization where appropriate, but preserve meaningful distinctions in identifiers, versions, and case-sensitive strings. Aggressive normalization can make technically different identifiers appear equivalent.

Chunking Strategy

Chunk size affects both retrieval precision and the amount of evidence available to the language model. Overly large chunks can introduce irrelevant content, while overly small chunks may separate important facts from their surrounding context.

Duplicate Results

The same source passage may appear in both dense and sparse results. Deduplicate candidates using stable document or chunk IDs before reranking, and consider diversity across source documents when assembling context.

Metadata Filtering and Access Control

Apply tenant, document-permission, language, and version filters consistently across retrieval paths. A hybrid pipeline must not expose documents simply because one retrieval branch failed to enforce access restrictions.

Latency and Cost

Hybrid retrieval adds work because multiple retrieval branches may run in parallel. Reranking adds another stage. Measure end-to-end latency, database load, embedding cost, and reranking throughput before choosing candidate counts.

Low-Confidence Queries

If the retrieved evidence is weak, conflicting, or incomplete, the generation layer should ask for clarification, qualify the response, or state that the available documents do not establish the answer.

9. Production Implementation Checklist

  • Generate embeddings for document chunks and store them in a vector index.
  • Build a lexical index using BM25 or a suitable full-text search implementation.
  • Run dense and sparse retrieval in parallel where appropriate.
  • Apply consistent access-control and metadata filters.
  • Deduplicate and merge candidates using RRF.
  • Add a cross-encoder reranker when its quality and latency trade-offs are justified.
  • Preserve source metadata and document-level citations.
  • Evaluate Recall@k, MRR, nDCG, latency, and answer faithfulness.
  • Test exact-match queries separately from semantic and paraphrased queries.
  • Monitor retrieval failures, stale indexes, and changes in document distribution.
  • Verify that the final generated answer is supported by retrieved evidence.

Conclusion

Pure vector search is effective for semantic retrieval, but it can miss exact identifiers and other lexical signals that are critical in technical documentation. Hybrid search combines dense semantic retrieval with sparse keyword retrieval, while Reciprocal Rank Fusion merges their rankings without requiring directly comparable relevance scores.

For production RAG systems, the complete retrieval strategy should be evaluated as a pipeline: dense retrieval + sparse retrieval + rank fusion + optional reranking + evidence-grounded generation.

The objective is not simply to retrieve more documents. It is to retrieve the right evidence, rank it appropriately, and ensure that the generated answer remains faithful to the available sources.

Editorial Transparency & Verification Standards

Provenance, research methodology & primary citations

Staff Architecture Explainer
Research Methodology

Staff-written distributed systems architectural explainer adhering to NexusBlog rigorous verification and reproducibility standards.

Technical Peer Review

All architectural diagrams, code snippets, and distributed protocol assertions are technically reviewed prior to release.

Spotted a technical inaccuracy or outdated code sample?
0
Track Roadmap: System Design: From Zero to Production
All Chapters (10)
D

Dev

@krish

Core technical contributor to NexusBlog.

Discussion & Technical Notes0

Peer architectural reviews, benchmark insights, and implementation Q&A

Join the Technical Discussion

Sign in to ask questions, share benchmark findings, or participate in architecture reviews.

Loading discussions...

Production-Grade RAG Pipelines with Hybrid Search | NexusNation | NexusBlog