Skip to main content

📝 Search & Retrieval

Description​

< What is it? >​

Search and retrieval finds relevant documents, passages, or items for a query. A search application returns these results to a user; a Retrieval Augmented Generation (RAG) system supplies retrieved content to a language model as context for an answer.

Retrieval selects candidates from a collection. Reranking optionally scores those candidates in more detail before selecting the final results. Sentence Transformers demonstrates this two-stage design.

  • Example: for “how do I reset my password?”, retrieve relevant help articles, rerank the candidates, and return the best matches or use them to ground a generated answer.

Key points​

< Indexing and querying >​

Prepare the searchable collection ahead of time, then reuse its index for incoming queries:

Indexing: documents → clean / split into passages → searchable index
Querying: query → retrieve candidates → optional reranking → top results

Keep document identifiers and source metadata with each indexed passage so results can link back to their sources. Query processing must be compatible with the index: lexical search uses consistent tokenization, and dense search uses compatible query and document embedding models.

< Keyword and semantic search >​

Keyword search, also called lexical search, matches terms. Term Frequency–Inverse Document Frequency (TF-IDF) assigns weights using term frequency within a document and rarity across the collection. BM25 is another lexical scoring method commonly used in search engines.

Dense semantic search represents queries and documents using learned embeddings and retrieves vectors with high similarity. These representations can capture relationships beyond shared words; Google's Machine Learning Crash Course introduces this motivation for embeddings.

  • Example: keyword matching can find an exact error code such as ERR_AUTH_42. A suitable semantic model can connect “forgot my password” with “account credential recovery” even when the wording differs.

< Hybrid search and reranking >​

Hybrid search combines lexical and semantic retrieval. A fusion method such as Reciprocal Rank Fusion (RRF) combines the ranked candidate lists without requiring their raw scores to share a scale. Elastic's hybrid search guide describes this approach.

After retrieval, a cross-encoder reranker can process each query–document pair jointly to estimate relevance. It operates on a limited candidate set because this detailed scoring is more expensive than initial retrieval. It cannot recover a document that was never retrieved. See Sentence Transformers' retrieve-and-rerank guide.

Comparison​

< Retrieval approaches >​

ApproachMatching signalUseful forMain limitation
Keyword / lexicalShared terms and their weightsNames, identifiers, exact terminologyWording differences can hide relevant results
Dense semanticSimilarity between learned embeddingsParaphrases and conceptual matchesQuality depends on the model and domain
HybridCombined lexical and semantic resultsQueries needing both exact and conceptual matchesRequires maintaining and combining two retrieval paths

Reranking is a later stage that can follow any of these approaches.

Implementation​

< A hybrid retrieval pipeline >​

This pseudocode shows the stages; the candidate counts are illustrative:

keyword_hits = lexical_index.search(query, top_k=100)
semantic_hits = vector_index.search(encode_query(query), top_k=100)
candidates = reciprocal_rank_fusion(keyword_hits, semantic_hits)
results = reranker.rank(query, candidates[:100])[:10]

Use stable document identifiers to merge duplicate results. For a small lexical baseline, start with the TF-IDF example before adding semantic retrieval or reranking.

Reference​