Choose and tune a search mode
This guide selects a search mode for a retrieval task and tunes the parameters of the mode: the number of results, the maximal marginal relevance (MMR) parameters, score thresholds and the distance strategy. A worked sequence of searches on one pipeline compares similarity search, MMR search and full-text search. Integrators who design the retrieval of a client application need this guide. Search and retrieval explains how each mode ranks fragments, and Search requests lists every field of the request.
Search modes
Each search mode suits a different kind of query:
| Retrieval task | Search mode | Reason |
|---|---|---|
| Questions in natural language, paraphrases and synonyms | Similarity search | The embedding model places texts with similar meaning close together, even when the texts share no words |
| Content with many near-duplicate fragments, such as chat history or successive versions of a policy | MMR search | MMR search selects fragments that are relevant to the query and different from one another, so the results do not repeat one piece of information |
| Exact terms: product codes, error messages, names and acronyms | Full-text search | Full-text search matches the words of the query after stemming, independently of how the embedding model represents rare terms |
| Questions and exact terms in the same search box | Similarity search and full-text search, fused by the client application | Combine full-text and vector results describes the pattern |
The pipeline determines which modes are available:
- Embedding model. Similarity and MMR search require a pipeline with an embedding model. Each vector search converts the query into a vector, so the latency of a vector search includes one embedding call.
- Full-text index. Full-text search requires a pipeline created with
has_full_text_searchset totrue. The setting cannot be changed after the pipeline is created, and a full-text search on any other pipeline returns HTTP 501.
Example pipeline
The examples use the support-kb pipeline and the three documents from First search, with the shell variables from that tutorial. The pipeline uses the all-MiniLM-L6-v2 embedding model and has full-text search enabled. Each document forms a single fragment. Every search names the classification internal, so the fragments of all three documents are candidates.
Similarity search baseline
Run a similarity search for a question in natural language. The request omits params, so the default parameters apply and k is 5:
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "How do customers get their money back?",
"type": "similarity",
"classification": "internal"
}'
Expected result: status 200 and the three fragments of the pipeline, ordered by ascending score:
[
{
"id": "01a0ed0d-e007-7c3d-8e4f-6a7b8c9d0e1f",
"classification": "internal",
"metadata": {"product": "billing"},
"document_id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b",
"page_content": "Refunds for annual plans are prorated. Support agents must record the refund reason in the ticket before approving a refund.",
"version": 1790683504447128,
"expired_at": null,
"score": 0.6445,
"order_id": 0
},
{
"id": "01a0ed0d-d814-7d6c-8b5a-4c3d2e1f0a9b",
"classification": "internal",
"metadata": {"product": "billing"},
"document_id": "01a0ed0d-d75f-7d5e-8f60-718293a4b5c6",
"page_content": "Invoices are generated on the first day of each month. Finance staff can export invoices as PDF or CSV files from the Billing page.",
"version": 1790683502431906,
"expired_at": null,
"score": 0.8043,
"order_id": 0
},
{
"id": "01a0ed0d-d291-7c5d-8e4f-3a2b1c0d9e8f",
"classification": "public",
"metadata": {"product": "accounts"},
"document_id": "01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f",
"page_content": "To reset a password, open Settings, choose Security and select Reset password. A reset link is sent to the registered email address and expires after 30 minutes.",
"version": 1790683500418273,
"expired_at": null,
"score": 0.8393,
"order_id": 0
}
]
The refund fragment is the closest match, although the question shares no significant word with the fragment. The invoice and password fragments are returned as well, because similarity search returns the k closest fragments whether or not a fragment answers the question. In this mode, score is the cosine distance from the query, so a lower score is a closer match.
Full-text search with a question
Send the same question as a full-text search:
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "How do customers get their money back?",
"type": "search",
"classification": "internal"
}'
Expected result: status 200 and an empty array:
[]
A fragment matches a full-text search only when the fragment contains every significant word of the query after stemming. Common words such as how and their are removed from the query, but the remaining words, such as customers and money, appear in no fragment. Questions in natural language therefore suit similarity search.
Full-text search with exact terms
Send a full-text search with terms that the content contains verbatim:
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "refund reason ticket",
"type": "search",
"classification": "internal",
"params": {"k": 5}
}'
Expected result: status 200 and one fragment, from the kb-refund-policy document:
[
{
"id": "01a0ed0d-e007-7c3d-8e4f-6a7b8c9d0e1f",
"classification": "internal",
"metadata": {"product": "billing"},
"document_id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b",
"page_content": "Refunds for annual plans are prorated. Support agents must record the refund reason in the ticket before approving a refund.",
"version": 1790683504447128,
"expired_at": null,
"score": 0.1059,
"order_id": 0
}
]
The fragment contains all three words, so the fragment matches. The other fragments lack at least one word and are not returned, however many results k allows. In a full-text search, score is a text search rank: a higher score is a better match, and the values are not comparable with the distances of a similarity search.
Number of results
k sets the maximum number of fragments that a search returns, from 1 to 65535. A larger value returns HTTP 422. The value depends on the use of the results:
- Retrieval for an LLM. The number of fragments that the prompt can hold. The fragment size of the text splitter multiplied by
kstays within the part of the LLM context that the prompt reserves for retrieved text. - Result lists. The number of results that the client application displays. The search request has no offset, so a request for more results repeats the search with a larger
k. - Hybrid search. A candidate pool larger than the displayed list, because the client application fuses two result lists, as described in Combine full-text and vector results.
A request that includes params sets k explicitly, and for MMR search also fetch_k, because a parameter omitted inside params takes the value 0 and a search with k of 0 returns no fragments. A search returns fewer than k fragments when fewer fragments satisfy the classifications, the filter and the threshold. A vector search with a selective filter can also return fewer than k fragments, because the vector index is approximate.
Similarity threshold
A threshold excludes fragments that are too far from the query. Repeat the baseline search with a threshold of 0.7:
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "How do customers get their money back?",
"type": "similarity",
"classification": "internal",
"params": {"k": 5, "threshold": 0.7}
}'
Expected result: status 200 and only the fragments whose score is 0.7 or lower. With the scores of the baseline, the response contains the refund fragment alone. The following response omits the fields other than document_id and score:
[
{"document_id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b", "score": 0.6445}
]
For similarity and MMR search, threshold is the maximum distance from the query, and PostgreSQL applies the threshold before k. A threshold is chosen from the pipeline's own content:
- Run the search for a set of representative queries without a threshold and with a large
k, such as 20. - Record the scores of the fragments that answer each query and of the fragments that do not.
- Set the threshold between the two groups of scores, closer to the answering fragments when the client application must avoid irrelevant results.
The scores depend on the embedding model and on the distance strategy, so the threshold is chosen again after either changes. Under the cosine strategy, a minimum cosine similarity of s corresponds to a threshold of 1 minus s: a minimum similarity of 0.6 is a threshold of 0.4.
For full-text search, threshold is a minimum rank instead. Ranks are small numbers without a fixed maximum, and a longer fragment ranks lower for the same matches, so a full-text threshold is also chosen from sample searches and reviewed when the text splitter changes. Without a threshold, a full-text search returns up to k fragments that contain every significant word of the query. Full-text languages describes the rank.
MMR parameters
Run an MMR search that selects 2 fragments from up to 10 candidates:
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "How do customers get their money back?",
"type": "mmr",
"classification": "internal",
"params": {"k": 2, "fetch_k": 10, "lambda_mult": 0.5}
}'
Expected result: status 200 and 2 fragments, ordered by ascending score. The following response omits the fields other than document_id and score:
[
{"document_id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b", "score": 0.6445},
{"document_id": "01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f", "score": 0.8393}
]
MMR search first selects the refund fragment, the candidate closest to the query. For the second selection, MMR search weighs each remaining candidate's similarity to the query against the candidate's similarity to the refund fragment. In this response, the invoice fragment is closer to the query, but the invoice fragment and the refund fragment both describe billing, so MMR search selects the password fragment. The response lists the selected fragments in order of distance, and score is the distance from the query, not the MMR score.
The parameters are tuned as follows:
k. The number of fragments to return, chosen as for similarity search.fetch_k. The number of candidates, at leastk, because MMR search returns at mostfetch_kfragments. A larger pool gives MMR search more fragments to choose from. The API server receives the vector of every candidate and compares the candidates in memory, so the cost of the selection grows withfetch_kandk. A request withoutparamsuses akof 5 and afetch_kof 10.lambda_mult. The balance between relevance and diversity, from 0 to 1. A value of 1 selects by relevance alone and returns the same fragments as a similarity search over the candidates. A value of 0 maximizes diversity. The default is 0.5, and lower values suit content that repeats itself more.threshold. The maximum distance of a candidate. The threshold applies to the candidates before the selection.- Distance strategy.
cosineordot-product. MMR selection witheuclidean-distancedoes not rank candidates as intended.
Distance strategy
The distance strategy matches the similarity function for which the embedding model was trained, as stated in the documentation of the model. Most sentence embedding models, including all-MiniLM-L6-v2, are trained for cosine similarity, and cosine is the default. The pipeline stores no distance strategy, so a client application that uses another strategy sends the same distance_strategy value in every similarity and MMR request.
distance_strategy | Use | Score at a cosine similarity of 0.5, for vectors of unit length |
|---|---|---|
cosine (default) | Models trained for cosine similarity | 0.5 |
dot-product | Models trained for inner product similarity | -0.5 |
euclidean-distance | Models trained for Euclidean distance, in similarity search only | 1.0 |
For an embedding model that produces vectors of unit length, such as all-MiniLM-L6-v2, the three strategies return the same fragments in the same order, and only the scale of score differs. The baseline search scores the refund fragment 0.6445 under cosine, -0.3555 under dot-product and 1.1353 under euclidean-distance. A threshold chosen under one strategy therefore does not apply under another strategy. For a model whose vectors differ in length, a strategy other than the trained strategy changes the ranking. Search requests lists the score range of each strategy.
Full-text query language
Full-text search reduces the words of each fragment and of each query to stems with the rules of a language. The language affects which fragments match:
- Detection per query. The API server detects the language of each query independently of the language of the fragments. A query of one or two words can be detected as a different language from the content. The query words are then stemmed with other rules than the indexed words, and fewer fragments match, or none. On the example pipeline, the query
Billing pagematches no fragment, although the invoice fragment contains both words, whilePDF CSVandinvoicesmatch the invoice fragment. - Query length. Queries of several words are detected more reliably. Each added word must also appear in a matching fragment, so a query of a few specific words balances reliable detection against the all-words rule.
- Common words. A query that consists only of common words, such as
theandandin English, produces no search terms and matches no fragment. - No language setting. No request field or pipeline setting selects the language.
A client application that sends one-word queries, such as product codes, runs a similarity search as well and combines the results, as described in Combine full-text and vector results. Full-text languages describes detection, the supported languages and the treatment of other languages.