Skip to main content

Combine full-text and vector results

This guide combines the results of a similarity search and a full-text search for one query into one ranked list, which is how hybrid search works today. The client application sends the two searches and fuses the two result lists by rank position with reciprocal rank fusion (RRF). Integrators who build a search box that serves questions and exact terms alike need this guide. Choose and tune a search mode describes each search mode separately.

Fusion pattern​

Hybrid search in a client application follows four steps:

  1. Send a similarity search, or a maximal marginal relevance (MMR) search, with the query, the reader's classification, the filter and a candidate count k.
  2. Send a full-text search with the same query, classification, filter and k.
  3. Fuse the two result lists by the rank of each fragment in each list.
  4. Display the first results of the fused list.

The two searches are independent requests, so the client application sends the two requests in parallel, and the latency of hybrid search is the latency of the slower search. The pipeline must have been created with has_full_text_search set to true.

The fusion uses rank positions rather than scores. The score of a similarity search is a distance, where lower is better, and the score of a full-text search is a text search rank, where higher is better. The two scales are unrelated, so no formula converts one score into the other. Score semantics describes both scores.

Reciprocal rank fusion​

RRF gives each fragment a fused score from the fragment's rank in each list. A fragment at rank r in a list, counting from 1, contributes w / (60 + r), where w is the weight of the list. The contributions from the two lists are added, and the fused list is sorted by descending fused score. A fragment that is missing from a list receives no contribution from that list.

The constant 60 flattens the differences between neighboring ranks. With equal weights, a fragment within the first 60 ranks of both lists therefore outranks a fragment that appears first in only one list:

FragmentSimilarity rankFull-text rankFused score, weights 1 and 1
A1Not returned1/61 = 0.01639
B511/65 + 1/61 = 0.03178
C2Not returned1/62 = 0.01613

Fragment B, which contains the exact terms of the query and is also close in meaning, leads the fused list.

Similarity search step​

The worked example uses the support-kb pipeline and the shell variables from First search. Send a similarity search for 20 candidates and save the response in a file. The -w option prints the HTTP status code, because the response body goes to the file:

curl -sS -o similarity.json -w '%{http_code}\n' -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "annual plan refunds",
"type": "similarity",
"classification": "internal",
"params": {"k": 20}
}'

Expected result: the command prints 200, and similarity.json holds the three fragments of the pipeline, ordered by ascending score:

[
{
"id": "01a0ed0d-e007-7c3d-8e4f-6a7b8c9d0e1f",
"classification": "internal",
"metadata": {"product": "billing"},
"document_id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b",
"page_content": "Refunds for annual plans are prorated. Support agents must record the refund reason in the ticket before approving a refund.",
"version": 1790683504447128,
"expired_at": null,
"score": 0.2773,
"order_id": 0
},
{
"id": "01a0ed0d-d814-7d6c-8b5a-4c3d2e1f0a9b",
"classification": "internal",
"metadata": {"product": "billing"},
"document_id": "01a0ed0d-d75f-7d5e-8f60-718293a4b5c6",
"page_content": "Invoices are generated on the first day of each month. Finance staff can export invoices as PDF or CSV files from the Billing page.",
"version": 1790683502431906,
"expired_at": null,
"score": 0.8543,
"order_id": 0
},
{
"id": "01a0ed0d-d291-7c5d-8e4f-3a2b1c0d9e8f",
"classification": "public",
"metadata": {"product": "accounts"},
"document_id": "01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f",
"page_content": "To reset a password, open Settings, choose Security and select Reset password. A reset link is sent to the registered email address and expires after 30 minutes.",
"version": 1790683500418273,
"expired_at": null,
"score": 0.9152,
"order_id": 0
}
]

Full-text search step​

Send a full-text search with the same query, classification and k:

curl -sS -o fulltext.json -w '%{http_code}\n' -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"query": "annual plan refunds",
"type": "search",
"classification": "internal",
"params": {"k": 20}
}'

Expected result: the command prints 200, and fulltext.json holds one fragment, the only fragment that contains the words annual, plan and refunds after stemming:

[
{
"id": "01a0ed0d-e007-7c3d-8e4f-6a7b8c9d0e1f",
"classification": "internal",
"metadata": {"product": "billing"},
"document_id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b",
"page_content": "Refunds for annual plans are prorated. Support agents must record the refund reason in the ticket before approving a refund.",
"version": 1790683504447128,
"expired_at": null,
"score": 0.0936,
"order_id": 0
}
]

The two searches return the same fragment with the same id, because both modes read the same fragments of the pipeline. The fusion therefore identifies a fragment across the lists by id.

Fusion step​

The following Python script uses only the standard library. The script reads the two files, fuses the lists with RRF and prints the fused score, the document identifier and the start of the text of each fragment. Save the script as fuse.py in the directory of the two files:

import json

RRF_K = 60


def fuse(result_lists, weights=None, key="id", limit=10):
"""Fuse ranked result lists by weighted reciprocal rank fusion."""
weights = weights or [1.0] * len(result_lists)
fused = {}
for results, weight in zip(result_lists, weights):
seen = set()
for fragment in results:
item_key = fragment[key]
if item_key in seen:
continue
seen.add(item_key)
rank = len(seen)
entry = fused.setdefault(item_key, {"fragment": fragment, "rrf": 0.0})
entry["rrf"] += weight / (RRF_K + rank)
ranked = sorted(fused.values(), key=lambda entry: entry["rrf"], reverse=True)
return ranked[:limit]


with open("similarity.json", encoding="utf-8") as f:
similarity = json.load(f)
with open("fulltext.json", encoding="utf-8") as f:
fulltext = json.load(f)

for entry in fuse([similarity, fulltext]):
fragment = entry["fragment"]
text = fragment.get("page_content", "")[:48].rstrip()
print(f'{entry["rrf"]:.5f} {fragment["document_id"]} {text}')

Run the script:

python3 fuse.py

Expected result: the fused list, with the refund fragment first:

0.03279 01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b Refunds for annual plans are prorated. Support a
0.01613 01a0ed0d-d75f-7d5e-8f60-718293a4b5c6 Invoices are generated on the first day of each
0.01587 01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f To reset a password, open Settings, choose Secur

The refund fragment ranks first in both lists and receives 1/61 + 1/61. The other two fragments appear in the similarity list only. Python's sort keeps the order of insertion for equal fused scores, so ties follow the similarity ranking. The script processes the lists in the order of the response arrays, which are sorted best first in both modes, so the script never reads the score field.

Fusion key​

The key argument selects the unit of the fused list:

  • id. One entry for each fragment. A document with several matching fragments can appear several times.
  • document_id. One entry for each document. Each list keeps the first, best-ranked fragment of each document and skips the later fragments of the document, so a document's rank in a list is the rank of the document's best fragment.

In the worked example each document has one fragment, so both keys give the same fused list. A result list for readers usually shows one entry for each document, and retrieval for an LLM prompt usually keeps fragments.

Weights and thresholds​

Weights shift the balance between the two lists. fuse([similarity, fulltext], weights=[0.7, 0.3]) favors meaning over exact terms, and weights=[0.3, 0.7] favors exact terms. Equal weights, the default in the script, treat both lists alike. In the worked example the refund fragment leads both lists, so the order stays the same with either pair of weights, and only the gap between the refund fragment and the other two fragments changes: weights=[0.7, 0.3] gives the fused scores 0.01639, 0.01129 and 0.01111, and weights=[0.3, 0.7] gives 0.01639, 0.00484 and 0.00476.

A similarity threshold removes distant fragments from the similarity list before fusion. The full-text search needs no threshold in most cases, because every full-text result already contains every significant word of the query. A fragment that the similarity threshold removes can therefore still reach the fused list through the full-text list, which suits exact terms such as error codes that an embedding model represents poorly. For example, the similarity search for PDF CSV scores the invoice fragment, which contains both terms, at 0.5645, so a similarity threshold of 0.5 returns no fragment. The full-text search for PDF CSV returns the invoice fragment, and the fused list holds the invoice fragment alone, with the fused score 0.01639. Choose and tune a search mode describes how to choose a threshold.

Candidate pool​

Each search returns at most k fragments, and fusion can only reorder the fragments that the two searches return. Each search therefore requests more fragments than the client application displays: a pool of three times the displayed count, with a minimum of 20 and a maximum of 100, gives each list room to contribute. The search request has no offset, so a client application that pages through fused results repeats both searches with a larger k and fuses again.

Failure handling​

The client application treats the two searches as independent sources:

  • Full-text search unavailable. A pipeline created without full-text search returns HTTP 501 to the full-text request. The client application then displays the similarity list alone, or checks has_full_text_search on the pipeline once and skips the full-text request.
  • One search fails. When one request fails or exceeds a timeout, the client application displays the list of the other search and records the failure.
  • Empty full-text list. A question in natural language often matches no fragment in full-text search, because a fragment must contain every significant word of the query. The full-text search for the question How do customers get their money back? returns [], for example. The fused list then equals the similarity list, and no special handling is needed.
  • Short queries. Full-text search detects the language of each query, and a query of one or two words can be detected as a different language from the content. For example, the full-text search for Billing page returns [], although the invoice fragment contains both words. Similarity search still returns results for such a query, with the invoice fragment first for Billing page, so the fused list is not empty. Full-text query language describes the language detection of full-text search.

The API defines the search type hybrid, which returns HTTP 501 with the message Hybrid placeholder is not implemented yet in the current release. Feature status lists native hybrid search, with both searches and the re-ranking inside Foundation4, as planned.