Search and retrieval
Foundation4 retrieves fragments from a pipeline in three search modes: similarity search, maximal marginal relevance (MMR) search and full-text search. This page describes the search request, how each mode selects and ranks fragments, what the score of a result means in each mode, and the scope that applies to every search. Integrators need this page to choose a search mode and to interpret search results.
Search request
All three modes use one endpoint, POST /pipelines/{id}/search. The API key requires read and execute permission on the pipeline. The request body contains the following fields.
| Field | Required | Description |
|---|---|---|
query | Yes | The search text, such as the end user's question or keywords |
type | Yes | The search mode: similarity, mmr or search (full-text search) |
classification | Yes | The classifications that the reader holds. The Classifications page describes the accepted forms. |
params | No | The parameters of the search mode, described in each mode's section |
filters | No | A metadata filter. The Metadata, filters and taxonomies page describes the filter language. |
distance_strategy | No | The vector distance function for similarity and MMR search. The default is cosine. |
The following request runs a similarity search for the five closest fragments:
{
"query": "How do I rotate the database credentials?",
"type": "similarity",
"classification": ["internal"],
"params": {"k": 5}
}
The response is an array of fragments, ordered from the best match to the weakest match. Each fragment carries the decrypted text in page_content and the match score in score. The Documents, versions and fragments page describes the other fragment fields.
The default parameter values apply when the request omits the params object. When the request includes params, the request sets k explicitly, and for MMR search also fetch_k, because a field omitted inside params takes the value 0 and a search with k of 0 returns no results. Search requests lists every field, parameter and error of the search request.
Search path
Rectangles are processing steps in the API server. Cylinders are indexes in PostgreSQL.
Similarity and MMR search first convert the query into a vector with the pipeline's embedding model, so the latency of a vector search includes one embedding call. The API server then resolves the classifications named in the request, including inherited classifications when matching is hierarchical. Full-text search resolves the classifications first and then detects the language of the query instead of embedding the query. Because a vector search embeds the query first, an invalid embedding parameter returns HTTP 400 before any classification is checked. PostgreSQL then selects the matching fragments from the index, restricted to the resolved classifications, the metadata filter and current versions, in a single query. The API server decrypts the text of the selected fragments and returns the fragments with their scores.
Similarity search
Similarity search returns the k fragments whose vectors are closest to the query vector. Similarity search suits questions phrased in natural language, because the embedding model places texts with similar meaning close together even when the texts share no words.
| Parameter | Default | Description |
|---|---|---|
k | 5 | The number of fragments to return. A value of 0 returns no fragments. |
threshold | None | The maximum distance. Fragments farther from the query than the threshold are excluded. |
Similarity search requires a pipeline with an embedding model. The vector index is a hierarchical navigable small world (HNSW) index, which performs approximate nearest-neighbor search: the index trades a small loss of recall for query times that remain short as the pipeline grows.
MMR search
MMR search returns k fragments that are relevant to the query and different from one another. MMR search suits content with many near-duplicate fragments, such as chat history or versioned documents, where the closest fragments often repeat the same information.
MMR search runs in two steps. The API server retrieves the fetch_k closest fragments, as in similarity search, and then selects k of those candidates one at a time. Each selection balances the candidate's similarity to the query against the candidate's similarity to the fragments already selected.
| Parameter | Default | Description |
|---|---|---|
k | 5 | The number of fragments to return. A value of 0 returns no fragments. |
fetch_k | 10 | The number of candidates. Set to at least k, because MMR search returns at most fetch_k fragments. |
lambda_mult | 0.5 | The balance between relevance and diversity, from 0 to 1. A value of 1 selects by relevance alone. A value of 0 maximizes diversity. Values outside 0 to 1 are not rejected. |
threshold | None | The maximum distance of a candidate from the query |
MMR search requires a pipeline with an embedding model. Use the cosine or dot-product distance strategy for MMR search: the API accepts euclidean-distance, but MMR selection with that strategy does not rank candidates as intended. The selected fragments are returned in order of distance from the query.
Distance strategies
The distance strategy determines how similarity and MMR search measure the distance between the query vector and a fragment vector. The strategy is set per request in distance_strategy. Each vector table carries an HNSW index for each supported strategy, so every strategy uses an index.
| Value | Distance | Score range |
|---|---|---|
cosine (default) | Cosine distance: 1 minus the cosine similarity | 0 (identical direction) to 2 |
euclidean-distance | Euclidean (L2) distance | 0 (identical vectors) upward |
dot-product | Negative inner product | Negative values for similar vectors. Lower is closer. |
The value max-inner-product is not supported. The distance strategy should match the strategy for which the embedding model was trained. Most sentence embedding models, including the default FastEmbed model, are trained for cosine similarity.
Full-text search
Full-text search returns the k fragments that contain every word of the query, ranked by relevance. Full-text search suits exact terms that an embedding model represents poorly: product codes, error messages, names and acronyms. Full-text search requires a pipeline created with full-text search enabled. A request to any other pipeline returns HTTP 501.
| Parameter | Default | Description |
|---|---|---|
k | 5 | The number of fragments to return. A value of 0 returns no fragments. |
threshold | None | The minimum rank. Fragments ranked below the threshold are excluded. |
Full-text search uses PostgreSQL text search. Language handling determines which fragments match:
- Language detection at ingestion. The worker detects the language of each fragment and builds the fragment's index entry with the PostgreSQL text search configuration for that language. The configuration reduces each word to a stem and removes common words, so the query
approve refundsmatches a fragment that contains "approving a refund". - Language detection at query time. The API server detects the language of the query independently and parses the query with that language's configuration.
- Supported languages. Arabic, Danish, Dutch, English, Finnish, French, German, Hungarian, Indonesian, Irish, Italian, Portuguese, Romanian, Russian, Spanish, Swedish and Turkish. Text in another language receives the configuration of the closest supported language, except text in scripts that no supported language uses, which receives the
simpleconfiguration: words are lowercased without stemming. Full-text languages describes detection in detail. - Word matching. Every remaining word of the query must appear in the fragment. The query syntax supports no phrase, prefix or Boolean operators.
A short query can be detected as a different language from the content, in which case the query words are stemmed differently from the indexed words and fewer fragments match. Queries of several words are detected more reliably than single words.
Score semantics
The meaning of score differs by mode, and scores from different modes cannot be compared.
| Mode | Score | Better match |
|---|---|---|
| Similarity | Distance from the query, under the request's distance strategy | Lower |
| MMR | Distance from the query, under the request's distance strategy | Lower |
| Full-text | PostgreSQL ts_rank, divided by 1 plus the logarithm of the fragment length | Higher |
A client application that combines vector results with full-text results therefore merges the two lists by rank position rather than by score. The Combine full-text and vector results guide describes this pattern, which is how hybrid search works today. The Feature status page lists the status of native hybrid search.
The API also defines the type values hybrid and history. Both return HTTP 501 in the current release.
Result scope
The following restrictions apply to every search mode:
- Classifications. Search returns only fragments that carry a classification resolved from the request.
- Metadata filter. Search returns only fragments whose metadata satisfies the filter.
- Current content. Search excludes expired versions. A new version becomes searchable when processing of that version completes.
Search always operates on current content. Point-in-time reads are available on the document and fragment endpoints, as described in Documents, versions and fragments.