Full-text languages
This page lists the languages that full-text search supports, how Foundation4 determines the language of fragments and queries, how text is reduced to searchable terms, the query syntax and the ranking. Integrators use this page to predict which fragments a full-text search matches. Search and retrieval explains full-text search, and Search requests describes the request.
Supported languages
Full-text search uses the text search of PostgreSQL. Each language has a PostgreSQL text search configuration, which defines how words are reduced to terms:
| Language | Configuration |
|---|---|
| Arabic | arabic |
| Danish | danish |
| Dutch | dutch |
| English | english |
| Finnish | finnish |
| French | french |
| German | german |
| Hungarian | hungarian |
| Indonesian | indonesian |
| Irish | irish |
| Italian | italian |
| Portuguese | portuguese |
| Romanian | romanian |
| Russian | russian |
| Spanish | spanish |
| Swedish | swedish |
| Turkish | turkish |
| No language detected | simple |
The configurations for Arabic, Indonesian and Irish require PostgreSQL 12 or later. The PostgreSQL images that Foundation4 provides run PostgreSQL 18.
Language detection
Foundation4 detects languages with the Lingua library, restricted to the 17 languages in the table. The language models are part of the Foundation4 binary, so detection needs no network access and no model files.
- Fragments. The worker detects the language of each fragment when the worker processes a document on a pipeline with full-text search. The fragments of one document can therefore receive different configurations.
- Queries. The API server detects the language of the whole query text in each full-text search.
- Independence. The language of a query is detected without regard to the language of the fragments. A query detected as a different language from the content is stemmed with a different configuration, and fewer fragments match. For example, the queries
billingandBilling pagematch nothing in an English fragment that contains "the billing page", whilethe billing pagematches the fragment. Queries of several words, and queries that contain common words of the language, are detected more reliably than single words. - Shared stems. Languages whose stemmers reduce a word to the same stem match each other's fragments. The French
facturesand the Spanishfacturasboth reduce tofactur, so either query matches French and Spanish fragments that contain the word. - No override. No request, pipeline setting or configuration key sets the language, and Foundation4 does not store or return the detected language.
Other languages
Because detection chooses among the 17 supported languages only, text in another language is treated as follows:
- Languages written in the Latin, Cyrillic or Arabic script. Detection assigns the closest supported language, such as Russian for Ukrainian text or Danish for Norwegian text. The text is then stemmed with the configuration of the assigned language.
- Other scripts and text without letters. Text in scripts that no supported language uses, such as Greek, Hebrew or Devanagari, and text that consists only of numbers, codes or punctuation, uses the
simpleconfiguration, which converts words to lowercase without stemming. - Chinese, Japanese and Korean. PostgreSQL text search does not divide these scripts into words, so a run of characters without spaces forms one term. Full-text search suits languages that separate words with spaces. Similarity search with a multilingual embedding model serves content in these languages.
Terms
PostgreSQL reduces the text of each fragment and each query to terms with the selected configuration:
- Case. Text is converted to lowercase, so matching is case-insensitive.
- Stems. Each word is reduced to a stem, so a query for
connectionsalso matchesconnectedandconnectingin English text. Thesimpleconfiguration does not stem. - Stop words. Common words, such as
theandandin English, are removed from the text and from the query in the configurations that define a stop word list. - Accents. Accents are kept, so
caféandcafeare different terms unless the stemmer of the language treats them as one. - Tokens. Words, numbers, hyphenated words, email addresses, URLs, host names and file paths each form terms. Other punctuation separates words.
PostgreSQL limits the size of a term and of the terms of one fragment. A word longer than 2,047 bytes is skipped, and a fragment whose terms exceed 1 megabyte cannot be indexed, which affects only very large fragments.
Query syntax
The query of a full-text search is plain text:
- All terms required. A fragment matches when the fragment contains every term of the query, in any order and at any distance.
- No operators. Quotation marks,
OR,-,NOT,*and other operators have no special meaning, and the characters separate words. - Empty queries. A query that consists only of stop words or punctuation produces no terms and matches no fragment.
A query of several specific words therefore narrows the results, and a query that contains a word absent from the relevant fragments returns nothing. A client application that needs looser matching runs a similarity search as well, as described in Search and retrieval.
Ranking
The score of a full-text result is the PostgreSQL ts_rank value of the fragment for the query, divided by 1 plus the logarithm of the fragment length:
- Frequency. A fragment that contains the query terms more often ranks higher.
- Length. A longer fragment ranks lower for the same matches.
- No corpus statistics. The rank does not consider how common a term is across the pipeline, so a rare term and a common term count the same.
- Direction. A higher score is a better match. The
thresholdparameter sets a minimum score. - Scale. Scores are small numbers without a fixed upper limit. A threshold is chosen from the scores of sample searches on the pipeline's content.