Providers and models
This page lists every embedding provider and text splitter provider: where each provider runs, the parameters with their defaults, the model names that each provider accepts, the images that carry models and packages, and the network access that each provider needs. Administrators use this page to configure embedding models and text splitters, and operators use this page to plan an air-gapped installation. Embedding models and text splitters explains both object types and shows the create requests.
Where providers run
Providers run in one of two places:
- API server and workers.
FastEmbedEmbeddingsandOpenAIEmbeddingsrun inside the Foundation4 process: the API server, each worker and the command-line interface (CLI). Each process loads a model on first use and keeps the model in memory. - gRPC service. The other embedding providers and every text splitter run in the gRPC service, a Python service that runs as a second container in each API server pod and each worker pod. The gRPC service keeps each loaded model and splitter in memory.
Provider names are case-sensitive. A request with any other provider name returns HTTP 400 with Unknown provider type: <provider> in details.error.
Embedding providers
provider | Runs in | Model source | Network access |
|---|---|---|---|
FastEmbedEmbeddings | API server and workers | ONNX model files in the FastEmbed model directory | None for models whose files are present. Other built-in models download on first use, unless a model image is mounted at the FastEmbed model directory |
OpenAIEmbeddings | API server and workers | A server that implements the OpenAI embeddings API | The server named in endpoint |
GPT4AllEmbeddings | gRPC service | GGUF model files in the GPT4All model directory | The GPT4All library can contact the GPT4All service while loading a model |
HuggingFaceEmbeddings | gRPC service | Sentence Transformers model files in the Hugging Face model directory | The Hugging Face library can contact the Hugging Face Hub while loading a model |
HuggingFaceEndpointEmbeddings | gRPC service | A Hugging Face inference endpoint | The inference endpoint, on every request |
The model directories are set by the paths configuration of the deployment. The defaults are /data/fastembed, /data/gpt4all and /data/huggingface.
Embedding model limits
- Dimensions. A pipeline stores vectors in a column with the dimension of the embedding model and builds hierarchical navigable small world (HNSW) indexes on the column. The indexes support up to 2,000 dimensions. A model with more dimensions can be created, but pipeline creation with that model fails.
- Distance. Every model on this page is trained for cosine similarity, so searches use the default
cosinedistance strategy. - Text as given. Foundation4 embeds the text of queries and fragments exactly as given, with no instruction or prefix. Models whose documentation asks for prefixes, such as
search_query:andsearch_document:fornomic-embed-text-v1.5, return vectors of lower quality without the prefixes.
FastEmbed
FastEmbedEmbeddings runs ONNX models with the FastEmbed library. The model field selects either a built-in model or a custom model:
- Built-in model.
parametersis empty ({}), andmodelis one of the names in the following table. - Custom ONNX model.
parametersis not empty, andmodelnames a model directory, as described in Custom ONNX models.
model | Dimensions | Files in the API server image |
|---|---|---|
Qdrant/all-MiniLM-L6-v2-onnx | 384 | Yes (seeded model) |
BAAI/bge-small-en-v1.5 | 384 | No |
nomic-ai/nomic-embed-text-v1.5 | 768 | Yes |
BAAI/bge-small-zh-v1.5 | 512 | No |
Snowflake/snowflake-arctic-embed-xs | 384 | No |
Xenova/all-MiniLM-L12-v2 | 384 | No |
jinaai/jina-embeddings-v2-base-en | 768 | No |
mixedbread-ai/mxbai-embed-large-v1 | 1024 | No |
intfloat/multilingual-e5-small | 384 | No |
onnx-community/embeddinggemma-300m-ONNX | 768 | No |
The API server image carries the files of the models marked Yes in /data/fastembed, and the workers run the same image. FastEmbed downloads the files of any other built-in model from Hugging Face the first time each process loads the model, into the container's own storage. A download fails in an air-gapped deployment, and the files are lost when the container restarts.
A model image mounted at /data/fastembed replaces the directory, so such an image includes every FastEmbed model that the deployment uses, including the seeded model. The mounted image is read-only, and FastEmbed does not download other models into it. With such an image, a built-in model whose files the image does not include fails to load, even when the deployment has internet access: the create request returns HTTP 400 with Invalid parameter, and details.error begins with Failed loading embedding model: Failed to retrieve.
A built-in model name that the table does not list returns HTTP 400 with Unknown FastEmbed embedding model: <name>. intfloat/multilingual-e5-small is the only multilingual built-in model.
Custom ONNX models
A custom model is any ONNX sentence embedding model stored in the FastEmbed model directory in the Hugging Face cache layout. FastEmbed reads custom models from local files only and never downloads them.
| Parameter | Type | Required | Values |
|---|---|---|---|
onnx | String | Yes | The path of the ONNX file inside the model snapshot, such as onnx/model.onnx |
pooling | String | No | cls or mean |
quantization | String | No | static or dynamic |
data_files | Array of strings | No | Accepted and ignored: Foundation4 reads only the ONNX file, so a model that stores weights in separate data files cannot be loaded |
Other parameters return HTTP 400. The model files are read from the following directory, where the model value has each / replaced by --:
<paths.fastembed>/models--<model>/snapshots/<snapshot>/
The snapshot directory holds the ONNX file and the files tokenizer.json, config.json, special_tokens_map.json and tokenizer_config.json. The model directory holds exactly one snapshot. A missing file returns HTTP 400 with Error reading <file> in details.error.
OpenAI-compatible embeddings
OpenAIEmbeddings calls a server that implements the OpenAI embeddings API, such as an inference server inside the deployment's network.
| Parameter | Type | Default | Description |
|---|---|---|---|
endpoint | String | https://api.openai.com/v1 | The base URL of the embeddings server. An air-gapped deployment sets an internal server. |
api_key | String | Empty | The key that the server requires. A server without authentication needs no key. |
The model field is the model name that the server expects. The provider has no parameter to reduce the number of dimensions, so the model's native dimension applies. Other parameters return HTTP 400.
GPT4All
GPT4AllEmbeddings runs GGUF embedding models with the GPT4All library. The model field is the name of a model file in the GPT4All model directory.
| Artifact | Content |
|---|---|
Package image gpt4all | The GPT4All Python package, mounted into the gRPC service |
Model image gpt4all | all-MiniLM-L6-v2.gguf2.f16.gguf (384 dimensions) and nomic-embed-text-v1.5.f16.gguf (768 dimensions) |
The standard gRPC service image does not include the GPT4All package, so the provider requires the package image. Parameters are passed to the LangChain GPT4AllEmbeddings class, such as n_threads and device. Foundation4 sets the model name and the model path, so parameters cannot change them.
Hugging Face
HuggingFaceEmbeddings runs Sentence Transformers models. The model field is the name of a model in the Hugging Face model directory, such as sentence-transformers/all-mpnet-base-v2.
| Artifact | Content |
|---|---|
Package image sentence-transformers | The Sentence Transformers package with PyTorch, mounted into the gRPC service |
Model image huggingface | sentence-transformers/all-mpnet-base-v2 (768 dimensions) |
Parameters are passed to the LangChain HuggingFaceEmbeddings class: model_kwargs (options for the model, such as device), encode_kwargs (options for encoding, such as normalize_embeddings), multi_process and show_progress. Foundation4 sets the model name and the cache directory, so parameters cannot change them.
HuggingFaceEndpointEmbeddings calls a Hugging Face inference endpoint. The parameter huggingfacehub_api_token is required, and a create request without the token returns HTTP 500 with API token is required for HuggingFaceEndpoint embeddings. in details.error. The model field names the model that the endpoint serves.
Text splitter providers
Every text splitter runs in the gRPC service and wraps a text splitter of the LangChain text splitters library.
provider | Availability | Network access |
|---|---|---|
RecursiveCharacterTextSplitter | Available | None |
CharacterTextSplitter | Available | None |
CodeTextSplitter | Available | None |
MarkdownHeaderTextSplitter | Available. The heading text is not kept in the fragments. | None |
TokenTextSplitter | Available with network access | The tiktoken library downloads the encoding files, which no image provides |
NLTKTextSplitter | Requires the nltk Python package, which no image provides | None, with the NLTK data from the nltk model image |
RecursiveJsonSplitter | Not available: creation fails with HTTP 500 in the current release | Not applicable |
Parameters are passed to the LangChain class as keyword arguments. A parameter that the class does not accept returns HTTP 400 when the splitter is created. A value that the class rejects, such as an invalid language or an overlap larger than the fragment size, returns HTTP 500 with Internal error, and details.error carries the message of the class.
Common parameters
The character, recursive character, code, token and NLTK splitters accept the following parameters. The defaults are those of the LangChain text splitters library, version 1.1.1.
| Parameter | Type | Default | Description |
|---|---|---|---|
chunk_size | Integer | 4000 | The maximum size of a fragment, in characters, or in tokens for TokenTextSplitter |
chunk_overlap | Integer | 200 | The maximum size of the overlap between consecutive fragments. Must not exceed chunk_size. |
keep_separator | Boolean, "start" or "end" | true for the recursive splitters, false for the others | Whether the separator stays in the fragments, and at which end |
strip_whitespace | Boolean | true | Whether leading and trailing whitespace is removed from each fragment |
Provider parameters
provider | Parameters |
|---|---|
RecursiveCharacterTextSplitter | separators: the list of separators, tried in order. The default is ["\n\n", "\n", " ", ""]. is_separator_regex: whether the separators are regular expressions. |
CharacterTextSplitter | separator: the separator. The default is "\n\n". Text without the separator stays one fragment, whatever the size. is_separator_regex. |
CodeTextSplitter | language (required): a LangChain language name, such as python, js, java, go, rust, markdown or html. An invalid language returns HTTP 500 with Unspecified or invalid language: <language> in details.error. |
MarkdownHeaderTextSplitter | headers_to_split_on (required): the headings that start a new fragment, such as [["#", "Header 1"], ["##", "Header 2"]]. Fragment size follows the sections of the document, with no size limit. |
TokenTextSplitter | encoding_name: the tiktoken encoding. The default is gpt2. model_name: selects the encoding of a model instead. |
NLTKTextSplitter | separator: the separator between sentences in a fragment. The default is "\n\n". language: the NLTK language. The default is english. |
Seeded objects
Every installation includes the following objects:
| Object | Identifier | Provider | Configuration |
|---|---|---|---|
Embedding model all-MiniLM-L6-v2 | 9ba4a409-8773-415b-aa9d-ae3a2bdbc775 | FastEmbedEmbeddings | Qdrant/all-MiniLM-L6-v2-onnx, 384 dimensions |
| Text splitter | 019252e9-b4a0-7713-9a69-d701b4f4a2d1 | RecursiveCharacterTextSplitter | Default parameters: 4,000 characters, 200 characters of overlap |
| Text splitter | 019252e9-da22-7f12-8c2f-36f4025fb0af | CharacterTextSplitter | Default parameters |
| Text splitter | 019252e9-f2b1-7482-930f-0cf9400bdd79 | NLTKTextSplitter | Default parameters. Requires the nltk package. |
Model and package images
Models and Python packages that the standard images do not include are delivered as separate images, which the Helm charts mount into the containers:
| Image | Content | Mounted at |
|---|---|---|
| API server image | The Foundation4 binary and the FastEmbed models marked Yes in the FastEmbed table, in /data/fastembed | Not applicable |
| gRPC service image | The Python service and the LangChain libraries | Not applicable |
Package image gpt4all | The GPT4All package | /packages/gpt4all in the gRPC service |
Package image sentence-transformers | Sentence Transformers and PyTorch | /packages/sentence-transformers in the gRPC service |
Model image gpt4all | Two GPT4All embedding models | /data/gpt4all |
Model image huggingface | sentence-transformers/all-mpnet-base-v2 | /data/huggingface |
Model image nltk | The NLTK sentence data punkt_tab | /data/nltk |
An air-gapped installation mirrors every image that the deployment uses. The CLI command foundation4ai data download-models downloads model files into the storage of the containers of the pod that runs the command, so the command does not prepare models for other pods or for an air-gapped installation. The worker container mounts the data subdirectory of each model image rather than the whole image, so the image layout decides which files the workers see. Models, packages and air-gapped installs describes the image configuration and a check of the files that each container sees.