Skip to main content

Providers and models

This page lists every embedding provider and text splitter provider: where each provider runs, the parameters with their defaults, the model names that each provider accepts, the images that carry models and packages, and the network access that each provider needs. Administrators use this page to configure embedding models and text splitters, and operators use this page to plan an air-gapped installation. Embedding models and text splitters explains both object types and shows the create requests.

Where providers run​

Providers run in one of two places:

  • API server and workers. FastEmbedEmbeddings and OpenAIEmbeddings run inside the Foundation4 process: the API server, each worker and the command-line interface (CLI). Each process loads a model on first use and keeps the model in memory.
  • gRPC service. The other embedding providers and every text splitter run in the gRPC service, a Python service that runs as a second container in each API server pod and each worker pod. The gRPC service keeps each loaded model and splitter in memory.

Provider names are case-sensitive. A request with any other provider name returns HTTP 400 with Unknown provider type: <provider> in details.error.

Embedding providers​

providerRuns inModel sourceNetwork access
FastEmbedEmbeddingsAPI server and workersONNX model files in the FastEmbed model directoryNone for models whose files are present. Other built-in models download on first use, unless a model image is mounted at the FastEmbed model directory
OpenAIEmbeddingsAPI server and workersA server that implements the OpenAI embeddings APIThe server named in endpoint
GPT4AllEmbeddingsgRPC serviceGGUF model files in the GPT4All model directoryThe GPT4All library can contact the GPT4All service while loading a model
HuggingFaceEmbeddingsgRPC serviceSentence Transformers model files in the Hugging Face model directoryThe Hugging Face library can contact the Hugging Face Hub while loading a model
HuggingFaceEndpointEmbeddingsgRPC serviceA Hugging Face inference endpointThe inference endpoint, on every request

The model directories are set by the paths configuration of the deployment. The defaults are /data/fastembed, /data/gpt4all and /data/huggingface.

Embedding model limits​

  • Dimensions. A pipeline stores vectors in a column with the dimension of the embedding model and builds hierarchical navigable small world (HNSW) indexes on the column. The indexes support up to 2,000 dimensions. A model with more dimensions can be created, but pipeline creation with that model fails.
  • Distance. Every model on this page is trained for cosine similarity, so searches use the default cosine distance strategy.
  • Text as given. Foundation4 embeds the text of queries and fragments exactly as given, with no instruction or prefix. Models whose documentation asks for prefixes, such as search_query: and search_document: for nomic-embed-text-v1.5, return vectors of lower quality without the prefixes.

FastEmbed​

FastEmbedEmbeddings runs ONNX models with the FastEmbed library. The model field selects either a built-in model or a custom model:

  • Built-in model. parameters is empty ({}), and model is one of the names in the following table.
  • Custom ONNX model. parameters is not empty, and model names a model directory, as described in Custom ONNX models.
modelDimensionsFiles in the API server image
Qdrant/all-MiniLM-L6-v2-onnx384Yes (seeded model)
BAAI/bge-small-en-v1.5384No
nomic-ai/nomic-embed-text-v1.5768Yes
BAAI/bge-small-zh-v1.5512No
Snowflake/snowflake-arctic-embed-xs384No
Xenova/all-MiniLM-L12-v2384No
jinaai/jina-embeddings-v2-base-en768No
mixedbread-ai/mxbai-embed-large-v11024No
intfloat/multilingual-e5-small384No
onnx-community/embeddinggemma-300m-ONNX768No

The API server image carries the files of the models marked Yes in /data/fastembed, and the workers run the same image. FastEmbed downloads the files of any other built-in model from Hugging Face the first time each process loads the model, into the container's own storage. A download fails in an air-gapped deployment, and the files are lost when the container restarts.

A model image mounted at /data/fastembed replaces the directory, so such an image includes every FastEmbed model that the deployment uses, including the seeded model. The mounted image is read-only, and FastEmbed does not download other models into it. With such an image, a built-in model whose files the image does not include fails to load, even when the deployment has internet access: the create request returns HTTP 400 with Invalid parameter, and details.error begins with Failed loading embedding model: Failed to retrieve.

A built-in model name that the table does not list returns HTTP 400 with Unknown FastEmbed embedding model: <name>. intfloat/multilingual-e5-small is the only multilingual built-in model.

Custom ONNX models​

A custom model is any ONNX sentence embedding model stored in the FastEmbed model directory in the Hugging Face cache layout. FastEmbed reads custom models from local files only and never downloads them.

ParameterTypeRequiredValues
onnxStringYesThe path of the ONNX file inside the model snapshot, such as onnx/model.onnx
poolingStringNocls or mean
quantizationStringNostatic or dynamic
data_filesArray of stringsNoAccepted and ignored: Foundation4 reads only the ONNX file, so a model that stores weights in separate data files cannot be loaded

Other parameters return HTTP 400. The model files are read from the following directory, where the model value has each / replaced by --:

<paths.fastembed>/models--<model>/snapshots/<snapshot>/

The snapshot directory holds the ONNX file and the files tokenizer.json, config.json, special_tokens_map.json and tokenizer_config.json. The model directory holds exactly one snapshot. A missing file returns HTTP 400 with Error reading <file> in details.error.

OpenAI-compatible embeddings​

OpenAIEmbeddings calls a server that implements the OpenAI embeddings API, such as an inference server inside the deployment's network.

ParameterTypeDefaultDescription
endpointStringhttps://api.openai.com/v1The base URL of the embeddings server. An air-gapped deployment sets an internal server.
api_keyStringEmptyThe key that the server requires. A server without authentication needs no key.

The model field is the model name that the server expects. The provider has no parameter to reduce the number of dimensions, so the model's native dimension applies. Other parameters return HTTP 400.

GPT4All​

GPT4AllEmbeddings runs GGUF embedding models with the GPT4All library. The model field is the name of a model file in the GPT4All model directory.

ArtifactContent
Package image gpt4allThe GPT4All Python package, mounted into the gRPC service
Model image gpt4allall-MiniLM-L6-v2.gguf2.f16.gguf (384 dimensions) and nomic-embed-text-v1.5.f16.gguf (768 dimensions)

The standard gRPC service image does not include the GPT4All package, so the provider requires the package image. Parameters are passed to the LangChain GPT4AllEmbeddings class, such as n_threads and device. Foundation4 sets the model name and the model path, so parameters cannot change them.

Hugging Face​

HuggingFaceEmbeddings runs Sentence Transformers models. The model field is the name of a model in the Hugging Face model directory, such as sentence-transformers/all-mpnet-base-v2.

ArtifactContent
Package image sentence-transformersThe Sentence Transformers package with PyTorch, mounted into the gRPC service
Model image huggingfacesentence-transformers/all-mpnet-base-v2 (768 dimensions)

Parameters are passed to the LangChain HuggingFaceEmbeddings class: model_kwargs (options for the model, such as device), encode_kwargs (options for encoding, such as normalize_embeddings), multi_process and show_progress. Foundation4 sets the model name and the cache directory, so parameters cannot change them.

HuggingFaceEndpointEmbeddings calls a Hugging Face inference endpoint. The parameter huggingfacehub_api_token is required, and a create request without the token returns HTTP 500 with API token is required for HuggingFaceEndpoint embeddings. in details.error. The model field names the model that the endpoint serves.

Text splitter providers​

Every text splitter runs in the gRPC service and wraps a text splitter of the LangChain text splitters library.

providerAvailabilityNetwork access
RecursiveCharacterTextSplitterAvailableNone
CharacterTextSplitterAvailableNone
CodeTextSplitterAvailableNone
MarkdownHeaderTextSplitterAvailable. The heading text is not kept in the fragments.None
TokenTextSplitterAvailable with network accessThe tiktoken library downloads the encoding files, which no image provides
NLTKTextSplitterRequires the nltk Python package, which no image providesNone, with the NLTK data from the nltk model image
RecursiveJsonSplitterNot available: creation fails with HTTP 500 in the current releaseNot applicable

Parameters are passed to the LangChain class as keyword arguments. A parameter that the class does not accept returns HTTP 400 when the splitter is created. A value that the class rejects, such as an invalid language or an overlap larger than the fragment size, returns HTTP 500 with Internal error, and details.error carries the message of the class.

Common parameters​

The character, recursive character, code, token and NLTK splitters accept the following parameters. The defaults are those of the LangChain text splitters library, version 1.1.1.

ParameterTypeDefaultDescription
chunk_sizeInteger4000The maximum size of a fragment, in characters, or in tokens for TokenTextSplitter
chunk_overlapInteger200The maximum size of the overlap between consecutive fragments. Must not exceed chunk_size.
keep_separatorBoolean, "start" or "end"true for the recursive splitters, false for the othersWhether the separator stays in the fragments, and at which end
strip_whitespaceBooleantrueWhether leading and trailing whitespace is removed from each fragment

Provider parameters​

providerParameters
RecursiveCharacterTextSplitterseparators: the list of separators, tried in order. The default is ["\n\n", "\n", " ", ""]. is_separator_regex: whether the separators are regular expressions.
CharacterTextSplitterseparator: the separator. The default is "\n\n". Text without the separator stays one fragment, whatever the size. is_separator_regex.
CodeTextSplitterlanguage (required): a LangChain language name, such as python, js, java, go, rust, markdown or html. An invalid language returns HTTP 500 with Unspecified or invalid language: <language> in details.error.
MarkdownHeaderTextSplitterheaders_to_split_on (required): the headings that start a new fragment, such as [["#", "Header 1"], ["##", "Header 2"]]. Fragment size follows the sections of the document, with no size limit.
TokenTextSplitterencoding_name: the tiktoken encoding. The default is gpt2. model_name: selects the encoding of a model instead.
NLTKTextSplitterseparator: the separator between sentences in a fragment. The default is "\n\n". language: the NLTK language. The default is english.

Seeded objects​

Every installation includes the following objects:

ObjectIdentifierProviderConfiguration
Embedding model all-MiniLM-L6-v29ba4a409-8773-415b-aa9d-ae3a2bdbc775FastEmbedEmbeddingsQdrant/all-MiniLM-L6-v2-onnx, 384 dimensions
Text splitter019252e9-b4a0-7713-9a69-d701b4f4a2d1RecursiveCharacterTextSplitterDefault parameters: 4,000 characters, 200 characters of overlap
Text splitter019252e9-da22-7f12-8c2f-36f4025fb0afCharacterTextSplitterDefault parameters
Text splitter019252e9-f2b1-7482-930f-0cf9400bdd79NLTKTextSplitterDefault parameters. Requires the nltk package.

Model and package images​

Models and Python packages that the standard images do not include are delivered as separate images, which the Helm charts mount into the containers:

ImageContentMounted at
API server imageThe Foundation4 binary and the FastEmbed models marked Yes in the FastEmbed table, in /data/fastembedNot applicable
gRPC service imageThe Python service and the LangChain librariesNot applicable
Package image gpt4allThe GPT4All package/packages/gpt4all in the gRPC service
Package image sentence-transformersSentence Transformers and PyTorch/packages/sentence-transformers in the gRPC service
Model image gpt4allTwo GPT4All embedding models/data/gpt4all
Model image huggingfacesentence-transformers/all-mpnet-base-v2/data/huggingface
Model image nltkThe NLTK sentence data punkt_tab/data/nltk

An air-gapped installation mirrors every image that the deployment uses. The CLI command foundation4ai data download-models downloads model files into the storage of the containers of the pod that runs the command, so the command does not prepare models for other pods or for an air-gapped installation. The worker container mounts the data subdirectory of each model image rather than the whole image, so the image layout decides which files the workers see. Models, packages and air-gapped installs describes the image configuration and a check of the files that each container sees.