Skip to main content

Embedding models and text splitters

Foundation4 divides every document into fragments with a text splitter. In pipelines with vector search, an embedding model then converts each fragment into a vector. This page describes the providers behind both object types, how embedding models and text splitters are configured, where each provider runs and what each provider needs in an air-gapped deployment. Administrators need this page to choose and configure models and splitters before creating pipelines.

Providers and configured objects​

A provider is an engine that Foundation4 supports, such as FastEmbed for embeddings or the LangChain recursive character splitter for splitting. An embedding model or a text splitter is a configured instance of a provider, with a model name and parameters. A deployment can hold several configured instances of the same provider, for example two text splitters with different fragment sizes.

GET /providers/embeddings and GET /providers/text-splitters list the providers of a deployment. GET /embedding-models and GET /text-splitters list the configured objects.

Creating an embedding model or a text splitter requires write permission on the object type and on its provider type: embedding-models and embedding-providers, or text-splitters and text-splitter-providers. A key that holds only read permission on the provider type receives HTTP 404 with Embedding Provider not found or Text Splitter not found, although the provider exists. The getting-started key of the tutorials holds read permission only on both provider types, so the create requests on this page need another key, such as one with 6 (read and write) on the object type and on its provider type.

Embedding providers​

Foundation4 supports five embedding providers. Each runs either inside the API server and worker processes or in the gRPC service, which runs as a sidecar container: a second container in each API server pod and each worker pod.

ProviderRuns inModel source
FastEmbedEmbeddingsAPI server and workersOpen Neural Network Exchange (ONNX) model files in the FastEmbed model directory
OpenAIEmbeddingsAPI server and workersAny server that implements the OpenAI embeddings API, hosted or on-premises
HuggingFaceEmbeddingsgRPC serviceSentence Transformers model files in the Hugging Face model directory
GPT4AllEmbeddingsgRPC serviceGPT4All model files in the GPT4All model directory
HuggingFaceEndpointEmbeddingsgRPC serviceA Hugging Face inference endpoint

The seeded embedding model uses FastEmbedEmbeddings, and the model files of the seeded model are included in the API server image, which the workers also run. The other local providers need model files and, for HuggingFaceEmbeddings and GPT4AllEmbeddings, Python packages that are delivered as separate images. Models, packages and air-gapped installs describes how to add model and package images.

Embedding models​

An embedding model is created with POST /embedding-models. The request contains the following fields.

FieldRequiredDescription
nameYesA name that is unique across the deployment
descriptionNoFree text for administrators
providerYesOne of the five provider names
modelYes, in practiceThe model name that the provider loads, such as Qdrant/all-MiniLM-L6-v2-onnx for FastEmbed
sizeYesThe number of dimensions of the vectors. The value 0 instructs Foundation4 to measure the size.
parametersNoProvider-specific settings, described below

When the model is created, Foundation4 embeds a test string with the provider. Creation therefore fails when the provider cannot load the model, and the stored size is always the measured size: a size of 0 stores the measured size, and a nonzero size that differs from the measured size returns HTTP 400. Sending 0 is the reliable choice.

The shell variables in the examples are described in Authenticate. The following request creates a FastEmbed model from the model files included in the API server image:

curl -X POST "$FOUNDATION4_URL/embedding-models" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"name": "minilm-l6",
"description": "FastEmbed all-MiniLM-L6-v2",
"provider": "FastEmbedEmbeddings",
"model": "Qdrant/all-MiniLM-L6-v2-onnx",
"size": 0
}'

The response has status 201 and contains the stored model, with size set to 384.

The size determines the vector column of every pipeline that uses the model. The hierarchical navigable small world (HNSW) indexes that pgvector builds for vector search support up to 2,000 dimensions, so a model with more dimensions cannot serve a pipeline.

Only the name and description of an embedding model can change after creation. A change request that also sends other fields, such as model or size, returns HTTP 200 and ignores those fields. A model that a pipeline uses cannot be deleted; the delete request returns HTTP 409 until every pipeline that uses the model is deleted. A pipeline uses one embedding model for the lifetime of the pipeline, as described in Pipelines.

POST /embedding-models/{id}/query embeds a text and returns the vector as an array of numbers. The endpoint serves to confirm that a model works and to compare the vectors of two texts.

Provider parameters​

The parameters object configures a provider beyond the model name.

  • FastEmbedEmbeddings. Empty parameters select a built-in FastEmbed model by name. A custom ONNX model is configured with onnx, the path of the ONNX file within the model directory, and optionally pooling (cls or mean) and quantization (static or dynamic). FastEmbed reads custom models only from local files, and only the ONNX file itself, so a data_files value is accepted but has no effect. Unknown parameters return HTTP 400.
  • OpenAIEmbeddings. endpoint sets the base URL of the embeddings server; the default is the OpenAI service, so an air-gapped deployment sets endpoint to an internal server. api_key sets the key, and a server without authentication needs no key. Unknown parameters return HTTP 400.
  • HuggingFaceEmbeddings and GPT4AllEmbeddings. Parameters are passed to the LangChain class of the provider. The model name always comes from the model field.
  • HuggingFaceEndpointEmbeddings. huggingfacehub_api_token is required. Other parameters are passed to the LangChain class.

Foundation4 stores parameters as plain configuration and returns the parameters with the model. A key with read permission on embedding models can therefore read any API key or token in the parameters. Security at a glance describes this exposure.

Text splitters​

A text splitter divides the text of a document into fragments before embedding. All text splitters run in the gRPC service and wrap the text splitters of the LangChain library.

ProviderSplitting method
RecursiveCharacterTextSplitterSplits on a list of separators in order (paragraphs, lines, words) until each fragment fits the size limit
CharacterTextSplitterSplits on one separator
TokenTextSplitterSplits by token count, using a tiktoken encoding
CodeTextSplitterSplits source code on the syntax boundaries of a programming language, named in the language parameter
MarkdownHeaderTextSplitterSplits Markdown at the headings named in the headers_to_split_on parameter
NLTKTextSplitterSplits on sentence boundaries detected by the Natural Language Toolkit (NLTK) library

The provider RecursiveJsonSplitter is listed by GET /providers/text-splitters but cannot be created in the current release.

A text splitter is created with POST /text-splitters. The request contains provider, an optional description and optional parameters. A text splitter has no name. Foundation4 passes the parameters unchanged to the LangChain class as keyword arguments, so the parameter names and defaults are those of LangChain, such as chunk_size and chunk_overlap for the character splitters. When the splitter is created, Foundation4 splits a test string, and a parameter that the LangChain class does not accept returns HTTP 400.

The following request creates a recursive character splitter with fragments of up to 1,000 characters that overlap by up to 200 characters:

curl -X POST "$FOUNDATION4_URL/text-splitters" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"provider": "RecursiveCharacterTextSplitter",
"description": "1,000-character fragments with 200 characters of overlap",
"parameters": {"chunk_size": 1000, "chunk_overlap": 200}
}'

The response has status 201 and contains the stored splitter. The splitter breaks text at paragraphs, lines and spaces, so fragments and overlaps end at word boundaries. For example, a single paragraph of 5,940 characters gives eight fragments: seven of 992 to 999 characters and a last fragment of 334 characters, with overlaps of 192 to 199 characters.

Only the description of a text splitter can change after creation, and a change request that also sends parameters ignores them. A splitter that is the default of a pipeline cannot be deleted; the delete request returns HTTP 409.

Splitter selection per document​

Each pipeline names a default text splitter, and a document can name a different splitter in the text_splitter_id field of the create request. A pipeline can therefore hold prose and source code, each split by a suitable splitter.

The worker selects the splitter when the worker processes the document: the document's own splitter if the document names one, otherwise the pipeline's default splitter at that time. Changing a pipeline's default splitter therefore affects documents that are still waiting for processing, and does not affect fragments already stored.

A splitter that documents name but that no pipeline uses as its default can be deleted. The fragments already stored are unaffected, and the documents still show the identifier of the deleted splitter in text_splitter_id. A document that is still waiting for processing needs the splitter that it names, so an administrator deletes such a splitter only after every document that names it has been processed.

Seeded defaults​

A new installation includes the following objects, so a first pipeline can be created without configuring a provider:

ObjectProviderConfiguration
Embedding model all-MiniLM-L6-v2FastEmbedEmbeddingsModel Qdrant/all-MiniLM-L6-v2-onnx, 384 dimensions, model files included in the API server image
Text splitterRecursiveCharacterTextSplitterLangChain default parameters
Text splitterCharacterTextSplitterLangChain default parameters
Text splitterNLTKTextSplitterLangChain default parameters. Requires the nltk Python package, which the standard gRPC service image does not include.

The seeded embedding model is named all-MiniLM-L6-v2; the value that FastEmbed loads is the model field, Qdrant/all-MiniLM-L6-v2-onnx. A new embedding model that uses the same files sets model to that value.

The seeded model reads only the beginning of each text. Measured with POST /embedding-models/{id}/query on English prose, the first 2,524 characters of a 5,940-character text produced exactly the vector of the whole text, so the remaining characters did not affect the vector. The seeded recursive splitter makes fragments of up to 4,000 characters, so with both defaults, the end of a long fragment is not represented in vector search. A splitter with a chunk_size of 2,000 characters or less keeps fragments of English text within what the model reads. The seeded text splitters have no description, and text splitters have no name, so the seeded splitters are identified by their provider and identifier.

Air-gapped operation​

An air-gapped deployment provides every model file, package and endpoint inside the network:

  • FastEmbed. FastEmbed reads model files from the FastEmbed model directory and makes no external calls when the files are present. The API server image includes the files of two built-in models: the seeded model and nomic-ai/nomic-embed-text-v1.5. FastEmbed downloads other built-in models from Hugging Face on first use, so an air-gapped deployment uses the included models or custom models from a model image. A model image mounted at the FastEmbed model directory also prevents downloads: with such an image, a built-in model whose files the image does not include fails to load, even with network access.
  • OpenAI-compatible embeddings. The endpoint parameter points to an embeddings server inside the network.
  • Hugging Face and GPT4All. Model files and the provider's Python package are added as model and package images. Both libraries can contact external services while loading a model.
  • Hugging Face inference endpoint. The inference endpoint must be reachable from the gRPC service.
  • Token splitter. tiktoken downloads the encoding files of the selected encoding unless the files are cached locally, and no image provides the files, so the token splitter needs network access.

Providers and models lists the network access of each provider, and Models, packages and air-gapped installs lists the images and the configuration for each provider.