Skip to main content

Models, packages and air-gapped installs

This page describes how Foundation4 delivers embedding models and Python packages, how an operator adds model and package images to a deployment, and which images an air-gapped installation mirrors. The page also states which embedding providers, text splitters and language model features work without network access. Operators who plan an air-gapped installation or add embedding models need this page. Providers and models lists the parameters and network access of every provider.

Model and package delivery​

Foundation4 runs the API server and the workers from the same image. Each API server pod and each worker pod also runs the gRPC service as a sidecar container, a second container in the same Kubernetes pod, which hosts the Python embedding providers and every text splitter.

SourceContentConsumers
API server imageThe Foundation4 binary and the FastEmbed model files listed in Default model shadowing, in /data/fastembedAPI server and worker processes
gRPC service imageThe Python service and the LangChain libraries, without the GPT4All, Sentence Transformers and NLTK packagesgRPC service
Package imageOne Python package with dependencies, such as gpt4all or sentence-transformersgRPC service, mounted at /packages/<name>
Model imageModel files for one provider, such as GPT4All GGUF files or NLTK dataEvery container of the API server and worker pods, mounted at the path in mount

Model and package images use the Kubernetes image volume source: the kubelet pulls an OCI image and mounts the image content read-only, without running the image. The cluster therefore needs image volume support in the API server, the kubelet and the container runtime. The charts declare no minimum Kubernetes version, so the operator confirms support before adding a model or package image. Deployment overview and requirements lists the cluster requirements. An installation that uses only the models in the API server image, such as the evaluation profile in Install for evaluation, needs no image volumes.

Values for model and package images​

Model and package images are declared under api-server in the values file of the application release:

api-server:
packages:
- name: gpt4all
image:
reference: <registry>/<package image repository>:<tag>
pullPolicy: IfNotPresent
models:
- name: gpt4all
image:
reference: <registry>/<model image repository>:<tag>
pullPolicy: IfNotPresent
mount: /data/gpt4all
KeyEffect
packages[].nameNames the volume and the mount directory /packages/<name> in both gRPC service containers
packages[].imageCopied unchanged into the image volume source: reference and pullPolicy
packages[].imagePullSecretsA list of name entries, added to the API server pod only
models[].nameNames the volume
models[].imageCopied unchanged into the image volume source: reference and pullPolicy
models[].mountThe mount path in every container of the API server and worker pods
models[].imagePullSecretsA list of name entries, added to the API server pod only

The worker pods and the hook Jobs take pull secrets from global.imagePullSecrets, or from api-server.imagePullSecrets when the global list is empty, and ignore the per-image lists. A pull secret that a model or package image needs therefore also goes in global.imagePullSecrets. Values reference describes every key of the charts.

A change to packages or models changes the pod templates, so helm upgrade replaces the API server and worker pods without a manual restart.

Container mounts​

The chart mounts model images differently in the worker container than in the other containers:

ContainerPodPackage imagesModel images
serverAPI serverNot mountedThe whole image at mount
grpcAPI server/packages/<name>The whole image at mount
workerWorkerNot mountedThe data subdirectory of the image at mount
grpcWorker/packages/<name>The whole image at mount

The current build of the model images places the model files at the image root and creates no data directory. The behavior of the worker container with such images has not been confirmed, and engineering has not yet confirmed which image layout is correct. The difference matters for FastEmbed model images, because the worker process reads FastEmbed files at /data/fastembed. GPT4All, Hugging Face and NLTK files are read only by the gRPC service, which mounts the whole image in both pods. After any change to models, the operator compares the files that the server and worker containers see, as shown in step 7 of Custom FastEmbed model. Subpath mounts of image volumes also need a Kubernetes version that supports subpaths on image volumes (unverified).

Model directories​

The paths configuration keys tell Foundation4 where each provider reads model files. The API server and worker processes read paths.fastembed directly and send the other paths to the gRPC service with each request.

KeyChart valueProviderContent
paths.fastembed/data/fastembedFastEmbedEmbeddingsONNX models in the Hugging Face cache layout
paths.gpt4all/data/gpt4allGPT4AllEmbeddingsGGUF model files
paths.huggingface/data/huggingfaceHuggingFaceEmbeddingsSentence Transformers models in the Hugging Face cache layout
paths.nltk/data/nltkNLTKTextSplitterNLTK data, such as punkt_tab

The mount value of a model image equals the paths value of the provider that reads the files. The chart values match the mount paths that this page uses for the model images, so a deployment changes a path only for a different directory layout. Configuration reference describes how a configs file sets these keys.

Default model shadowing​

A model image mounted at /data/fastembed replaces the directory of the API server image. The FastEmbed files in the API server image are then hidden in every container that mounts the model image.

The seeded embedding model all-MiniLM-L6-v2 loads Qdrant/all-MiniLM-L6-v2-onnx. A FastEmbed model image without those files leaves the seeded model, and every pipeline that uses the seeded model, unable to embed text. A FastEmbed model image therefore contains every FastEmbed model that the deployment uses, including the seeded model.

The API server image holds the following entries in /data/fastembed:

EntryModel
models--qdrant--all-MiniLM-L6-v2-onnxQdrant/all-MiniLM-L6-v2-onnx, the seeded model
models--Qdrant--all-MiniLM-L6-v2-onnxA link to the previous entry, for the capitalized name
models--Xenova--bge-small-en-v1.5Files of Xenova/bge-small-en-v1.5, which the built-in name BAAI/bge-small-en-v1.5 does not load
models--nomic-ai--nomic-embed-text-v1.5nomic-ai/nomic-embed-text-v1.5

The following command lists the directory of a running API server:

kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- ls -la /data/fastembed

FastEmbed models​

FastEmbed accepts two kinds of models:

  • Built-in models. Two of the ten built-in models, the seeded model and nomic-ai/nomic-embed-text-v1.5, load from the files in the API server image. FastEmbed downloads the files of the other eight from Hugging Face the first time each process loads the model, so those models fail in an air-gapped deployment and in any container whose FastEmbed directory is a read-only model image.
  • Custom ONNX models. FastEmbed reads a custom model only from local files in the FastEmbed model directory and never downloads the model. A custom model is the supported way to add a FastEmbed model to an air-gapped deployment.

A custom model occupies one directory in the Hugging Face cache layout, where the model name has each / replaced by --:

<paths.fastembed>/models--<model>/snapshots/<snapshot>/

The snapshot directory holds the ONNX file named by the onnx parameter and the files tokenizer.json, config.json, special_tokens_map.json and tokenizer_config.json. The snapshots directory holds exactly one snapshot directory. The loader reads the ONNX file into memory, so an ONNX export that stores weights in separate external data files is not expected to load (unverified); a single-file ONNX export avoids the question. Providers and models lists the parameters.

Custom FastEmbed model​

The following procedure builds a FastEmbed model image that contains the files of the API server image and one custom model, example-org/example-embed-onnx. The build runs on a host with Docker and access to the registry that the cluster uses.

  1. Place the model files in the build directory and list the files:

    find models--example-org--example-embed-onnx -type f

    Expected result: five files under models--example-org--example-embed-onnx/snapshots/<snapshot>/: the ONNX file, such as onnx/model.onnx, and the four JSON files.

  2. Create a file named Dockerfile in the build directory. The first stage reads the FastEmbed directory of the API server image, so the new image keeps the seeded model and the other two included models:

    FROM <registry>/foundation4ai-api:<api server image tag> AS api

    FROM scratch
    COPY --from=api /data/fastembed/ /
    COPY models--example-org--example-embed-onnx/ /models--example-org--example-embed-onnx/

    Expected result: the build directory holds Dockerfile and the model directory.

  3. Build and push the image:

    docker build -t <registry>/fastembed-models:1.0.0 .
    docker push <registry>/fastembed-models:1.0.0

    Expected result: the push ends with a line that reports the image digest.

  4. Add the image to the values file of the application release:

    api-server:
    models:
    - name: fastembed
    image:
    reference: <registry>/fastembed-models:1.0.0
    pullPolicy: IfNotPresent
    mount: /data/fastembed

    Expected result: models lists the entry once, next to any other model images of the deployment.

  5. Upgrade the application release:

    helm upgrade --install foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed.

  6. Check the pods:

    kubectl -n foundation4ai get pods

    Expected result: new API server and worker pods show 2/2 containers ready and Running.

  7. Compare the FastEmbed directory in the server and worker containers:

    kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- ls /data/fastembed
    kubectl -n foundation4ai exec deploy/foundation4ai-api-server-worker -c worker -- ls /data/fastembed

    Expected result: both listings show the four entries of the API server image and models--example-org--example-embed-onnx. A worker listing that differs from the server listing, or a worker pod that does not start, is the mount difference described in Container mounts; the operator removes the model image entry and reports the result to the Foundation4 provider before continuing.

  8. Create the embedding model. The shell variables are described in Authenticate:

    curl -X POST "$FOUNDATION4_URL/embedding-models" \
    -H "x-api-key: $FOUNDATION4_API_KEY" \
    -H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
    -H "Content-Type: application/json" \
    -d '{
    "name": "example-embed",
    "description": "Custom ONNX model from the FastEmbed model image",
    "provider": "FastEmbedEmbeddings",
    "model": "example-org/example-embed-onnx",
    "size": 0,
    "parameters": {"onnx": "onnx/model.onnx", "pooling": "mean"}
    }'

    Expected result: status 201 and a body whose size is the dimension of the model. A missing file returns HTTP 400 with Error reading <file> in details.error.

Providers in the gRPC service​

The gRPC service loads Python packages from /packages at startup. For each directory in /packages, the service adds /packages/<name>/lib to the Python path when that directory exists, and /packages/<name> otherwise. A custom package image therefore contains the package directories of a Python 3.13 site-packages directory at the image root, built for the processor architecture of the cluster nodes.

Provider or splitterPackage imageModel image and mount
GPT4AllEmbeddingsgpt4allgpt4all, at /data/gpt4all
HuggingFaceEmbeddingssentence-transformershuggingface, at /data/huggingface
NLTKTextSplitterNone available: no image provides the nltk packagenltk, at /data/nltk

The seeded NLTKTextSplitter text splitter exists in every installation but cannot split text until an image provides the nltk package.

GPT4All provider​

  1. Add the package and model images to the values file of the application release:

    api-server:
    packages:
    - name: gpt4all
    image:
    reference: <registry>/<gpt4all package image>:<tag>
    pullPolicy: IfNotPresent
    models:
    - name: gpt4all
    image:
    reference: <registry>/<gpt4all model image>:<tag>
    pullPolicy: IfNotPresent
    mount: /data/gpt4all

    Expected result: packages and models each list one gpt4all entry.

  2. Upgrade the application release:

    helm upgrade --install foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed, and the new API server and worker pods show 2/2 containers ready.

  3. Confirm that the gRPC service loaded the package:

    kubectl -n foundation4ai logs deploy/foundation4ai-api-server -c grpc | grep "Added package directory"

    Expected result: INFO:processor:Added package directory to site: /packages/gpt4all or the same line ending in /packages/gpt4all/lib.

  4. Confirm that the model files are present:

    kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c grpc -- ls /data/gpt4all

    Expected result: all-MiniLM-L6-v2.gguf2.f16.gguf and nomic-embed-text-v1.5.f16.gguf.

The GPT4All library can contact the GPT4All service while loading a model, so an air-gapped deployment creates a GPT4All embedding model in a test before relying on the provider. The Hugging Face provider follows the same steps with the sentence-transformers package image and the huggingface model image, with the same caution for the Hugging Face Hub.

Offline capabilities​

The following table states what works without access to the internet. The conditions match Providers and models.

CapabilityWithout internet accessCondition or reason
FastEmbed, the two models in the API server imageYesFiles present in the API server image, or in the FastEmbed model image that replaces the directory
FastEmbed, the other eight built-in modelsNoFastEmbed downloads the files from Hugging Face on first use in each process
FastEmbed, custom ONNX modelsYesFiles in a FastEmbed model image
OpenAI-compatible embeddingsYesendpoint names an embeddings server inside the network
Hugging Face Sentence TransformersNot verifiedThe library can contact the Hugging Face Hub while loading a model
GPT4AllNot verifiedThe library can contact the GPT4All service while loading a model
Hugging Face inference endpointOnly with an internal endpointThe endpoint must be reachable from the gRPC service
Recursive character, character, code and Markdown header splittersYesNone
Token splitterNotiktoken downloads the encoding files, which no image provides
NLTK splitterNot availableNo image provides the nltk package
Agents, direct queries and the OpenAI-compatible endpointYes, with an internal model serverThe model server must be reachable from the API server pods
foundation4ai data download-modelsNoThe command downloads into the storage of the containers of the pod that runs the command and does not prepare other pods

Language models in an air-gapped deployment​

Foundation4 includes no language model and has no default model server. An administrator registers a model server inside the network with POST /llms, as described in LLMs. The API server calls the model server, so the network allows traffic from the API server pods to the model server.

Agents and direct queries require a model server that implements the OpenAI Responses API, and the registration carries an API key even when the model server does not check keys. A model server that implements only chat completions serves the OpenAI-compatible endpoint of Foundation4 but not agents or direct queries.

Network access in an air-gapped deployment​

With the capabilities marked Yes in Offline capabilities, Foundation4 components connect only to services inside the network:

ComponentDestinations
API serverPostgreSQL, the Redis-compatible cache, NATS JetStream, the gRPC service in the same pod, registered model servers and embedding servers
WorkerPostgreSQL, the Redis-compatible cache, NATS JetStream, the gRPC service in the same pod, embedding servers
gRPC serviceA Hugging Face inference endpoint, when a model uses one
DashboardNone; the browser calls the API server

A model server or embedding server that uses TLS presents a certificate that the API server image trusts. The charts provide no value that adds a certificate authority to the images, so a server certificate from a private certificate authority is an open point to settle with the Foundation4 provider before installation.

Images to mirror​

An air-gapped installation copies every image that the releases use into a registry inside the network and points each values key at the copy. The chart directories include the archives of the PostgreSQL, Valkey, NATS and Prometheus subcharts, so the installation reads no Helm repository.

RoleValues keyImage
API server, workers and hook Jobsapi-server.image.repository with global.serverImageTagSupplied by the Foundation4 provider
gRPC serviceapi-server.image.repositoryGrpc with global.grpcImageTagSupplied by the Foundation4 provider
Dashboarddashboard.image.repository with global.dashboardImageTagSupplied by the Foundation4 provider
Model imagesapi-server.models[].image.referenceSupplied by the Foundation4 provider, or built as in Custom FastEmbed model
Package imagesapi-server.packages[].image.referenceSupplied by the Foundation4 provider
NATS servernats.container.imagenats, tag 2.12.4-alpine
NATS configuration reloadernats.reloader.imagenatsio/nats-server-config-reloader, tag 0.21.1
NATS utility containernats.natsBox.container.imagenatsio/nats-box, tag 0.19.3
Redis-compatible cache (Valkey)redis.imagevalkey/valkey, the chart default tag 9.0.1
Prometheus serverprometheus.server.imageprometheus/prometheus, the chart default tag v3.9.1
Prometheus configuration reloaderprometheus.configmapReload.prometheus.imageprometheus-operator/prometheus-config-reloader, tag v0.89.0
PostgreSQL, evaluation profile onlypostgres.imagepgvector/pgvector, tag pg18-trixie

The dependency checks of the deployment files (just check) start three further utility images: natsio/nats-box:latest, redis:alpine and curlimages/curl. An air-gapped installation mirrors these images too, or runs the equivalent checks in Pre-installation checks. Operator tools, such as kubectl, Helm and a container runtime for the mirroring host, are also brought into the network.

Image list from a staging installation​

A staging installation with internet access gives the exact list of images, including tags that the charts derive from the chart versions:

  1. List every container image and image volume reference in the namespace:

    kubectl get pods -n foundation4ai -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.image}{"\n"}{end}{range .spec.volumes[*]}{.image.reference}{"\n"}{end}{end}' \
    | grep -v '^$' | sort -u > images.txt
    cat images.txt

    Expected result: one line per image, including the NATS, Valkey, Prometheus and Foundation4 images and every model and package image.

  2. Pull the images and write one archive:

    xargs -n 1 docker pull < images.txt
    docker save -o foundation4ai-images.tar $(cat images.txt)

    Expected result: foundation4ai-images.tar exists and docker save reports no error.

  3. Inside the network, load the archive, then tag and push each image to the internal registry:

    docker load -i foundation4ai-images.tar
    docker tag <source image>:<tag> <registry>/<repository>:<tag>
    docker push <registry>/<repository>:<tag>

    Expected result: docker load reports Loaded image for each image, and each push reports a digest.

Registry values​

The following values point every chart at an internal registry that keeps the source repository paths. The NATS chart prefixes global.image.registry to each NATS image, and the Valkey chart prefixes global.imageRegistry.

global:
serverImageTag: "<api server image tag>"
grpcImageTag: "<grpc service image tag>"
dashboardImageTag: "<dashboard image tag>"
imagePullSecrets:
- name: <pull secret>
image:
registry: <registry>
pullSecretNames:
- <pull secret>
imageRegistry: <registry>

api-server:
image:
repository: <registry>/foundation4ai-api
repositoryGrpc: <registry>/foundation4ai-grpc

dashboard:
image:
repository: <registry>/foundation4ai-dashboard

prometheus:
imagePullSecrets:
- name: <pull secret>
server:
image:
repository: <registry>/prometheus/prometheus
configmapReload:
prometheus:
image:
repository: <registry>/prometheus-operator/prometheus-config-reloader

postgres:
image:
registry: <registry>
imagePullSecrets:
- name: <pull secret>

redis:
imagePullSecrets:
- <pull secret>

The charts read pull secrets in three shapes. The Foundation4 charts, Prometheus and PostgreSQL read a list of name entries. NATS reads global.image.pullSecretNames, a list of names. Valkey reads global.imagePullSecrets and redis.imagePullSecrets as lists of names, so the name entries that the Foundation4 charts need produce an unusable pull secret reference in the Valkey pod, and the redis.imagePullSecrets entry supplies the usable one, as described in Charts, images and installation bundle. A registry inside the network that requires no pull secret needs none of the pull secret keys.