Models, packages and air-gapped installs
This page describes how Foundation4 delivers embedding models and Python packages, how an operator adds model and package images to a deployment, and which images an air-gapped installation mirrors. The page also states which embedding providers, text splitters and language model features work without network access. Operators who plan an air-gapped installation or add embedding models need this page. Providers and models lists the parameters and network access of every provider.
Model and package delivery
Foundation4 runs the API server and the workers from the same image. Each API server pod and each worker pod also runs the gRPC service as a sidecar container, a second container in the same Kubernetes pod, which hosts the Python embedding providers and every text splitter.
| Source | Content | Consumers |
|---|---|---|
| API server image | The Foundation4 binary and the FastEmbed model files listed in Default model shadowing, in /data/fastembed | API server and worker processes |
| gRPC service image | The Python service and the LangChain libraries, without the GPT4All, Sentence Transformers and NLTK packages | gRPC service |
| Package image | One Python package with dependencies, such as gpt4all or sentence-transformers | gRPC service, mounted at /packages/<name> |
| Model image | Model files for one provider, such as GPT4All GGUF files or NLTK data | Every container of the API server and worker pods, mounted at the path in mount |
Model and package images use the Kubernetes image volume source: the kubelet pulls an OCI image and mounts the image content read-only, without running the image. The cluster therefore needs image volume support in the API server, the kubelet and the container runtime. The charts declare no minimum Kubernetes version, so the operator confirms support before adding a model or package image. Deployment overview and requirements lists the cluster requirements. An installation that uses only the models in the API server image, such as the evaluation profile in Install for evaluation, needs no image volumes.
Values for model and package images
Model and package images are declared under api-server in the values file of the application release:
api-server:
packages:
- name: gpt4all
image:
reference: <registry>/<package image repository>:<tag>
pullPolicy: IfNotPresent
models:
- name: gpt4all
image:
reference: <registry>/<model image repository>:<tag>
pullPolicy: IfNotPresent
mount: /data/gpt4all
| Key | Effect |
|---|---|
packages[].name | Names the volume and the mount directory /packages/<name> in both gRPC service containers |
packages[].image | Copied unchanged into the image volume source: reference and pullPolicy |
packages[].imagePullSecrets | A list of name entries, added to the API server pod only |
models[].name | Names the volume |
models[].image | Copied unchanged into the image volume source: reference and pullPolicy |
models[].mount | The mount path in every container of the API server and worker pods |
models[].imagePullSecrets | A list of name entries, added to the API server pod only |
The worker pods and the hook Jobs take pull secrets from global.imagePullSecrets, or from api-server.imagePullSecrets when the global list is empty, and ignore the per-image lists. A pull secret that a model or package image needs therefore also goes in global.imagePullSecrets. Values reference describes every key of the charts.
A change to packages or models changes the pod templates, so helm upgrade replaces the API server and worker pods without a manual restart.
Container mounts
The chart mounts model images differently in the worker container than in the other containers:
| Container | Pod | Package images | Model images |
|---|---|---|---|
server | API server | Not mounted | The whole image at mount |
grpc | API server | /packages/<name> | The whole image at mount |
worker | Worker | Not mounted | The data subdirectory of the image at mount |
grpc | Worker | /packages/<name> | The whole image at mount |
The current build of the model images places the model files at the image root and creates no data directory. The behavior of the worker container with such images has not been confirmed, and engineering has not yet confirmed which image layout is correct. The difference matters for FastEmbed model images, because the worker process reads FastEmbed files at /data/fastembed. GPT4All, Hugging Face and NLTK files are read only by the gRPC service, which mounts the whole image in both pods. After any change to models, the operator compares the files that the server and worker containers see, as shown in step 7 of Custom FastEmbed model. Subpath mounts of image volumes also need a Kubernetes version that supports subpaths on image volumes (unverified).
Model directories
The paths configuration keys tell Foundation4 where each provider reads model files. The API server and worker processes read paths.fastembed directly and send the other paths to the gRPC service with each request.
| Key | Chart value | Provider | Content |
|---|---|---|---|
paths.fastembed | /data/fastembed | FastEmbedEmbeddings | ONNX models in the Hugging Face cache layout |
paths.gpt4all | /data/gpt4all | GPT4AllEmbeddings | GGUF model files |
paths.huggingface | /data/huggingface | HuggingFaceEmbeddings | Sentence Transformers models in the Hugging Face cache layout |
paths.nltk | /data/nltk | NLTKTextSplitter | NLTK data, such as punkt_tab |
The mount value of a model image equals the paths value of the provider that reads the files. The chart values match the mount paths that this page uses for the model images, so a deployment changes a path only for a different directory layout. Configuration reference describes how a configs file sets these keys.
Default model shadowing
A model image mounted at /data/fastembed replaces the directory of the API server image. The FastEmbed files in the API server image are then hidden in every container that mounts the model image.
The seeded embedding model all-MiniLM-L6-v2 loads Qdrant/all-MiniLM-L6-v2-onnx. A FastEmbed model image without those files leaves the seeded model, and every pipeline that uses the seeded model, unable to embed text. A FastEmbed model image therefore contains every FastEmbed model that the deployment uses, including the seeded model.
The API server image holds the following entries in /data/fastembed:
| Entry | Model |
|---|---|
models--qdrant--all-MiniLM-L6-v2-onnx | Qdrant/all-MiniLM-L6-v2-onnx, the seeded model |
models--Qdrant--all-MiniLM-L6-v2-onnx | A link to the previous entry, for the capitalized name |
models--Xenova--bge-small-en-v1.5 | Files of Xenova/bge-small-en-v1.5, which the built-in name BAAI/bge-small-en-v1.5 does not load |
models--nomic-ai--nomic-embed-text-v1.5 | nomic-ai/nomic-embed-text-v1.5 |
The following command lists the directory of a running API server:
kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- ls -la /data/fastembed
FastEmbed models
FastEmbed accepts two kinds of models:
- Built-in models. Two of the ten built-in models, the seeded model and
nomic-ai/nomic-embed-text-v1.5, load from the files in the API server image. FastEmbed downloads the files of the other eight from Hugging Face the first time each process loads the model, so those models fail in an air-gapped deployment and in any container whose FastEmbed directory is a read-only model image. - Custom ONNX models. FastEmbed reads a custom model only from local files in the FastEmbed model directory and never downloads the model. A custom model is the supported way to add a FastEmbed model to an air-gapped deployment.
A custom model occupies one directory in the Hugging Face cache layout, where the model name has each / replaced by --:
<paths.fastembed>/models--<model>/snapshots/<snapshot>/
The snapshot directory holds the ONNX file named by the onnx parameter and the files tokenizer.json, config.json, special_tokens_map.json and tokenizer_config.json. The snapshots directory holds exactly one snapshot directory. The loader reads the ONNX file into memory, so an ONNX export that stores weights in separate external data files is not expected to load (unverified); a single-file ONNX export avoids the question. Providers and models lists the parameters.
Custom FastEmbed model
The following procedure builds a FastEmbed model image that contains the files of the API server image and one custom model, example-org/example-embed-onnx. The build runs on a host with Docker and access to the registry that the cluster uses.
-
Place the model files in the build directory and list the files:
find models--example-org--example-embed-onnx -type fExpected result: five files under
models--example-org--example-embed-onnx/snapshots/<snapshot>/: the ONNX file, such asonnx/model.onnx, and the four JSON files. -
Create a file named
Dockerfilein the build directory. The first stage reads the FastEmbed directory of the API server image, so the new image keeps the seeded model and the other two included models:FROM <registry>/foundation4ai-api:<api server image tag> AS apiFROM scratchCOPY --from=api /data/fastembed/ /COPY models--example-org--example-embed-onnx/ /models--example-org--example-embed-onnx/Expected result: the build directory holds
Dockerfileand the model directory. -
Build and push the image:
docker build -t <registry>/fastembed-models:1.0.0 .docker push <registry>/fastembed-models:1.0.0Expected result: the push ends with a line that reports the image digest.
-
Add the image to the values file of the application release:
api-server:models:- name: fastembedimage:reference: <registry>/fastembed-models:1.0.0pullPolicy: IfNotPresentmount: /data/fastembedExpected result:
modelslists the entry once, next to any other model images of the deployment. -
Upgrade the application release:
helm upgrade --install foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployed. -
Check the pods:
kubectl -n foundation4ai get podsExpected result: new API server and worker pods show
2/2containers ready andRunning. -
Compare the FastEmbed directory in the
serverandworkercontainers:kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- ls /data/fastembedkubectl -n foundation4ai exec deploy/foundation4ai-api-server-worker -c worker -- ls /data/fastembedExpected result: both listings show the four entries of the API server image and
models--example-org--example-embed-onnx. A worker listing that differs from the server listing, or a worker pod that does not start, is the mount difference described in Container mounts; the operator removes the model image entry and reports the result to the Foundation4 provider before continuing. -
Create the embedding model. The shell variables are described in Authenticate:
curl -X POST "$FOUNDATION4_URL/embedding-models" \-H "x-api-key: $FOUNDATION4_API_KEY" \-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \-H "Content-Type: application/json" \-d '{"name": "example-embed","description": "Custom ONNX model from the FastEmbed model image","provider": "FastEmbedEmbeddings","model": "example-org/example-embed-onnx","size": 0,"parameters": {"onnx": "onnx/model.onnx", "pooling": "mean"}}'Expected result: status 201 and a body whose
sizeis the dimension of the model. A missing file returns HTTP 400 withError reading <file>indetails.error.
Providers in the gRPC service
The gRPC service loads Python packages from /packages at startup. For each directory in /packages, the service adds /packages/<name>/lib to the Python path when that directory exists, and /packages/<name> otherwise. A custom package image therefore contains the package directories of a Python 3.13 site-packages directory at the image root, built for the processor architecture of the cluster nodes.
| Provider or splitter | Package image | Model image and mount |
|---|---|---|
GPT4AllEmbeddings | gpt4all | gpt4all, at /data/gpt4all |
HuggingFaceEmbeddings | sentence-transformers | huggingface, at /data/huggingface |
NLTKTextSplitter | None available: no image provides the nltk package | nltk, at /data/nltk |
The seeded NLTKTextSplitter text splitter exists in every installation but cannot split text until an image provides the nltk package.
GPT4All provider
-
Add the package and model images to the values file of the application release:
api-server:packages:- name: gpt4allimage:reference: <registry>/<gpt4all package image>:<tag>pullPolicy: IfNotPresentmodels:- name: gpt4allimage:reference: <registry>/<gpt4all model image>:<tag>pullPolicy: IfNotPresentmount: /data/gpt4allExpected result:
packagesandmodelseach list onegpt4allentry. -
Upgrade the application release:
helm upgrade --install foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployed, and the new API server and worker pods show2/2containers ready. -
Confirm that the gRPC service loaded the package:
kubectl -n foundation4ai logs deploy/foundation4ai-api-server -c grpc | grep "Added package directory"Expected result:
INFO:processor:Added package directory to site: /packages/gpt4allor the same line ending in/packages/gpt4all/lib. -
Confirm that the model files are present:
kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c grpc -- ls /data/gpt4allExpected result:
all-MiniLM-L6-v2.gguf2.f16.ggufandnomic-embed-text-v1.5.f16.gguf.
The GPT4All library can contact the GPT4All service while loading a model, so an air-gapped deployment creates a GPT4All embedding model in a test before relying on the provider. The Hugging Face provider follows the same steps with the sentence-transformers package image and the huggingface model image, with the same caution for the Hugging Face Hub.
Offline capabilities
The following table states what works without access to the internet. The conditions match Providers and models.
| Capability | Without internet access | Condition or reason |
|---|---|---|
| FastEmbed, the two models in the API server image | Yes | Files present in the API server image, or in the FastEmbed model image that replaces the directory |
| FastEmbed, the other eight built-in models | No | FastEmbed downloads the files from Hugging Face on first use in each process |
| FastEmbed, custom ONNX models | Yes | Files in a FastEmbed model image |
| OpenAI-compatible embeddings | Yes | endpoint names an embeddings server inside the network |
| Hugging Face Sentence Transformers | Not verified | The library can contact the Hugging Face Hub while loading a model |
| GPT4All | Not verified | The library can contact the GPT4All service while loading a model |
| Hugging Face inference endpoint | Only with an internal endpoint | The endpoint must be reachable from the gRPC service |
| Recursive character, character, code and Markdown header splitters | Yes | None |
| Token splitter | No | tiktoken downloads the encoding files, which no image provides |
| NLTK splitter | Not available | No image provides the nltk package |
| Agents, direct queries and the OpenAI-compatible endpoint | Yes, with an internal model server | The model server must be reachable from the API server pods |
foundation4ai data download-models | No | The command downloads into the storage of the containers of the pod that runs the command and does not prepare other pods |
Language models in an air-gapped deployment
Foundation4 includes no language model and has no default model server. An administrator registers a model server inside the network with POST /llms, as described in LLMs. The API server calls the model server, so the network allows traffic from the API server pods to the model server.
Agents and direct queries require a model server that implements the OpenAI Responses API, and the registration carries an API key even when the model server does not check keys. A model server that implements only chat completions serves the OpenAI-compatible endpoint of Foundation4 but not agents or direct queries.
Network access in an air-gapped deployment
With the capabilities marked Yes in Offline capabilities, Foundation4 components connect only to services inside the network:
| Component | Destinations |
|---|---|
| API server | PostgreSQL, the Redis-compatible cache, NATS JetStream, the gRPC service in the same pod, registered model servers and embedding servers |
| Worker | PostgreSQL, the Redis-compatible cache, NATS JetStream, the gRPC service in the same pod, embedding servers |
| gRPC service | A Hugging Face inference endpoint, when a model uses one |
| Dashboard | None; the browser calls the API server |
A model server or embedding server that uses TLS presents a certificate that the API server image trusts. The charts provide no value that adds a certificate authority to the images, so a server certificate from a private certificate authority is an open point to settle with the Foundation4 provider before installation.
Images to mirror
An air-gapped installation copies every image that the releases use into a registry inside the network and points each values key at the copy. The chart directories include the archives of the PostgreSQL, Valkey, NATS and Prometheus subcharts, so the installation reads no Helm repository.
| Role | Values key | Image |
|---|---|---|
| API server, workers and hook Jobs | api-server.image.repository with global.serverImageTag | Supplied by the Foundation4 provider |
| gRPC service | api-server.image.repositoryGrpc with global.grpcImageTag | Supplied by the Foundation4 provider |
| Dashboard | dashboard.image.repository with global.dashboardImageTag | Supplied by the Foundation4 provider |
| Model images | api-server.models[].image.reference | Supplied by the Foundation4 provider, or built as in Custom FastEmbed model |
| Package images | api-server.packages[].image.reference | Supplied by the Foundation4 provider |
| NATS server | nats.container.image | nats, tag 2.12.4-alpine |
| NATS configuration reloader | nats.reloader.image | natsio/nats-server-config-reloader, tag 0.21.1 |
| NATS utility container | nats.natsBox.container.image | natsio/nats-box, tag 0.19.3 |
| Redis-compatible cache (Valkey) | redis.image | valkey/valkey, the chart default tag 9.0.1 |
| Prometheus server | prometheus.server.image | prometheus/prometheus, the chart default tag v3.9.1 |
| Prometheus configuration reloader | prometheus.configmapReload.prometheus.image | prometheus-operator/prometheus-config-reloader, tag v0.89.0 |
| PostgreSQL, evaluation profile only | postgres.image | pgvector/pgvector, tag pg18-trixie |
The dependency checks of the deployment files (just check) start three further utility images: natsio/nats-box:latest, redis:alpine and curlimages/curl. An air-gapped installation mirrors these images too, or runs the equivalent checks in Pre-installation checks. Operator tools, such as kubectl, Helm and a container runtime for the mirroring host, are also brought into the network.
Image list from a staging installation
A staging installation with internet access gives the exact list of images, including tags that the charts derive from the chart versions:
-
List every container image and image volume reference in the namespace:
kubectl get pods -n foundation4ai -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.image}{"\n"}{end}{range .spec.volumes[*]}{.image.reference}{"\n"}{end}{end}' \| grep -v '^$' | sort -u > images.txtcat images.txtExpected result: one line per image, including the NATS, Valkey, Prometheus and Foundation4 images and every model and package image.
-
Pull the images and write one archive:
xargs -n 1 docker pull < images.txtdocker save -o foundation4ai-images.tar $(cat images.txt)Expected result:
foundation4ai-images.tarexists anddocker savereports no error. -
Inside the network, load the archive, then tag and push each image to the internal registry:
docker load -i foundation4ai-images.tardocker tag <source image>:<tag> <registry>/<repository>:<tag>docker push <registry>/<repository>:<tag>Expected result:
docker loadreportsLoaded imagefor each image, and each push reports a digest.
Registry values
The following values point every chart at an internal registry that keeps the source repository paths. The NATS chart prefixes global.image.registry to each NATS image, and the Valkey chart prefixes global.imageRegistry.
global:
serverImageTag: "<api server image tag>"
grpcImageTag: "<grpc service image tag>"
dashboardImageTag: "<dashboard image tag>"
imagePullSecrets:
- name: <pull secret>
image:
registry: <registry>
pullSecretNames:
- <pull secret>
imageRegistry: <registry>
api-server:
image:
repository: <registry>/foundation4ai-api
repositoryGrpc: <registry>/foundation4ai-grpc
dashboard:
image:
repository: <registry>/foundation4ai-dashboard
prometheus:
imagePullSecrets:
- name: <pull secret>
server:
image:
repository: <registry>/prometheus/prometheus
configmapReload:
prometheus:
image:
repository: <registry>/prometheus-operator/prometheus-config-reloader
postgres:
image:
registry: <registry>
imagePullSecrets:
- name: <pull secret>
redis:
imagePullSecrets:
- <pull secret>
The charts read pull secrets in three shapes. The Foundation4 charts, Prometheus and PostgreSQL read a list of name entries. NATS reads global.image.pullSecretNames, a list of names. Valkey reads global.imagePullSecrets and redis.imagePullSecrets as lists of names, so the name entries that the Foundation4 charts need produce an unusable pull secret reference in the Valkey pod, and the redis.imagePullSecrets entry supplies the usable one, as described in Charts, images and installation bundle. A registry inside the network that requires no pull secret needs none of the pull secret keys.