Deployment overview and requirements
This page describes what a Foundation4 installation deploys on Kubernetes, the order in which the parts are installed, the two deployment profiles and the requirements for the cluster, the database and the operator's workstation. Operators read this page before planning an installation. Security reviewers use the requirements to assess the footprint of Foundation4 in a cluster. The procedure itself is described in Install on Kubernetes.
Components
A Foundation4 installation consists of two Helm releases in the namespace foundation4ai. The core release installs the services that Foundation4 depends on. The application release installs Foundation4 and reads the connection settings that the core release produces.
| Release | Chart | Installs |
|---|---|---|
foundation4ai-core | foundation4ai-core 0.3.0 | NATS JetStream (3 replicas and a NATS utility pod), Valkey as the Redis-compatible cache, the Prometheus server, optionally PostgreSQL with pgvector, and the Secret foundation4ai-core with the connection URLs, the license, the application secret and the master key |
foundation4ai | foundation4ai 0.3.0 | The API server (1 replica), the workers (3 replicas), the dashboard (1 replica), three installation Jobs, and optionally an Ingress or an HTTPRoute |
The releases have the following characteristics:
- Sidecar containers. The API server pod and each worker pod run two containers: the Foundation4 process and the gRPC service. The gRPC service is deployed as a sidecar container, a second container in the same Kubernetes pod, and hosts the Python text splitters and embedding providers.
- Installation Jobs. The application release runs the database migration, the license check and the creation of the master key as Helm hook Jobs before the Deployments start. All three Jobs run the API server image.
- Fixed names. The application release reads the Secret
foundation4ai-coreby name. The core release is therefore always namedfoundation4ai-coreand installed in the same namespace as the application release. - Rendering against the cluster. Both charts read Secrets from the cluster while Helm renders the templates. Installation therefore uses the Helm command line with access to the cluster. Tools that render charts offline, such as
helm templateor some GitOps controllers, fail to render the charts or produce Secrets without credentials.
Installation order
The installation follows a fixed order, because each step consumes the output of the step before:
- The operator creates the Secret
foundation4ai-secretsfrom the secrets filefoundation4ai.secrets.env. - The core release installs NATS JetStream, Valkey and Prometheus, and copies the connection URLs, the license, the application secret and the master key into the Secret
foundation4ai-core. - A first installation of the application release migrates the database. The license check Job then writes the system ID to the Job log and waits, so this installation stops at the Helm timeout.
- The operator sends the system ID to the Foundation4 provider and receives the license. The operator adds the license to the secrets file, updates the Secret and upgrades the core release, which copies the license into the Secret
foundation4ai-core. - The operator installs the application release again. The migration, license check and master key Jobs complete, and the API server, workers and dashboard start.
- The operator signs in with the master key and creates one API key for each client application. The master key stays with the operators and serves only for administration.
The license is bound to the system ID, which Foundation4 derives from the database catalog after the migration creates the schema. Licensing describes the system ID and the license.
Solid arrows show actions and requests; dashed arrows show information returned to the operator.
Later changes, such as new image tags or a renewed license, use helm upgrade of the core release and then of the application release.
Deployment profiles
Foundation4 supports two deployment profiles. Install for evaluation installs the evaluation profile; Install on Kubernetes installs the production profile.
| Aspect | Evaluation | Production |
|---|---|---|
| Nodes | One node, such as a single-node k3s cluster | Several nodes, sized for the workload |
| PostgreSQL | Bundled chart (postgres.enabled: true) on a temporary pod volume | External PostgreSQL with pgvector, reached through POSTGRES_URL |
| Images | A registry, or images loaded onto the node | A registry that every node can pull from, mirrored when the cluster has no internet access |
| Resources | No requests or limits | Requests and limits for every component |
| Workers | 1 replica | 3 replicas (the default) or more |
| Monitoring | Bundled Prometheus | Bundled Prometheus or the cluster's monitoring stack |
| Access | Port forward, with an ingress for the dashboard | Ingress or HTTPRoute with Transport Layer Security (TLS) |
The evaluation profile is for evaluation only. The bundled PostgreSQL stores data on a temporary pod volume, so deleting or rescheduling the PostgreSQL pod deletes every pipeline and document. The new database also has a new system ID, which requires a new license.
Kubernetes requirements
The charts declare no minimum Kubernetes version. The bundled Prometheus chart requires Kubernetes 1.19 or later. The charts use the following resources and features:
| Resource or feature | Used by | Required when |
|---|---|---|
| Jobs run as Helm hooks | Migration, license check and master key Jobs | Always |
| gRPC liveness probes (stable from Kubernetes 1.27) | gRPC service containers | Always |
| Default StorageClass | NATS JetStream and Prometheus volume claims | Always, unless the values name a storage class |
| ClusterRole and ClusterRoleBinding | Prometheus server service discovery | The bundled Prometheus is enabled (the default) |
| PodDisruptionBudget | NATS JetStream | Always |
| Image volumes | Model and package images | api-server.models or api-server.packages is set |
HorizontalPodAutoscaler autoscaling/v2 | API server and worker autoscaling | Autoscaling is enabled |
Ingress networking.k8s.io/v1 and an ingress controller | External access | ingress.enabled is true |
Gateway API gateway.networking.k8s.io/v1 and a Gateway | External access | httpRoute.enabled is true |
The following constraints also apply:
- Image volumes. Image volumes need support in both Kubernetes and the container runtime. The minimum Kubernetes version is being confirmed. Models, packages and air-gapped installs describes the model and package images.
- Node architecture. The Foundation4 images are built for the
amd64architecture. The API server, worker and dashboard pods run onamd64nodes. - Pod security. The Foundation4 charts set no pod security context by default, so a namespace that enforces the
restrictedPod Security Standard rejects the Foundation4 pods. Security hardening describes the settings. - Cluster permissions. The installing account creates a namespace, namespaced resources and, for the bundled Prometheus, a ClusterRole and a ClusterRoleBinding.
Storage requirements
The core release requests 38 GiB in 4 volume claims. The claims name no storage class unless the values set one, so the cluster needs a default StorageClass.
| Component | Volume claims | Size | Values keys |
|---|---|---|---|
| NATS JetStream | 3, one per replica | 10 GiB each | nats.config.jetstream.fileStore.pvc.size, nats.config.jetstream.fileStore.pvc.storageClassName |
| Prometheus server | 1 | 8 GiB | prometheus.server.persistentVolume.size, prometheus.server.persistentVolume.storageClass |
| Valkey | None | Not applicable | The cache keeps data in memory only |
| Bundled PostgreSQL | None | Not applicable | The bundled chart uses a temporary pod volume |
The NATS JetStream volumes hold documents that wait for a worker. Deleting a NATS JetStream volume claim deletes the queued documents.
PostgreSQL requirements
The production profile uses an external PostgreSQL database with the pgvector extension. PostgreSQL is the primary data store and holds every pipeline, document, fragment and vector.
- Version. The bundled chart and the Foundation4 test environment use PostgreSQL 18.
- pgvector. Version 0.8.0 or later. Similarity search sets the pgvector parameter
hnsw.iterative_scan, which pgvector 0.8.0 introduced. - Extension. The migration runs
CREATE EXTENSION IF NOT EXISTS vector. The extension exists in the database before installation, or the database user can create the extension. - Schemas. Foundation4 runs
CREATE SCHEMA IF NOT EXISTSat every start for the schemaspublicandembeddings, the defaults ofdatabase.schemaanddatabase.embeddings_schema. The database user holds theCREATEprivilege on the database. A database owned by the Foundation4 database user meets this requirement. - Connection. The secrets file carries the connection URL in
POSTGRES_URL.
Database describes the connection settings, TLS and the connection budget.
Operator tools
| Tool | Purpose | Required when |
|---|---|---|
kubectl | Namespace, Secrets, Job logs, port forward | Always |
| Helm 3 | Both releases | Always |
kubectl kustomize or kustomize | The Secret foundation4ai-secrets from the secrets file | Always |
openssl and uuidgen | Passwords, secrets and the master key identifier | Always |
curl | Access validation | Always |
just, kustomize and yq | The recipes in the .justfile of the deployment files | The recipes are used |
The .justfile checks for kubectl, kustomize, helm and yq when just loads the file, so every recipe fails when one of the four tools is missing from the PATH.
Sizing
No measured sizing figures are available. The charts establish the following starting points:
- Resource settings. The charts set no resource requests or limits for any component.
- Shared resources value. The value
api-server.resourcessets the requests and limits of the API server container, the worker container and both gRPC service containers. The value is sized for the largest of the four containers. - Dashboard resources. The value
dashboard.resourceshas no effect in chart version 0.3.0. - Replicas. 1 API server, 3 workers and 1 dashboard. Autoscaling is disabled by default; when enabled, the API server scales from 1 to 100 replicas and the workers from 3 to 100, at a CPU target of 80 percent.
- Worker batches. Each worker takes documents from the queue in batches of up to 100.
- Queue retention. NATS JetStream keeps a queued document for up to 24 hours. A document that waits longer than 24 hours is removed from the queue without processing, so the workers are sized to keep the queue shorter than 24 hours of submissions.
- Database connections. Each API server replica opens up to 3 database connections, and each worker opens 1.
Scaling and performance describes replica counts, autoscaling and the queue.