Skip to main content

Deployment overview and requirements

This page describes what a Foundation4 installation deploys on Kubernetes, the order in which the parts are installed, the two deployment profiles and the requirements for the cluster, the database and the operator's workstation. Operators read this page before planning an installation. Security reviewers use the requirements to assess the footprint of Foundation4 in a cluster. The procedure itself is described in Install on Kubernetes.

Components​

A Foundation4 installation consists of two Helm releases in the namespace foundation4ai. The core release installs the services that Foundation4 depends on. The application release installs Foundation4 and reads the connection settings that the core release produces.

ReleaseChartInstalls
foundation4ai-corefoundation4ai-core 0.3.0NATS JetStream (3 replicas and a NATS utility pod), Valkey as the Redis-compatible cache, the Prometheus server, optionally PostgreSQL with pgvector, and the Secret foundation4ai-core with the connection URLs, the license, the application secret and the master key
foundation4aifoundation4ai 0.3.0The API server (1 replica), the workers (3 replicas), the dashboard (1 replica), three installation Jobs, and optionally an Ingress or an HTTPRoute

The releases have the following characteristics:

  • Sidecar containers. The API server pod and each worker pod run two containers: the Foundation4 process and the gRPC service. The gRPC service is deployed as a sidecar container, a second container in the same Kubernetes pod, and hosts the Python text splitters and embedding providers.
  • Installation Jobs. The application release runs the database migration, the license check and the creation of the master key as Helm hook Jobs before the Deployments start. All three Jobs run the API server image.
  • Fixed names. The application release reads the Secret foundation4ai-core by name. The core release is therefore always named foundation4ai-core and installed in the same namespace as the application release.
  • Rendering against the cluster. Both charts read Secrets from the cluster while Helm renders the templates. Installation therefore uses the Helm command line with access to the cluster. Tools that render charts offline, such as helm template or some GitOps controllers, fail to render the charts or produce Secrets without credentials.

Installation order​

The installation follows a fixed order, because each step consumes the output of the step before:

  1. The operator creates the Secret foundation4ai-secrets from the secrets file foundation4ai.secrets.env.
  2. The core release installs NATS JetStream, Valkey and Prometheus, and copies the connection URLs, the license, the application secret and the master key into the Secret foundation4ai-core.
  3. A first installation of the application release migrates the database. The license check Job then writes the system ID to the Job log and waits, so this installation stops at the Helm timeout.
  4. The operator sends the system ID to the Foundation4 provider and receives the license. The operator adds the license to the secrets file, updates the Secret and upgrades the core release, which copies the license into the Secret foundation4ai-core.
  5. The operator installs the application release again. The migration, license check and master key Jobs complete, and the API server, workers and dashboard start.
  6. The operator signs in with the master key and creates one API key for each client application. The master key stays with the operators and serves only for administration.

The license is bound to the system ID, which Foundation4 derives from the database catalog after the migration creates the schema. Licensing describes the system ID and the license.

Solid arrows show actions and requests; dashed arrows show information returned to the operator.

Later changes, such as new image tags or a renewed license, use helm upgrade of the core release and then of the application release.

Deployment profiles​

Foundation4 supports two deployment profiles. Install for evaluation installs the evaluation profile; Install on Kubernetes installs the production profile.

AspectEvaluationProduction
NodesOne node, such as a single-node k3s clusterSeveral nodes, sized for the workload
PostgreSQLBundled chart (postgres.enabled: true) on a temporary pod volumeExternal PostgreSQL with pgvector, reached through POSTGRES_URL
ImagesA registry, or images loaded onto the nodeA registry that every node can pull from, mirrored when the cluster has no internet access
ResourcesNo requests or limitsRequests and limits for every component
Workers1 replica3 replicas (the default) or more
MonitoringBundled PrometheusBundled Prometheus or the cluster's monitoring stack
AccessPort forward, with an ingress for the dashboardIngress or HTTPRoute with Transport Layer Security (TLS)

The evaluation profile is for evaluation only. The bundled PostgreSQL stores data on a temporary pod volume, so deleting or rescheduling the PostgreSQL pod deletes every pipeline and document. The new database also has a new system ID, which requires a new license.

Kubernetes requirements​

The charts declare no minimum Kubernetes version. The bundled Prometheus chart requires Kubernetes 1.19 or later. The charts use the following resources and features:

Resource or featureUsed byRequired when
Jobs run as Helm hooksMigration, license check and master key JobsAlways
gRPC liveness probes (stable from Kubernetes 1.27)gRPC service containersAlways
Default StorageClassNATS JetStream and Prometheus volume claimsAlways, unless the values name a storage class
ClusterRole and ClusterRoleBindingPrometheus server service discoveryThe bundled Prometheus is enabled (the default)
PodDisruptionBudgetNATS JetStreamAlways
Image volumesModel and package imagesapi-server.models or api-server.packages is set
HorizontalPodAutoscaler autoscaling/v2API server and worker autoscalingAutoscaling is enabled
Ingress networking.k8s.io/v1 and an ingress controllerExternal accessingress.enabled is true
Gateway API gateway.networking.k8s.io/v1 and a GatewayExternal accesshttpRoute.enabled is true

The following constraints also apply:

  • Image volumes. Image volumes need support in both Kubernetes and the container runtime. The minimum Kubernetes version is being confirmed. Models, packages and air-gapped installs describes the model and package images.
  • Node architecture. The Foundation4 images are built for the amd64 architecture. The API server, worker and dashboard pods run on amd64 nodes.
  • Pod security. The Foundation4 charts set no pod security context by default, so a namespace that enforces the restricted Pod Security Standard rejects the Foundation4 pods. Security hardening describes the settings.
  • Cluster permissions. The installing account creates a namespace, namespaced resources and, for the bundled Prometheus, a ClusterRole and a ClusterRoleBinding.

Storage requirements​

The core release requests 38 GiB in 4 volume claims. The claims name no storage class unless the values set one, so the cluster needs a default StorageClass.

ComponentVolume claimsSizeValues keys
NATS JetStream3, one per replica10 GiB eachnats.config.jetstream.fileStore.pvc.size, nats.config.jetstream.fileStore.pvc.storageClassName
Prometheus server18 GiBprometheus.server.persistentVolume.size, prometheus.server.persistentVolume.storageClass
ValkeyNoneNot applicableThe cache keeps data in memory only
Bundled PostgreSQLNoneNot applicableThe bundled chart uses a temporary pod volume

The NATS JetStream volumes hold documents that wait for a worker. Deleting a NATS JetStream volume claim deletes the queued documents.

PostgreSQL requirements​

The production profile uses an external PostgreSQL database with the pgvector extension. PostgreSQL is the primary data store and holds every pipeline, document, fragment and vector.

  • Version. The bundled chart and the Foundation4 test environment use PostgreSQL 18.
  • pgvector. Version 0.8.0 or later. Similarity search sets the pgvector parameter hnsw.iterative_scan, which pgvector 0.8.0 introduced.
  • Extension. The migration runs CREATE EXTENSION IF NOT EXISTS vector. The extension exists in the database before installation, or the database user can create the extension.
  • Schemas. Foundation4 runs CREATE SCHEMA IF NOT EXISTS at every start for the schemas public and embeddings, the defaults of database.schema and database.embeddings_schema. The database user holds the CREATE privilege on the database. A database owned by the Foundation4 database user meets this requirement.
  • Connection. The secrets file carries the connection URL in POSTGRES_URL.

Database describes the connection settings, TLS and the connection budget.

Operator tools​

ToolPurposeRequired when
kubectlNamespace, Secrets, Job logs, port forwardAlways
Helm 3Both releasesAlways
kubectl kustomize or kustomizeThe Secret foundation4ai-secrets from the secrets fileAlways
openssl and uuidgenPasswords, secrets and the master key identifierAlways
curlAccess validationAlways
just, kustomize and yqThe recipes in the .justfile of the deployment filesThe recipes are used

The .justfile checks for kubectl, kustomize, helm and yq when just loads the file, so every recipe fails when one of the four tools is missing from the PATH.

Sizing​

No measured sizing figures are available. The charts establish the following starting points:

  • Resource settings. The charts set no resource requests or limits for any component.
  • Shared resources value. The value api-server.resources sets the requests and limits of the API server container, the worker container and both gRPC service containers. The value is sized for the largest of the four containers.
  • Dashboard resources. The value dashboard.resources has no effect in chart version 0.3.0.
  • Replicas. 1 API server, 3 workers and 1 dashboard. Autoscaling is disabled by default; when enabled, the API server scales from 1 to 100 replicas and the workers from 3 to 100, at a CPU target of 80 percent.
  • Worker batches. Each worker takes documents from the queue in batches of up to 100.
  • Queue retention. NATS JetStream keeps a queued document for up to 24 hours. A document that waits longer than 24 hours is removed from the queue without processing, so the workers are sized to keep the queue shorter than 24 hours of submissions.
  • Database connections. Each API server replica opens up to 3 database connections, and each worker opens 1.

Scaling and performance describes replica counts, autoscaling and the queue.