Skip to main content

Production readiness checklist

This checklist confirms that a Foundation4 deployment is ready for production use. The checks are grouped by area: database, secrets, license, capacity, observability, security and availability. Each check names the command or setting that verifies the check, or the page that describes the check, and the expected result. Operators run the checklist before the first production use and after each upgrade, and security reviewers use the checklist as evidence of the configuration. The commands use the namespace and release names of Install on Kubernetes, and Operator command reference explains each command.

Database​

CheckCommand or settingExpected result
External PostgreSQL in usekubectl get pods -n foundation4ai -l app.kubernetes.io/name=postgresNo resources found in foundation4ai namespace. The bundled PostgreSQL of the evaluation profile is not persistent.
Database reachablekubectl exec -n foundation4ai deploy/foundation4ai-api-server -c server -- ./foundation4ai database check-connectionOk.
TLS with certificate checksgrep -o 'sslmode=[a-z-]*' foundation4ai.secrets.env, and database.ca_cert in api-server.configs, as in Security hardeningsslmode=verify-full, and every Foundation4 connection shows ssl as t in pg_stat_ssl
pgvector installedpsql "<administrator-connection-url>" -c "SELECT extversion FROM pg_extension WHERE extname = 'vector';"One row with the version that Database requires
Connection budgetSHOW max_connections; and the budget in DatabaseAPI server replicas multiplied by database.pool_size, plus 1 per worker, at the maximum replica counts of any autoscaler, stays below max_connections with room for administration sessions
Backups and restore testBackup, restore and upgradesScheduled backups, and one completed restore test that includes the license step for a restored database

Secrets​

CheckCommand or settingExpected result
Template passwords replacedgrep -Fx -f foundation4ai.secrets.env.template foundation4ai.secrets.envNo printed line contains PASSWORD. Comments, blank lines and unchanged names may be printed.
Secrets file outside version controlgit check-ignore -v foundation4ai.secrets.env, in a Git working copy of the deployment filesThe rule /*.env of .gitignore is printed
Recovery copiesThe organization's secret store, as described in Secrets and keysThe store holds the current FOUNDATION4AI_APP_SECRET, FOUNDATION4AI_APP_MASTER_KEY, FOUNDATION4AI_APP_MASTER_SECRET and FOUNDATION4AI_APP_LICENSE
Secret access restrictedkubectl auth can-i get secrets -n foundation4ai --as <user> for a user outside the operator groupno
Encryption at restEncryption configuration of the Kubernetes API server, or the key management option of the platformSecrets are encrypted at rest

License​

CheckCommand or settingExpected result
License validkubectl logs -n foundation4ai job/foundation4ai-api-server-license-checkChecking for a valid license... OK
Expiry recordedlicense details, as described in Operator command referenceThe expiry date, or (perpetual), is recorded, with a renewal reminder ahead of the date in the operations calendar
Limits knownThe Limits: lines of license detailsThe planned numbers of pipelines, documents and other objects stay below the limits, as described in Licensing

Capacity​

CheckCommand or settingExpected result
Container resourceskubectl get pods -n foundation4ai -o custom-columns=NAME:.metadata.name,MEMORY:.spec.containers[*].resources.limits.memoryA memory limit for every container of the API server, worker and dashboard pods, as described in Scaling and performance. The dashboard chart does not read dashboard.resources; Values reference describes how the dashboard container receives a limit.
Storage class and sizeskubectl get pvc -n foundation4ai -o custom-columns=NAME:.metadata.name,CLASS:.spec.storageClassName,SIZE:.spec.resources.requests.storageThe NATS JetStream and Prometheus claims use the planned storage class and sizes
JetStream storage headroomkubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats account infoStorage in use leaves room for 24 hours of submitted documents at peak rate
Worker countkubectl get deployment foundation4ai-api-server-worker -n foundation4aiREADY equals the planned count, 3 by default
Connection budgetSee the database checksMet

Observability​

CheckCommand or settingExpected result
Worker metrics collectedA Prometheus query for foundation4ai_document_queue_processed, as described in Monitoring and loggingOne series per worker pod. The worker pod annotations name port 8000, while the worker serves metrics on port 9090, so the query returns no series until the scrape configuration reads port 9090.
Alert on failed documentsAn alert rule on the growth of foundation4ai_document_queue_failedThe rule is loaded and fires in a test
Alert on queue ageA scheduled check of the first message time in nats stream info DOCUMENTS, or a metric of a NATS exporterAn alert fires well before a job reaches the 24-hour queue limit
Availability probeAn external check that sends POST /login through the ingress with a dedicated keyStatus 201. /healthz does not reflect the state of the components.
Log collectionThe log collector of the clusterLogs of the server, worker and grpc containers are kept outside the pods
License expiry alertThe expiry date from the license checksA reminder or alert before the expiry date

Security​

CheckCommand or settingExpected result
TLS at the ingress or Gatewaykubectl get ingress foundation4ai -n foundation4ai -o jsonpath='{.spec.tls}', or the Gateway listenerA TLS entry for the host name, and 201 from the request in step 3 of the ingress procedure in Security hardening
Least-privilege API keysThe permission review of Manage API keys and permissionsOne key per client application with permissions on individual objects, and no client application that uses the master key
Log levelhelm get values foundation4ai -n foundation4aiNo log_level value other than warn
Network policieskubectl get networkpolicy -n foundation4aiThe policies of Security hardening, including default-deny
Pod Security Admissionkubectl get namespace foundation4ai --show-labelspod-security.kubernetes.io/enforce=baseline
Service account tokenskubectl get serviceaccounts -n foundation4ai -o custom-columns=NAME:.metadata.name,AUTOMOUNT:.automountServiceAccountTokenfalse for default, foundation4ai-api-server and foundation4ai-dashboard
Dashboard exposuredashboard.enabled in the application valuesfalse when no operator uses the dashboard
Model serversThe endpoint of each embedding model and LLMInternal addresses only, reachable from the Foundation4 pods only

Availability​

CheckCommand or settingExpected result
API server replicaskubectl get deployment foundation4ai-api-server -n foundation4aiREADY of 2/2 or more
Disruption budgetskubectl get pdb -n foundation4aifoundation4ai-api-server, foundation4ai-api-server-worker and foundation4ai-core-nats, each with ALLOWED DISRUPTIONS of 1
NATS JetStream spreadkubectl get pods -n foundation4ai -l app.kubernetes.io/component=nats -o wideThe 3 NATS pods run on 3 different nodes. nats.podTemplate.topologySpreadConstraints sets the spread when the scheduler does not.
Priority classkubectl get pods -n foundation4ai -o custom-columns=NAME:.metadata.name,PRIORITY:.spec.priorityClassNameThe class of Security hardening on the Foundation4 and core pods
Queue maintenance plankubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats stream info DOCUMENTSReplicas: 1, and a maintenance procedure for the node that holds the stream
Restart after changesThe change procedure in Secrets and keysThe operations runbook restarts the API server and the workers after each change of a Secret, the license or api-server.configs