Production readiness checklist
This checklist confirms that a Foundation4 deployment is ready for production use. The checks are grouped by area: database, secrets, license, capacity, observability, security and availability. Each check names the command or setting that verifies the check, or the page that describes the check, and the expected result. Operators run the checklist before the first production use and after each upgrade, and security reviewers use the checklist as evidence of the configuration. The commands use the namespace and release names of Install on Kubernetes, and Operator command reference explains each command.
Database
| Check | Command or setting | Expected result |
|---|---|---|
| External PostgreSQL in use | kubectl get pods -n foundation4ai -l app.kubernetes.io/name=postgres | No resources found in foundation4ai namespace. The bundled PostgreSQL of the evaluation profile is not persistent. |
| Database reachable | kubectl exec -n foundation4ai deploy/foundation4ai-api-server -c server -- ./foundation4ai database check-connection | Ok. |
| TLS with certificate checks | grep -o 'sslmode=[a-z-]*' foundation4ai.secrets.env, and database.ca_cert in api-server.configs, as in Security hardening | sslmode=verify-full, and every Foundation4 connection shows ssl as t in pg_stat_ssl |
| pgvector installed | psql "<administrator-connection-url>" -c "SELECT extversion FROM pg_extension WHERE extname = 'vector';" | One row with the version that Database requires |
| Connection budget | SHOW max_connections; and the budget in Database | API server replicas multiplied by database.pool_size, plus 1 per worker, at the maximum replica counts of any autoscaler, stays below max_connections with room for administration sessions |
| Backups and restore test | Backup, restore and upgrades | Scheduled backups, and one completed restore test that includes the license step for a restored database |
Secrets
| Check | Command or setting | Expected result |
|---|---|---|
| Template passwords replaced | grep -Fx -f foundation4ai.secrets.env.template foundation4ai.secrets.env | No printed line contains PASSWORD. Comments, blank lines and unchanged names may be printed. |
| Secrets file outside version control | git check-ignore -v foundation4ai.secrets.env, in a Git working copy of the deployment files | The rule /*.env of .gitignore is printed |
| Recovery copies | The organization's secret store, as described in Secrets and keys | The store holds the current FOUNDATION4AI_APP_SECRET, FOUNDATION4AI_APP_MASTER_KEY, FOUNDATION4AI_APP_MASTER_SECRET and FOUNDATION4AI_APP_LICENSE |
| Secret access restricted | kubectl auth can-i get secrets -n foundation4ai --as <user> for a user outside the operator group | no |
| Encryption at rest | Encryption configuration of the Kubernetes API server, or the key management option of the platform | Secrets are encrypted at rest |
License
| Check | Command or setting | Expected result |
|---|---|---|
| License valid | kubectl logs -n foundation4ai job/foundation4ai-api-server-license-check | Checking for a valid license... OK |
| Expiry recorded | license details, as described in Operator command reference | The expiry date, or (perpetual), is recorded, with a renewal reminder ahead of the date in the operations calendar |
| Limits known | The Limits: lines of license details | The planned numbers of pipelines, documents and other objects stay below the limits, as described in Licensing |
Capacity
| Check | Command or setting | Expected result |
|---|---|---|
| Container resources | kubectl get pods -n foundation4ai -o custom-columns=NAME:.metadata.name,MEMORY:.spec.containers[*].resources.limits.memory | A memory limit for every container of the API server, worker and dashboard pods, as described in Scaling and performance. The dashboard chart does not read dashboard.resources; Values reference describes how the dashboard container receives a limit. |
| Storage class and sizes | kubectl get pvc -n foundation4ai -o custom-columns=NAME:.metadata.name,CLASS:.spec.storageClassName,SIZE:.spec.resources.requests.storage | The NATS JetStream and Prometheus claims use the planned storage class and sizes |
| JetStream storage headroom | kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats account info | Storage in use leaves room for 24 hours of submitted documents at peak rate |
| Worker count | kubectl get deployment foundation4ai-api-server-worker -n foundation4ai | READY equals the planned count, 3 by default |
| Connection budget | See the database checks | Met |
Observability
| Check | Command or setting | Expected result |
|---|---|---|
| Worker metrics collected | A Prometheus query for foundation4ai_document_queue_processed, as described in Monitoring and logging | One series per worker pod. The worker pod annotations name port 8000, while the worker serves metrics on port 9090, so the query returns no series until the scrape configuration reads port 9090. |
| Alert on failed documents | An alert rule on the growth of foundation4ai_document_queue_failed | The rule is loaded and fires in a test |
| Alert on queue age | A scheduled check of the first message time in nats stream info DOCUMENTS, or a metric of a NATS exporter | An alert fires well before a job reaches the 24-hour queue limit |
| Availability probe | An external check that sends POST /login through the ingress with a dedicated key | Status 201. /healthz does not reflect the state of the components. |
| Log collection | The log collector of the cluster | Logs of the server, worker and grpc containers are kept outside the pods |
| License expiry alert | The expiry date from the license checks | A reminder or alert before the expiry date |
Security
| Check | Command or setting | Expected result |
|---|---|---|
| TLS at the ingress or Gateway | kubectl get ingress foundation4ai -n foundation4ai -o jsonpath='{.spec.tls}', or the Gateway listener | A TLS entry for the host name, and 201 from the request in step 3 of the ingress procedure in Security hardening |
| Least-privilege API keys | The permission review of Manage API keys and permissions | One key per client application with permissions on individual objects, and no client application that uses the master key |
| Log level | helm get values foundation4ai -n foundation4ai | No log_level value other than warn |
| Network policies | kubectl get networkpolicy -n foundation4ai | The policies of Security hardening, including default-deny |
| Pod Security Admission | kubectl get namespace foundation4ai --show-labels | pod-security.kubernetes.io/enforce=baseline |
| Service account tokens | kubectl get serviceaccounts -n foundation4ai -o custom-columns=NAME:.metadata.name,AUTOMOUNT:.automountServiceAccountToken | false for default, foundation4ai-api-server and foundation4ai-dashboard |
| Dashboard exposure | dashboard.enabled in the application values | false when no operator uses the dashboard |
| Model servers | The endpoint of each embedding model and LLM | Internal addresses only, reachable from the Foundation4 pods only |
Availability
| Check | Command or setting | Expected result |
|---|---|---|
| API server replicas | kubectl get deployment foundation4ai-api-server -n foundation4ai | READY of 2/2 or more |
| Disruption budgets | kubectl get pdb -n foundation4ai | foundation4ai-api-server, foundation4ai-api-server-worker and foundation4ai-core-nats, each with ALLOWED DISRUPTIONS of 1 |
| NATS JetStream spread | kubectl get pods -n foundation4ai -l app.kubernetes.io/component=nats -o wide | The 3 NATS pods run on 3 different nodes. nats.podTemplate.topologySpreadConstraints sets the spread when the scheduler does not. |
| Priority class | kubectl get pods -n foundation4ai -o custom-columns=NAME:.metadata.name,PRIORITY:.spec.priorityClassName | The class of Security hardening on the Foundation4 and core pods |
| Queue maintenance plan | kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats stream info DOCUMENTS | Replicas: 1, and a maintenance procedure for the node that holds the stream |
| Restart after changes | The change procedure in Secrets and keys | The operations runbook restarts the API server and the workers after each change of a Secret, the license or api-server.configs |