Backup, restore and upgrades
This page describes what an operator backs up for a Foundation4 installation, how a backup is restored, and how Foundation4 is upgraded and rolled back. Operators and database administrators who plan recovery and maintenance use this page.
The commands on this page read the production values file foundation4ai.values.yaml, created in Install on Kubernetes, and read database connection URLs from shell variables: DATABASE_URL for the Foundation4 database and RESTORE_DATABASE_URL for a restore target. An evaluation installation uses evaluation.values.yaml and the kubectl exec forms of the database commands.
Backup set
The persistent state of an installation is the PostgreSQL database together with the secrets and settings that match the database. Every other component holds data that Foundation4 rebuilds or that expires.
| Item | Location | Backed up | Reason |
|---|---|---|---|
PostgreSQL database, schemas public and embeddings | PostgreSQL | Yes | Primary data store |
| Application secret | FOUNDATION4AI_APP_SECRET in the secrets file | Yes, with each database backup | Decrypts the LLM API keys stored in the database |
| Master key identifier and secret | Secrets file | Yes, with each database backup | Must match the master key stored in the database |
Secrets file and kustomization.yaml | Operator workstation | Yes, in the organization's secret store | Recreate foundation4ai-secrets |
| Values files | Operator workstation or version control | Yes | Recreate both releases |
| Chart version and image tags | Values files and Helm history | Recorded | Reinstall the same version |
| License | Secrets file | Recorded | Valid only for the system ID of the database |
| NATS JetStream volumes | Volume claims of the core release | No | Queued jobs and document text expire after 24 hours |
| Redis-compatible cache | Valkey memory | No | Rebuilt from the database |
| Prometheus data | Volume claim of the core release | Optional | Metrics history only |
Each database backup is paired with the secrets that were in effect when the backup was taken. The secrets are stored separately from the backup files, because the application secret protects the LLM API keys inside the backup. Secrets and keys describes the application secret.
Database backup
A production database is backed up with the tools of the database service: physical backups, storage snapshots or point-in-time recovery, and logical dumps with pg_dump. The backup type decides the license after a restore:
- Physical backups and snapshots. A physical restore keeps the PostgreSQL catalog and is therefore expected to keep the system ID and the license. This expectation has not been tested.
- Logical dumps. A restore of a
pg_dumpfile creates new catalog entries, so the restored database has a new system ID and needs a new license.
The following command writes a logical dump of a production database:
pg_dump --format=custom --no-comments \
--file="foundation4ai-$(date +%Y%m%d%H%M).dump" "$DATABASE_URL"
Expected result: the command prints nothing, exits with status 0, and the dump file exists. pg_restore --list <dump file> | grep "SCHEMA - embeddings" prints one line.
The following command writes a logical dump of the bundled database of an evaluation installation to the operator workstation:
kubectl exec -n foundation4ai foundation4ai-core-postgres-0 -c postgres -- \
pg_dump --format=custom --no-comments -U foundation4ai -d foundation4ai \
> foundation4ai-evaluation.dump
Expected result: the command prints nothing, and foundation4ai-evaluation.dump is larger than zero bytes.
pg_dump reads one consistent snapshot of the database. Documents that are pending when the dump starts remain pending after a restore, because the queued jobs are not part of the backup.
Logical restore
A logical restore replaces the database of an installation with a dump. The procedure applies to a restore into a new, empty database prepared as described in Database preparation. The secrets file holds the application secret and the master key of the installation that produced the dump.
-
Stop the API server and the workers:
kubectl scale -n foundation4ai deployment/foundation4ai-api-server \deployment/foundation4ai-api-server-worker --replicas=0Expected result:
deployment.apps/foundation4ai-api-server scaledanddeployment.apps/foundation4ai-api-server-worker scaled. -
Restore the dump and check the result:
pg_restore --no-owner --no-privileges --exit-on-error \--dbname="$RESTORE_DATABASE_URL" <dump file>psql "$RESTORE_DATABASE_URL" -c 'SELECT count(*) FROM pipelines'Expected result:
pg_restoreprints nothing, and the pipeline count equals the count of the source installation. -
Set
POSTGRES_URLinfoundation4ai.secrets.envto the restored database, apply the Secret and upgrade the core release:kubectl kustomize . | kubectl apply -f -helm upgrade foundation4ai-core ./charts/foundation4ai-core \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result:
secret/foundation4ai-secrets configured, and Helm reportsSTATUS: deployed. -
Run the application upgrade to obtain the new system ID:
helm upgrade foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --timeout 3mkubectl logs -n foundation4ai job/foundation4ai-api-server-license-checkExpected result: after 3 minutes, Helm reports that the pre-upgrade hook failed. The log shows
SystemID: <system ID>andChecking for a valid license... Invalid. -
Obtain a license for the new system ID from the Foundation4 provider. Set the license in
foundation4ai.secrets.env, then apply the Secret, upgrade the core release, delete the waiting Job and upgrade the application release:kubectl kustomize . | kubectl apply -f -helm upgrade foundation4ai-core ./charts/foundation4ai-core \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mkubectl delete job -n foundation4ai foundation4ai-api-server-license-checkhelm upgrade foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployedfor both releases, and the license check log ends withChecking for a valid license... OK. -
Clear the cache, which holds API key verification results and objects from before the restore:
kubectl rollout restart -n foundation4ai deployment/foundation4ai-core-rediskubectl rollout status -n foundation4ai deployment/foundation4ai-core-redisExpected result:
deployment "foundation4ai-core-redis" successfully rolled out. An external Redis-compatible cache is cleared with the tools of that service instead. -
Start the API server and the workers with the replica counts of the values file. The chart defaults are 1 and 3:
kubectl scale -n foundation4ai deployment/foundation4ai-api-server --replicas=1kubectl scale -n foundation4ai deployment/foundation4ai-api-server-worker --replicas=3kubectl rollout status -n foundation4ai deployment/foundation4ai-api-serverkubectl rollout status -n foundation4ai deployment/foundation4ai-api-server-workerExpected result: both Deployments report
successfully rolled out.
After the restore, the client application submits again the documents that remain pending, as described in Ingest documents reliably.
A restore into a new cluster installs both releases as described in Install on Kubernetes, with POSTGRES_URL pointing to the restored database and the secrets of the source installation.
Evaluation restore
The bundled database of the evaluation profile is recreated empty whenever the PostgreSQL pod is recreated, so a restore uses a new pod as the empty database:
-
Remove the application release and recreate the PostgreSQL pod:
helm uninstall foundation4ai -n foundation4aikubectl delete pod -n foundation4ai foundation4ai-core-postgres-0kubectl wait -n foundation4ai --for=condition=Ready \pod/foundation4ai-core-postgres-0 --timeout=5mExpected result:
release "foundation4ai" uninstalled,pod "foundation4ai-core-postgres-0" deletedandpod/foundation4ai-core-postgres-0 condition met. -
Restore the dump into the new database:
kubectl exec -i -n foundation4ai foundation4ai-core-postgres-0 -c postgres -- \pg_restore --no-owner --no-privileges --exit-on-error \-U foundation4ai -d foundation4ai < foundation4ai-evaluation.dumpExpected result: the command prints nothing and exits with status 0.
-
Install the application release and the license as in steps 4 to 6 of Install for evaluation. The restored database has a new system ID.
Expected result: in step 4, the license check log shows a system ID that differs from the system ID of the previous license, followed by
Invalid; after step 6, the license check log ends withChecking for a valid license... OK.
Upgrade behavior
An upgrade installs new chart versions or image tags with helm upgrade, first for the core release and then for the application release. The application upgrade runs in the following order:
- Helm writes the ConfigMap and the Secret of the application release.
- The Job
foundation4ai-api-server-db-migrationapplies the new migrations. - The Job
foundation4ai-api-server-license-checkchecks the license. - The Job
foundation4ai-api-server-create-admin-api-keyconfirms the master key. - Helm updates the Deployments, and Kubernetes replaces the pods with a rolling update.
The following consequences apply:
- Old pods during migration. The pods of the previous version keep serving while the Jobs run, against a database that the migration is changing. The upgrade therefore runs in a maintenance window in which client applications do not write.
- Migration duration. A migration that rewrites fragment tables takes time in proportion to the number of fragments. The Helm timeout covers the migration, so large databases use a longer timeout than 10 minutes.
- Forward only. Migrations cannot be reverted, and the previous version does not start against a migrated database. Rollback describes the consequences.
- Evaluation profile. A core release upgrade that changes the PostgreSQL pod, for example a new chart version, recreates the bundled database empty.
- Configuration changes. An upgrade that changes image tags restarts the pods. An upgrade that changes only Secrets or configuration files needs a restart, as described in Secrets and keys.
The Release notes list the changes of each version.
Pre-upgrade checks
-
Record the release history and the values of both releases:
helm history foundation4ai -n foundation4aihelm history foundation4ai-core -n foundation4aihelm get values foundation4ai -n foundation4ai > foundation4ai.previous-values.yamlExpected result: the last revision of each release has the status
deployed, and the values file is written. -
Back up the database and record the secrets, as described in Database backup.
Expected result: a dump file or a database service backup taken after client writes stopped.
-
List the applied migrations:
psql "$DATABASE_URL" -c 'SELECT version FROM schema_migrations ORDER BY version'Expected result: one row for each applied migration, starting with
m20250120_030022_initial_migration. -
When the list does not contain
m20260827_140000_vector_embeddings, check every pipeline for duplicate fragment rows, which stop that migration:cat > duplicate-check.sql <<'SQL'DO $$DECLAREp record;n bigint;BEGINFOR p IN SELECT id FROM pipelines LOOPEXECUTE format('SELECT count(*) FROM (SELECT 1 FROM embeddings.%I GROUP BY document_id, version, order_id HAVING count(*) > 1) AS d',p.id::text) INTO n;RAISE NOTICE 'pipeline %: % duplicate groups', p.id, n;END LOOP;END $$;SQLpsql "$DATABASE_URL" -f duplicate-check.sqlExpected result: one line
NOTICE: pipeline <pipeline ID>: 0 duplicate groupsfor each pipeline, followed byDO. A pipeline with duplicate groups stops the upgrade; the operator sends the pipeline identifiers and counts to the Foundation4 provider before upgrading. An evaluation installation runs the file withkubectl exec -i -n foundation4ai foundation4ai-core-postgres-0 -c postgres -- psql -U foundation4ai -d foundation4ai < duplicate-check.sql.
Upgrade procedure
-
Set the new image tags, and any new values, in
foundation4ai.values.yaml.Expected result:
grep ImageTag foundation4ai.values.yamlshows the new tags. -
Upgrade the core release:
helm upgrade foundation4ai-core ./charts/foundation4ai-core \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployed, andkubectl get pods -n foundation4aishows the NATS, Valkey and Prometheus podsRunning. -
Upgrade the application release, with a timeout that covers the migration:
helm upgrade foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 30mExpected result: Helm reports
STATUS: deployed. -
Verify the migration and the pods:
kubectl logs -n foundation4ai job/foundation4ai-api-server-db-migrationpsql "$DATABASE_URL" -c 'SELECT version FROM schema_migrations ORDER BY version'kubectl get pods -n foundation4aiExpected result: the migration log shows
Migrated to latest schema., the migration list contains the migrations of the new version, and the API server and worker pods show2/2containers ready.POST /loginwith the master key returns status 201, as described in API and dashboard access.
An upgrade never uses helm uninstall followed by helm install. An uninstall keeps the database but loses the release history that a rollback needs, and an evaluation installation loses the bundled database.
Rollback
helm rollback restores the manifests of an earlier revision, such as the previous image tags. The charts define no rollback hooks, so a rollback runs no Job and reverts no migration. The safe rollback path depends on how far the upgrade progressed:
| Situation | Rollback path |
|---|---|
| The upgrade stopped before the migration Job applied a new migration, for example at an image pull error or the license check | helm rollback foundation4ai <revision> -n foundation4ai, or correct the cause and upgrade again |
| The migration Job failed | The failed migration left no changes. Correct the cause and upgrade again. When the same upgrade applied earlier migrations, the running pods keep serving, but a restarted pod of the previous version does not start. |
| The upgrade completed | Restore the pre-upgrade backup, then run helm rollback. A rollback without the restore leaves the API server and the workers unable to start. |
The migration list from the pre-upgrade checks shows which migrations the upgrade applied. A logical restore needs a new license, as described in Logical restore.
Troubleshooting
Upgrade failing on duplicate fragment rows
- Symptoms. The application upgrade stops at a failed pre-upgrade hook, and
kubectl logs -n foundation4ai job/foundation4ai-api-server-db-migrationshowsFailed applying migrationswithduplicate key value violates unique constraint. - Diagnosis. The duplicate check of the pre-upgrade checks names the pipelines with duplicate groups.
- Cause. The migration copies fragment rows into a new table with a unique index on the document, version and position of each fragment. Duplicate rows in a pipeline stop the copy.
- Resolution. The failed migration left no changes. Send the pipeline identifiers and counts to the Foundation4 provider, and keep the pods of the previous version running until the upgrade completes or the backup is restored.
- Actions to avoid. Deleting fragment rows by hand, and restarting the API server or the workers of the previous version, which do not start when the upgrade applied earlier migrations.
Pods failing after a rollback with a missing migration
- Symptoms. After
helm rollback, theserverandworkercontainers restart repeatedly. - Diagnosis.
kubectl logs -n foundation4ai deploy/foundation4ai-api-server -c server --previousshowsMigration file of version '<version>' is missing, this migration has been applied but its file is missing. - Cause. The database schema belongs to the newer version. Migrations cannot be reverted, and the previous version refuses to start against migrations that the previous version does not know.
- Resolution. Return to the newer version with
helm rollback foundation4ai <newer revision> -n foundation4ai, or restore the pre-upgrade backup and then roll back. A logical restore needs a new license. - Actions to avoid. Deleting rows from
schema_migrations. The next upgrade then applies the migrations again to a schema that already contains the changes.
Restored installation waiting at the license check
- Symptoms. After a restore, the application upgrade stops at the license check, and the API server does not start.
- Diagnosis.
kubectl logs -n foundation4ai job/foundation4ai-api-server-license-checkshows a system ID that differs from the system ID of the license. - Cause. The restored database has a new system ID.
- Resolution. Complete step 5 of Logical restore with a license for the new system ID.
- Actions to avoid. Restoring the dump again into another new database, which produces another system ID.