Skip to main content

Backup, restore and upgrades

This page describes what an operator backs up for a Foundation4 installation, how a backup is restored, and how Foundation4 is upgraded and rolled back. Operators and database administrators who plan recovery and maintenance use this page.

The commands on this page read the production values file foundation4ai.values.yaml, created in Install on Kubernetes, and read database connection URLs from shell variables: DATABASE_URL for the Foundation4 database and RESTORE_DATABASE_URL for a restore target. An evaluation installation uses evaluation.values.yaml and the kubectl exec forms of the database commands.

Backup set​

The persistent state of an installation is the PostgreSQL database together with the secrets and settings that match the database. Every other component holds data that Foundation4 rebuilds or that expires.

ItemLocationBacked upReason
PostgreSQL database, schemas public and embeddingsPostgreSQLYesPrimary data store
Application secretFOUNDATION4AI_APP_SECRET in the secrets fileYes, with each database backupDecrypts the LLM API keys stored in the database
Master key identifier and secretSecrets fileYes, with each database backupMust match the master key stored in the database
Secrets file and kustomization.yamlOperator workstationYes, in the organization's secret storeRecreate foundation4ai-secrets
Values filesOperator workstation or version controlYesRecreate both releases
Chart version and image tagsValues files and Helm historyRecordedReinstall the same version
LicenseSecrets fileRecordedValid only for the system ID of the database
NATS JetStream volumesVolume claims of the core releaseNoQueued jobs and document text expire after 24 hours
Redis-compatible cacheValkey memoryNoRebuilt from the database
Prometheus dataVolume claim of the core releaseOptionalMetrics history only

Each database backup is paired with the secrets that were in effect when the backup was taken. The secrets are stored separately from the backup files, because the application secret protects the LLM API keys inside the backup. Secrets and keys describes the application secret.

Database backup​

A production database is backed up with the tools of the database service: physical backups, storage snapshots or point-in-time recovery, and logical dumps with pg_dump. The backup type decides the license after a restore:

  • Physical backups and snapshots. A physical restore keeps the PostgreSQL catalog and is therefore expected to keep the system ID and the license. This expectation has not been tested.
  • Logical dumps. A restore of a pg_dump file creates new catalog entries, so the restored database has a new system ID and needs a new license.

The following command writes a logical dump of a production database:

pg_dump --format=custom --no-comments \
--file="foundation4ai-$(date +%Y%m%d%H%M).dump" "$DATABASE_URL"

Expected result: the command prints nothing, exits with status 0, and the dump file exists. pg_restore --list <dump file> | grep "SCHEMA - embeddings" prints one line.

The following command writes a logical dump of the bundled database of an evaluation installation to the operator workstation:

kubectl exec -n foundation4ai foundation4ai-core-postgres-0 -c postgres -- \
pg_dump --format=custom --no-comments -U foundation4ai -d foundation4ai \
> foundation4ai-evaluation.dump

Expected result: the command prints nothing, and foundation4ai-evaluation.dump is larger than zero bytes.

pg_dump reads one consistent snapshot of the database. Documents that are pending when the dump starts remain pending after a restore, because the queued jobs are not part of the backup.

Logical restore​

A logical restore replaces the database of an installation with a dump. The procedure applies to a restore into a new, empty database prepared as described in Database preparation. The secrets file holds the application secret and the master key of the installation that produced the dump.

  1. Stop the API server and the workers:

    kubectl scale -n foundation4ai deployment/foundation4ai-api-server \
    deployment/foundation4ai-api-server-worker --replicas=0

    Expected result: deployment.apps/foundation4ai-api-server scaled and deployment.apps/foundation4ai-api-server-worker scaled.

  2. Restore the dump and check the result:

    pg_restore --no-owner --no-privileges --exit-on-error \
    --dbname="$RESTORE_DATABASE_URL" <dump file>
    psql "$RESTORE_DATABASE_URL" -c 'SELECT count(*) FROM pipelines'

    Expected result: pg_restore prints nothing, and the pipeline count equals the count of the source installation.

  3. Set POSTGRES_URL in foundation4ai.secrets.env to the restored database, apply the Secret and upgrade the core release:

    kubectl kustomize . | kubectl apply -f -
    helm upgrade foundation4ai-core ./charts/foundation4ai-core \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: secret/foundation4ai-secrets configured, and Helm reports STATUS: deployed.

  4. Run the application upgrade to obtain the new system ID:

    helm upgrade foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --timeout 3m
    kubectl logs -n foundation4ai job/foundation4ai-api-server-license-check

    Expected result: after 3 minutes, Helm reports that the pre-upgrade hook failed. The log shows SystemID: <system ID> and Checking for a valid license... Invalid.

  5. Obtain a license for the new system ID from the Foundation4 provider. Set the license in foundation4ai.secrets.env, then apply the Secret, upgrade the core release, delete the waiting Job and upgrade the application release:

    kubectl kustomize . | kubectl apply -f -
    helm upgrade foundation4ai-core ./charts/foundation4ai-core \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m
    kubectl delete job -n foundation4ai foundation4ai-api-server-license-check
    helm upgrade foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed for both releases, and the license check log ends with Checking for a valid license... OK.

  6. Clear the cache, which holds API key verification results and objects from before the restore:

    kubectl rollout restart -n foundation4ai deployment/foundation4ai-core-redis
    kubectl rollout status -n foundation4ai deployment/foundation4ai-core-redis

    Expected result: deployment "foundation4ai-core-redis" successfully rolled out. An external Redis-compatible cache is cleared with the tools of that service instead.

  7. Start the API server and the workers with the replica counts of the values file. The chart defaults are 1 and 3:

    kubectl scale -n foundation4ai deployment/foundation4ai-api-server --replicas=1
    kubectl scale -n foundation4ai deployment/foundation4ai-api-server-worker --replicas=3
    kubectl rollout status -n foundation4ai deployment/foundation4ai-api-server
    kubectl rollout status -n foundation4ai deployment/foundation4ai-api-server-worker

    Expected result: both Deployments report successfully rolled out.

After the restore, the client application submits again the documents that remain pending, as described in Ingest documents reliably.

A restore into a new cluster installs both releases as described in Install on Kubernetes, with POSTGRES_URL pointing to the restored database and the secrets of the source installation.

Evaluation restore​

The bundled database of the evaluation profile is recreated empty whenever the PostgreSQL pod is recreated, so a restore uses a new pod as the empty database:

  1. Remove the application release and recreate the PostgreSQL pod:

    helm uninstall foundation4ai -n foundation4ai
    kubectl delete pod -n foundation4ai foundation4ai-core-postgres-0
    kubectl wait -n foundation4ai --for=condition=Ready \
    pod/foundation4ai-core-postgres-0 --timeout=5m

    Expected result: release "foundation4ai" uninstalled, pod "foundation4ai-core-postgres-0" deleted and pod/foundation4ai-core-postgres-0 condition met.

  2. Restore the dump into the new database:

    kubectl exec -i -n foundation4ai foundation4ai-core-postgres-0 -c postgres -- \
    pg_restore --no-owner --no-privileges --exit-on-error \
    -U foundation4ai -d foundation4ai < foundation4ai-evaluation.dump

    Expected result: the command prints nothing and exits with status 0.

  3. Install the application release and the license as in steps 4 to 6 of Install for evaluation. The restored database has a new system ID.

    Expected result: in step 4, the license check log shows a system ID that differs from the system ID of the previous license, followed by Invalid; after step 6, the license check log ends with Checking for a valid license... OK.

Upgrade behavior​

An upgrade installs new chart versions or image tags with helm upgrade, first for the core release and then for the application release. The application upgrade runs in the following order:

  1. Helm writes the ConfigMap and the Secret of the application release.
  2. The Job foundation4ai-api-server-db-migration applies the new migrations.
  3. The Job foundation4ai-api-server-license-check checks the license.
  4. The Job foundation4ai-api-server-create-admin-api-key confirms the master key.
  5. Helm updates the Deployments, and Kubernetes replaces the pods with a rolling update.

The following consequences apply:

  • Old pods during migration. The pods of the previous version keep serving while the Jobs run, against a database that the migration is changing. The upgrade therefore runs in a maintenance window in which client applications do not write.
  • Migration duration. A migration that rewrites fragment tables takes time in proportion to the number of fragments. The Helm timeout covers the migration, so large databases use a longer timeout than 10 minutes.
  • Forward only. Migrations cannot be reverted, and the previous version does not start against a migrated database. Rollback describes the consequences.
  • Evaluation profile. A core release upgrade that changes the PostgreSQL pod, for example a new chart version, recreates the bundled database empty.
  • Configuration changes. An upgrade that changes image tags restarts the pods. An upgrade that changes only Secrets or configuration files needs a restart, as described in Secrets and keys.

The Release notes list the changes of each version.

Pre-upgrade checks​

  1. Record the release history and the values of both releases:

    helm history foundation4ai -n foundation4ai
    helm history foundation4ai-core -n foundation4ai
    helm get values foundation4ai -n foundation4ai > foundation4ai.previous-values.yaml

    Expected result: the last revision of each release has the status deployed, and the values file is written.

  2. Back up the database and record the secrets, as described in Database backup.

    Expected result: a dump file or a database service backup taken after client writes stopped.

  3. List the applied migrations:

    psql "$DATABASE_URL" -c 'SELECT version FROM schema_migrations ORDER BY version'

    Expected result: one row for each applied migration, starting with m20250120_030022_initial_migration.

  4. When the list does not contain m20260827_140000_vector_embeddings, check every pipeline for duplicate fragment rows, which stop that migration:

    cat > duplicate-check.sql <<'SQL'
    DO $$
    DECLARE
    p record;
    n bigint;
    BEGIN
    FOR p IN SELECT id FROM pipelines LOOP
    EXECUTE format(
    'SELECT count(*) FROM (SELECT 1 FROM embeddings.%I GROUP BY document_id, version, order_id HAVING count(*) > 1) AS d',
    p.id::text) INTO n;
    RAISE NOTICE 'pipeline %: % duplicate groups', p.id, n;
    END LOOP;
    END $$;
    SQL
    psql "$DATABASE_URL" -f duplicate-check.sql

    Expected result: one line NOTICE: pipeline <pipeline ID>: 0 duplicate groups for each pipeline, followed by DO. A pipeline with duplicate groups stops the upgrade; the operator sends the pipeline identifiers and counts to the Foundation4 provider before upgrading. An evaluation installation runs the file with kubectl exec -i -n foundation4ai foundation4ai-core-postgres-0 -c postgres -- psql -U foundation4ai -d foundation4ai < duplicate-check.sql.

Upgrade procedure​

  1. Set the new image tags, and any new values, in foundation4ai.values.yaml.

    Expected result: grep ImageTag foundation4ai.values.yaml shows the new tags.

  2. Upgrade the core release:

    helm upgrade foundation4ai-core ./charts/foundation4ai-core \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed, and kubectl get pods -n foundation4ai shows the NATS, Valkey and Prometheus pods Running.

  3. Upgrade the application release, with a timeout that covers the migration:

    helm upgrade foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 30m

    Expected result: Helm reports STATUS: deployed.

  4. Verify the migration and the pods:

    kubectl logs -n foundation4ai job/foundation4ai-api-server-db-migration
    psql "$DATABASE_URL" -c 'SELECT version FROM schema_migrations ORDER BY version'
    kubectl get pods -n foundation4ai

    Expected result: the migration log shows Migrated to latest schema., the migration list contains the migrations of the new version, and the API server and worker pods show 2/2 containers ready. POST /login with the master key returns status 201, as described in API and dashboard access.

An upgrade never uses helm uninstall followed by helm install. An uninstall keeps the database but loses the release history that a rollback needs, and an evaluation installation loses the bundled database.

Rollback​

helm rollback restores the manifests of an earlier revision, such as the previous image tags. The charts define no rollback hooks, so a rollback runs no Job and reverts no migration. The safe rollback path depends on how far the upgrade progressed:

SituationRollback path
The upgrade stopped before the migration Job applied a new migration, for example at an image pull error or the license checkhelm rollback foundation4ai <revision> -n foundation4ai, or correct the cause and upgrade again
The migration Job failedThe failed migration left no changes. Correct the cause and upgrade again. When the same upgrade applied earlier migrations, the running pods keep serving, but a restarted pod of the previous version does not start.
The upgrade completedRestore the pre-upgrade backup, then run helm rollback. A rollback without the restore leaves the API server and the workers unable to start.

The migration list from the pre-upgrade checks shows which migrations the upgrade applied. A logical restore needs a new license, as described in Logical restore.

Troubleshooting​

Upgrade failing on duplicate fragment rows​

  • Symptoms. The application upgrade stops at a failed pre-upgrade hook, and kubectl logs -n foundation4ai job/foundation4ai-api-server-db-migration shows Failed applying migrations with duplicate key value violates unique constraint.
  • Diagnosis. The duplicate check of the pre-upgrade checks names the pipelines with duplicate groups.
  • Cause. The migration copies fragment rows into a new table with a unique index on the document, version and position of each fragment. Duplicate rows in a pipeline stop the copy.
  • Resolution. The failed migration left no changes. Send the pipeline identifiers and counts to the Foundation4 provider, and keep the pods of the previous version running until the upgrade completes or the backup is restored.
  • Actions to avoid. Deleting fragment rows by hand, and restarting the API server or the workers of the previous version, which do not start when the upgrade applied earlier migrations.

Pods failing after a rollback with a missing migration​

  • Symptoms. After helm rollback, the server and worker containers restart repeatedly.
  • Diagnosis. kubectl logs -n foundation4ai deploy/foundation4ai-api-server -c server --previous shows Migration file of version '<version>' is missing, this migration has been applied but its file is missing.
  • Cause. The database schema belongs to the newer version. Migrations cannot be reverted, and the previous version refuses to start against migrations that the previous version does not know.
  • Resolution. Return to the newer version with helm rollback foundation4ai <newer revision> -n foundation4ai, or restore the pre-upgrade backup and then roll back. A logical restore needs a new license.
  • Actions to avoid. Deleting rows from schema_migrations. The next upgrade then applies the migrations again to a schema that already contains the changes.

Restored installation waiting at the license check​

  • Symptoms. After a restore, the application upgrade stops at the license check, and the API server does not start.
  • Diagnosis. kubectl logs -n foundation4ai job/foundation4ai-api-server-license-check shows a system ID that differs from the system ID of the license.
  • Cause. The restored database has a new system ID.
  • Resolution. Complete step 5 of Logical restore with a license for the new system ID.
  • Actions to avoid. Restoring the dump again into another new database, which produces another system ID.