Skip to main content

Troubleshooting

This page lists the problems that operators meet while they install, start, reach, load and upgrade Foundation4, in that order. Each entry names the observed symptom, the commands that confirm the cause, the cause, the resolution and the actions that make the problem worse or hide the cause. The commands use the namespace foundation4ai, the releases foundation4ai-core and foundation4ai, and the values file foundation4ai.values.yaml of Install on Kubernetes. Operator command reference describes each command and the output of a healthy deployment.

First checks​

The following commands show the state of every pod, Job and volume claim, the latest events and the state of both releases:

kubectl get pods,jobs,pvc -n foundation4ai
kubectl get events -n foundation4ai --sort-by=.lastTimestamp | tail -20
helm list -n foundation4ai

Expected result on a healthy deployment: every pod is Running with all containers ready, the three installation Jobs show 1/1 completions, every volume claim is Bound, and both releases show deployed.

Installation​

Pods or volume claims in pending status​

  • Symptoms. After the core release installation, NATS JetStream or Prometheus server pods remain Pending, kubectl get pvc -n foundation4ai shows claims in Pending status, and Helm waits until the timeout.

  • Diagnosis.

    kubectl get storageclass
    kubectl describe pvc -n foundation4ai foundation4ai-core-nats-js-foundation4ai-core-nats-0
    kubectl describe pod -n foundation4ai foundation4ai-core-nats-0

    A list without a class marked (default), or claim events that report that no storage class is set, point to storage. Pod events that report Insufficient cpu or Insufficient memory point to node capacity instead.

  • Cause. The NATS JetStream claims (3 of 10 GiB) and the Prometheus claim (8 GiB) name no storage class unless the values name one, so the cluster needs a default StorageClass.

  • Resolution. In the evaluation profile, mark a StorageClass as the default or install a local volume provisioner. A production installation names the class in foundation4ai.values.yaml before the first installation, then runs the core release installation again:

    nats:
    config:
    jetstream:
    fileStore:
    pvc:
    storageClassName: <storage-class>
    prometheus:
    server:
    persistentVolume:
    storageClass: <storage-class>

    Missing node capacity is resolved with more nodes or larger nodes; Scaling and performance describes the resource requests of each component.

  • Actions to avoid. Deleting the NATS JetStream volume claims after documents have been submitted, which deletes the queued documents. Changing the storage class of an existing installation in the values, which the NATS JetStream StatefulSet does not apply to existing claims.

Image pull errors​

  • Symptoms. Hook Job pods or application pods report ErrImagePull, ImagePullBackOff or InvalidImageName, and the installation stops at the timeout. On a running deployment, new or restarted pods fail to pull while older pods keep running.

  • Diagnosis.

    kubectl describe pod -n foundation4ai <pod-name>
    kubectl get pod -n foundation4ai <pod-name> -o jsonpath='{.spec.imagePullSecrets}{"\n"}'

    The events show the image reference and the response of the registry, such as not found or unauthorized. The second command lists the pull secrets of the pod.

  • Cause. A repository or tag value is missing or wrong, the registry requires credentials that no pull secret provides, or the registry token in the pull secret has expired. Some registries issue tokens that expire after a few hours. A pull secret listed only under a model or package image reaches the API server pods, but not the worker pods or the installation Jobs.

  • Resolution. Correct the repository and tag values; the charts provide no default tags. Create the pull secret and list the secret under global.imagePullSecrets of the application release, which the API server, worker, Job and dashboard pods use. The core release takes pull secrets from the per-chart values listed in Charts, images and installation bundle. For short-lived registry tokens, refresh the pull secret on a schedule or pull from a mirror registry, as described on the same page. Then delete the failing pods, or repeat the failed installation step.

  • Actions to avoid. Setting the pull policy to Always with short-lived tokens, which makes every pod restart depend on a valid token. Omitting the tags in the expectation of a default.

Image volume errors on older Kubernetes versions​

  • Symptoms. With api-server.models or api-server.packages set, the application installation fails with a validation error on a volume of the API server or worker Deployment, or the API server and worker pods remain in ContainerCreating.

  • Diagnosis.

    kubectl version
    kubectl get nodes -o wide
    kubectl describe pod -n foundation4ai -l app.kubernetes.io/name=api-server-worker

    kubectl version shows the server version, and the CONTAINER-RUNTIME column shows the runtime of each node. A Helm error that a volume must specify a volume type means that the Kubernetes API server dropped the image volume source. Pod events that report a failure to mount an image volume point to the kubelet or the container runtime.

  • Cause. Model and package images use the Kubernetes image volume source, which needs support in the Kubernetes API server, the kubelet and the container runtime. The charts declare no minimum Kubernetes version.

  • Resolution. Install on a Kubernetes version and a container runtime that support image volumes, as listed in Deployment overview and requirements, or enable the ImageVolume feature gate on a version that provides the feature behind the gate.

  • Actions to avoid. Removing models from the values of a deployment whose pipelines use the models from the images. Those pipelines then fail to embed text.

Credentials missing after installation through a GitOps tool​

  • Symptoms. A GitOps tool fails to render the charts, or the installation Jobs fail with an empty database URL or connection errors.

  • Diagnosis.

    kubectl get secret foundation4ai-core -n foundation4ai \
    -o jsonpath='{.data.FOUNDATION4AI_DATABASE_URL}' | base64 -d | cut -d@ -f2; echo

    An empty line means that the database URL is missing. The tool may instead report a template error in secrets.yaml or secret.yaml.

  • Cause. Both charts read Secrets from the cluster while Helm renders the templates. Tools that render charts without cluster access, such as helm template and Argo CD, receive no Secret data.

  • Resolution. Install and upgrade both releases with the Helm command line against the cluster, as described in Secrets and keys.

  • Actions to avoid. Storing credentials in a values file in a Git repository to replace the cluster lookup, which exposes every credential to the readers of the repository.

Database migration failure​

  • Symptoms. The application installation or upgrade stops in the pre-install or pre-upgrade hook, and the pods of job/foundation4ai-api-server-db-migration end in Error.

  • Diagnosis.

    kubectl logs -n foundation4ai job/foundation4ai-api-server-db-migration
    kubectl get secret foundation4ai-core -n foundation4ai \
    -o jsonpath='{.data.FOUNDATION4AI_DATABASE_URL}' | base64 -d | cut -d@ -f2; echo

    Error: "Failed connecting to database" carries no reason, because the command reports every connection failure with the same text. The second command prints the host, the database and the parameters of the URL without the credentials, or an empty line when the URL is missing. A PostgreSQL client started in the namespace, as described in Operator command reference, shows the reason of a connection failure. Error: "Failed applying migrations: <reason>" means that the connection succeeded and a migration failed; a reason with extension "vector" is not available or permission denied to create extension "vector" points to the pgvector extension.

  • Cause. The URL is empty, because postgres.enabled is false and POSTGRES_URL is missing. A password contains characters that break the URL, because the charts insert passwords without encoding. The network, DNS, a network policy or pg_hba.conf blocks the connection. The TLS settings do not match the server, such as sslmode=verify-full with a server certificate from a private certificate authority that database.ca_cert does not contain. The database lacks pgvector, or the database user may not create the extension.

  • Resolution. Correct the value in foundation4ai.secrets.env (hexadecimal passwords contain no URL characters), the TLS settings described in Security hardening, the network path, or the database prerequisites described in Database. Apply the Secret, upgrade the core release and repeat the application installation or upgrade.

  • Actions to avoid. Deleting the bundled PostgreSQL pod to force a new connection: in the evaluation profile, a new pod starts with an empty database and a new system ID. Granting the Foundation4 database user superuser rights to create the extension, where an administrator can create the extension once.

Install waiting on the license check​

  • Symptoms. helm upgrade --install of the application release waits until the timeout and reports that the pre-install or pre-upgrade hook failed. The pod of job/foundation4ai-api-server-license-check stays Running after Helm stops.

  • Diagnosis.

    kubectl logs -n foundation4ai job/foundation4ai-api-server-license-check

    The log shows the system ID and the result of the check:

    SystemID: <system ID>
    Checking for a valid license... Invalid
    ResultMeaning
    OKThe license is valid; the hook completes
    MissingFOUNDATION4AI_APP_LICENSE is empty
    InvalidThe license is malformed or truncated, or was issued for another system ID
    ExpiredThe license expiry date has passed
  • Cause. After any result other than OK, the check waits 365 days so that the system ID stays in the log. The hook therefore never completes, and Helm stops at the timeout. The Job has no deletion policy, so the waiting pod continues after Helm stops. A new database, including a logical restore, has a new system ID, so a license issued before the change reads as Invalid.

  • Resolution. Send the system ID to the Foundation4 provider, as described in Licensing. Replace the value of FOUNDATION4AI_APP_LICENSE in foundation4ai.secrets.env, apply the Secret, upgrade the core release and delete the waiting Job:

    kubectl kustomize . | kubectl apply -f -
    helm upgrade foundation4ai-core ./charts/foundation4ai-core \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m
    kubectl delete job -n foundation4ai foundation4ai-api-server-license-check

    Expected result: secret/foundation4ai-secrets configured, STATUS: deployed for foundation4ai-core, and job.batch "foundation4ai-api-server-license-check" deleted. Then repeat the application installation or upgrade. A first installation that failed is removed with helm uninstall foundation4ai -n foundation4ai before the application release is installed again. The license check log then ends with Checking for a valid license... OK.

  • Actions to avoid. Running license check through kubectl exec, which waits 365 days after a failure; license details prints the result and returns. Raising the Helm timeout, which only delays the failure. Recreating the database, which changes the system ID again.

Master key Job failure​

  • Symptoms. The installation or upgrade stops after the license check, the pods of job/foundation4ai-api-server-create-admin-api-key end in Error, and Helm reports that the hook failed.

  • Diagnosis.

    kubectl logs -n foundation4ai job/foundation4ai-api-server-create-admin-api-key

    The Job reports every connection failure as Error: "Failed to create database connection: <reason>", and the reason names the component:

    Reason begins withComponent
    ORM error:PostgreSQL
    Invalid license, Expired licenseLicense
    Redis error:Redis-compatible cache: address or password
    Nats error:NATS JetStream
    Decryption error: Error decoding secret keyApplication secret
    gRPC error:Address of the gRPC service in grpc.url

    Error: "Failed to parse config: <reason>" means that a value of foundation4ai.secrets.env cannot be parsed. A panic message that contains Result::unwrap(), with no Error: line, means that the cache URL is malformed, for example because a password contains characters that break the URL.

  • Cause. The master key Job connects to PostgreSQL, checks the license, and connects to the cache and NATS JetStream before the Job creates the master key, so every core service and a valid license must be available.

  • Resolution. Correct the component named in the log, then delete the failed Job and repeat the application installation or upgrade. Secrets and keys describes the values and the change procedure. The log of a successful run ends with Admin API key created successfully.

  • Actions to avoid. Generating a new master key or master secret for an existing installation. Creating a replacement administration key with admin generate-admin-api-key, because the API server and the workers still authenticate with the configured master key at startup.

Prometheus server restarting after installation​

  • Symptoms. After the core release installation, the pod of the Deployment foundation4ai-core-prometheus-server restarts repeatedly.

  • Diagnosis.

    kubectl logs -n foundation4ai deploy/foundation4ai-core-prometheus-server -c prometheus-server --previous
    kubectl get configmap foundation4ai-core-prometheus-server -n foundation4ai -o yaml \
    | grep -c "job_name: kubernetes-pods"

    The log reports an error loading the configuration file with more than one scrape job named kubernetes-pods, and the count is 2.

  • Cause. The core chart defines a kubernetes-pods scrape job and removes the default job of the same name from the Prometheus chart. When the removal does not take effect, the configuration holds two jobs with one name, and Prometheus refuses the configuration.

  • Resolution. Disable the default job in foundation4ai.values.yaml and upgrade the core release:

    prometheus:
    scrapeConfigs:
    kubernetes-pods:
    enabled: false
    helm upgrade foundation4ai-core ./charts/foundation4ai-core \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: STATUS: deployed, the Prometheus server pod is Running, and the count of the diagnosis is 1.

  • Actions to avoid. Removing the scrape job of the core chart, which is the job that collects the Foundation4 metrics.

Startup​

API server or worker containers restarting​

  • Symptoms. After an installation, an upgrade or a restart, the server or worker container restarts repeatedly and the pod shows CrashLoopBackOff.

  • Diagnosis.

    kubectl logs -n foundation4ai deploy/foundation4ai-api-server -c server --previous
    kubectl logs -n foundation4ai deploy/foundation4ai-api-server-worker -c worker --previous

    A license failure produces a warning with the system ID, followed by the error that stops the process (timestamps omitted):

    WARN foundation4ai_lib::connection: Failed to validate license for system id: <system ID>
    Error: "Failed to initialize server: Failed to create database connection: Invalid license"

    The worker reports the same failures after Failed to connect to Foundation4.ai: instead of Failed to initialize server: Failed to create database connection:. The text after that prefix names the component:

    Log textCause
    Invalid license or Expired licenseThe license does not match the system ID in the warning, or has expired
    ORM error:PostgreSQL connection or migration
    Redis error:Redis-compatible cache
    Nats error:NATS JetStream
    Decryption error: Error decoding secret keyApplication secret format
    Failed to initialize Foundation4aiCore with master key (API server), Failed to initialize Foundation4.ai: (worker)Master key
    Failed to parse config:Configuration file or value
    Failed to bind to addressListen address or port of the API server
    Migration file of version '<version>' is missingThe image is older than the database schema, for example after helm rollback (inferred)
  • Cause. At startup, the API server and each worker connect to PostgreSQL, apply pending migrations, check the license, connect to the cache and NATS JetStream, then authenticate with the master key. A failure of any step stops the process, and Kubernetes restarts the container. The API server reports every failure of the connection steps, including a license failure, as Failed to create database connection.

  • Resolution. Correct the component named in the log: request a license for the system ID in the warning, as described in Licensing; restore PostgreSQL, the cache or NATS JetStream; restore the original application secret or master key values; or correct the configuration. A value in foundation4ai.secrets.env changes through the change procedure in Secrets and keys. An image older than the database is replaced by the current version, as described in Backup, restore and upgrades.

  • Actions to avoid. Generating a new master secret for an existing installation. Deleting the bundled PostgreSQL pod, which in the evaluation profile creates an empty database with a new system ID.

Access​

Login returning HTTP 401​

  • Symptoms. POST /login returns HTTP 401, and the dashboard sign-in fails.

  • Diagnosis.

    curl -sS -X POST "$FOUNDATION4_URL/login" \
    -H "x-api-key: $FOUNDATION4_API_KEY" \
    -H "x-api-key-secret: $FOUNDATION4_API_SECRET"
    kubectl exec -n foundation4ai deploy/foundation4ai-api-server -c server -- \
    ./foundation4ai database check-connection

    The message of the error body names the cause, as listed in Errors: a missing header, a key identifier that is not a UUID, or Invalid API Key or Secret. The connection check prints Ok. when the configuration of the API server reaches PostgreSQL, and Error: "Failed connecting to database" otherwise.

  • Cause. Invalid API Key or Secret covers an unknown key, a wrong secret, an inactive or expired key, and a database failure while the API server checks the key. A license problem does not produce HTTP 401: an invalid license stops the API server at startup, and an expired license returns HTTP 403 Invalid license: Expired on the operations that check the license.

  • Resolution. Send the key identifier in x-api-key and the secret in x-api-key-secret, as described in Authenticate. The master key identifier and secret are the values of FOUNDATION4AI_APP_MASTER_KEY and FOUNDATION4AI_APP_MASTER_SECRET in foundation4ai.secrets.env. When a correct key fails for every client, restore the database connection of the API server.

  • Actions to avoid. Replacing the master key values in foundation4ai.secrets.env to regain access, which stops the API server and the workers at the next restart.

Ingress configured but unreachable​

  • Symptoms. Requests to the host name time out, or the ingress controller returns HTTP 404 or 503. Through a Gateway, GET / returns a redirect to /dashboard.

  • Diagnosis.

    kubectl get ingressclass
    kubectl describe ingress foundation4ai -n foundation4ai
    kubectl get endpoints foundation4ai-api-server foundation4ai-dashboard -n foundation4ai
    kubectl get httproute foundation4ai -n foundation4ai -o yaml

    kubectl describe ingress shows the class, the host and the backends of the Ingress. An endpoints entry without addresses means that no ready pod serves the Service. The status of the HTTPRoute lists each parent with an Accepted condition; False with a reason such as NoMatchingParent means that the parent reference matches no Gateway listener.

  • Cause. ingress.className names a class that no controller serves. The request host differs from ingress.hosts, whose chart default is chart-example.local. No API server pod is ready. The default httpRoute.parentRefs (Gateway gateway, listener http) match no Gateway. A network policy blocks the controller. Through the chart's HTTPRoute, an exact request for / redirects to /dashboard, so GET / never reaches the API.

  • Resolution. Set className to an IngressClass from the list, set the host names in ingress.hosts or httpRoute.hostnames and send requests to those names, and set httpRoute.parentRefs to the Gateway and the listener. Allow the controller namespace in the network policies described in Security hardening. Check the API with POST /login rather than GET /.

  • Actions to avoid. Changing the API server Service to LoadBalancer or NodePort to bypass the controller, which also bypasses the TLS of the ingress.

Dashboard sign-in failure through a port forward​

  • Symptoms. A port forward to the dashboard Service shows the dashboard page, but the sign-in fails with a correct key.
  • Diagnosis. The network panel of the browser's developer tools shows the sign-in request POST /login sent to the forwarded local address, where the dashboard server answers instead of the API server.
  • Cause. The dashboard is served under /dashboard/ and calls the API on the origin of the dashboard page, so the dashboard works only where one host name serves both /dashboard and the API. The chart's Ingress and HTTPRoute provide that host name; a port forward to one Service does not.
  • Resolution. Open the dashboard through the Ingress or the HTTPRoute host name, as described in API and dashboard access.
  • Actions to avoid. Exposing the dashboard or the API server through a LoadBalancer or NodePort Service to replace the port forward, which bypasses the TLS of the ingress.

Ingestion​

Documents remaining in pending status​

  • Symptoms. The pending count does not fall, or documents remain pending long after submission.

  • Diagnosis.

    kubectl logs -n foundation4ai -l app.kubernetes.io/name=api-server-worker -c worker --tail=500 --prefix \
    | grep -E "Error processing documents|Failed processing document"
    kubectl get pods -n foundation4ai -l app.kubernetes.io/name=api-server-worker
    kubectl logs -n foundation4ai deploy/foundation4ai-api-server-worker -c grpc --tail=100
    kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- \
    nats consumer info DOCUMENTS document-processor

    A line with Failed processing document names one document, the pipeline and the reason. A line with Error processing documents for pipeline names a failure of a whole pipeline group, such as an error of the gRPC service, of the embedding model or Expired license. The same group error every 10 seconds for one pipeline means that the group returns to the queue on each attempt. The gRPC service log shows the errors of Python text splitters and embedding providers, and the consumer information shows the jobs that wait and the jobs delivered more than once.

  • Cause. Processing fails for the document, for the text splitter or embedding model of the pipeline, or for the license, or no worker runs. A document that the text splitter or embedding model rejects fails every document of the same pipeline in the batch, and a group failure returns every job of the batch to the queue. The worker retries every 10 seconds, the document stays pending while processing fails, and the queue removes a job 24 hours after submission, after which the document is not processed again.

  • Resolution. Remove the cause shown in the log: restore the gRPC service or the model server, install a valid license as described in the next entry, or start the workers. Documents that are pending for less than 24 hours are then processed without action. Documents that are pending more than 24 hours after creation are submitted again, as described in Resubmission.

  • Actions to avoid. Submitting documents again while the cause persists, which adds versions and jobs that fail in the same way. Purging or recreating the DOCUMENTS stream, which removes the jobs of every queued document and leaves those documents pending.

Processing stopped after license expiry​

  • Symptoms. The worker log shows Error processing documents for pipeline <id>: Expired license every 10 seconds, operations that check the license return HTTP 403 Invalid license: Expired, and restarted API server or worker containers stop at startup with Expired license.

  • Diagnosis.

    kubectl logs -n foundation4ai -l app.kubernetes.io/name=api-server-worker -c worker --tail=200 --prefix \
    | grep -c "Expired license"
    LICENSE=$(grep '^FOUNDATION4AI_APP_LICENSE=' foundation4ai.secrets.env | cut -d= -f2-)
    kubectl exec -n foundation4ai deploy/foundation4ai-api-server -c server -- \
    ./foundation4ai license details "$LICENSE"

    A count above zero and the output Expired confirm the cause.

  • Cause. The worker checks the license before each pipeline group, so an expired license fails every group. The processes read the license at startup, so a renewed license takes effect only after a restart, and a process that restarts with an expired license stops at startup. The queue removes jobs 24 hours after submission, so documents submitted more than 24 hours before the renewal remain pending.

  • Resolution. Obtain a renewed license, as described in Licensing, and apply the license with the change procedure in Secrets and keys, which ends with a restart of the API server and the workers. The license check log of the application upgrade shows Checking for a valid license... OK. Then submit again the documents that are still pending more than 24 hours after creation, as described in Resubmission.

  • Actions to avoid. Upgrading a release for other changes before the renewal, which stops at the license check. Purging the DOCUMENTS stream.

Workers timing out on NATS JetStream​

  • Symptoms. Worker containers restart at startup with Error: "Failed to connect to Foundation4.ai: Nats error: <reason>", where the reason reports a JetStream timeout or that JetStream is unavailable. The API server may report the same reason after Failed to create database connection:, and document submission returns HTTP 500 NATS error.

  • Diagnosis.

    kubectl get pods -n foundation4ai -l app.kubernetes.io/component=nats
    kubectl logs -n foundation4ai foundation4ai-core-nats-0 -c nats --tail=50
    kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats account info
    kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats stream info DOCUMENTS

    The 3 NATS pods are expected Running with every container ready. nats account info shows whether JetStream is available to the account and how much storage is used, and nats stream info shows the server that holds the stream. A NATS log that reports no JetStream leader, or storage limits reached, points to the NATS cluster.

  • Cause. At startup, each worker and API server opens or creates the documents object store and the DOCUMENTS stream. JetStream answers only while a majority of the 3 NATS servers runs, and refuses new data when the JetStream volume is full, at 10 GiB per server by default. A network policy that blocks port 4222 or 6222 has the same effect.

  • Resolution. Restore the NATS pods, free or enlarge the JetStream storage described in Deployment overview and requirements, and allow ports 4222 and 6222 in the network policies. Then restart the workers:

    kubectl rollout restart deployment/foundation4ai-api-server-worker -n foundation4ai
    kubectl rollout status deployment/foundation4ai-api-server-worker -n foundation4ai

    Expected result: successfully rolled out, and the worker log shows no further NATS error.

  • Actions to avoid. Deleting the NATS JetStream volume claims or the DOCUMENTS stream, which removes the queued jobs. Reducing NATS to one replica on an existing installation, which can remove the server that holds the stream.

Worker restarts on queue errors​

  • Symptoms. Worker containers restart during processing, not at startup, and the pod shows no OOMKilled reason.

  • Diagnosis.

    kubectl get pods -n foundation4ai -l app.kubernetes.io/name=api-server-worker
    kubectl logs -n foundation4ai <worker-pod-name> -c worker --previous | tail -5

    The last line of the previous container is one of the following: Error: "Error retrieving messages from message queue.", Error: "Error processing message: Decryption error: Invalid message queue document" or Error: "Message not found for document <id> version <version>".

  • Cause. The worker ends the process when a read from the queue fails, when a job message cannot be decoded, or when a processing result names a document that the batch does not contain, and Kubernetes restarts the container. A transient NATS JetStream failure causes one restart per worker. A message that no worker can decode returns to the queue after the acknowledgement timeout and can restart each worker in turn until the queue removes the message 24 hours after submission (inferred).

  • Resolution. After Error retrieving messages from message queue., check NATS JetStream as in the previous entry; the workers resume after the restart. When a decoding error repeats, collect the worker logs and the output of nats stream info DOCUMENTS and nats consumer info DOCUMENTS document-processor, and contact the Foundation4 provider.

  • Actions to avoid. Purging the DOCUMENTS stream, which removes every queued job. Scaling the workers to 0 for more than 24 hours during the investigation, after which the queue has removed every waiting job.

Load​

Workers restarted for exceeding memory​

  • Symptoms. Worker pods restart during ingestion, and kubectl describe pod shows Last State: Terminated with Reason: OOMKilled and exit code 137, or the pod status is Evicted.

  • Diagnosis.

    kubectl get pods -n foundation4ai -l app.kubernetes.io/name=api-server-worker \
    -o jsonpath='{range .items[*]}{.metadata.name}{": "}{range .status.containerStatuses[*]}{.name}={.lastState.terminated.reason}{" "}{end}{"\n"}{end}'
    kubectl top pods -n foundation4ai --containers

    worker=OOMKilled or grpc=OOMKilled names the container that exceeded the memory limit. kubectl top, which needs the metrics server, shows the memory of each container during a batch.

  • Cause. Each worker takes up to 100 jobs per batch, and the memory of a batch grows with the size of the documents and with the embedding model. FastEmbed models run in the worker process, and Python embedding providers and text splitters run in the gRPC service. The charts set no resources, and one api-server.resources value applies to the server, worker and grpc containers, each with a separate limit. A worker that stops during a batch leaves the jobs unacknowledged, the queue delivers the jobs again after the acknowledgement timeout, and the same batch can exceed the limit again.

  • Resolution. Set requests and a memory limit that covers the largest container, then upgrade the application release. The values below are an example, not a sizing recommendation; Scaling and performance describes how to size the containers:

    api-server:
    resources:
    requests:
    cpu: "1"
    memory: 2Gi
    limits:
    memory: 4Gi
  • Actions to avoid. Removing the memory limit on shared nodes, which lets one batch take memory from other pods. Adding worker replicas to reduce memory use, which does not change the size of a batch.

Low ingestion throughput​

  • Symptoms. The pending count grows during steady submission, and the message count of the DOCUMENTS stream rises.

  • Diagnosis. Read the stream state, then forward the metrics port of a worker in a separate terminal and read the worker metrics:

    kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats stream info DOCUMENTS
    kubectl port-forward -n foundation4ai deploy/foundation4ai-api-server-worker 9090:9090
    curl -s http://localhost:9090/metrics | grep foundation4ai_document_queue

    The counters foundation4ai_document_queue_processed, _success, _skip and _failed count documents. The values _processed_latency, _text_splitter_latency and _embedding_latency are running totals in seconds. Two readings a few minutes apart give the documents per second of that worker and the share of the time spent in the text splitter and the embedding model. A rising _failed counter means that the backlog comes from failures, as described in the entry on pending documents. The port forward reaches one worker pod.

  • Cause. Each worker processes one batch at a time, and the pipeline groups of a batch one after another. Throughput is limited by the number of workers, by the CPU available to the worker and gRPC service containers, and by the time the embedding model needs per fragment.

  • Resolution. Add worker replicas with api-server.replicaCount.worker, set CPU requests for the containers, or use an embedding model that runs on a dedicated model server, as described in Scaling and performance.

  • Actions to avoid. Adding API server replicas, which do not process documents. Submitting more documents than the workers process in 24 hours, because the queue removes jobs after 24 hours.

HTTP 500 responses under load​

  • Symptoms. Under load, requests return HTTP 500 after about 10 seconds with the message Database connection error: Failed to acquire connection from pool: Connection pool timed out, or with Failed to acquire connection from pool: Connection pool timed out.

  • Diagnosis.

    psql "<administrator-connection-url>" -c "SHOW max_connections;"
    psql "<administrator-connection-url>" -c \
    "SELECT usename, state, count(*) FROM pg_stat_activity WHERE datname = '<database>' GROUP BY 1, 2;"
    kubectl get deployments,hpa -n foundation4ai

    The connection counts per state show whether the API server pools are full, and the replica counts give the number of pools.

  • Cause. Each API server process holds at most database.pool_size connections, 3 in the chart defaults, and a request waits up to 10 seconds for a free connection. Bursts or long searches exhaust the pool. Each additional API server replica adds a pool, so replicas, including those that an autoscaler adds, can reach max_connections of PostgreSQL, which then refuses new connections. Each worker holds 1 connection.

  • Resolution. Raise the pool size in a configs file, keep the total below max_connections as described in Database, then upgrade the application release and restart the API server:

    api-server:
    configs:
    f4ai-10-database-pool.yaml: |
    database:
    pool_size: 10
    helm upgrade foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m
    kubectl rollout restart deployment/foundation4ai-api-server -n foundation4ai

    Expected result: Helm reports STATUS: deployed, then deployment.apps/foundation4ai-api-server restarted.

  • Actions to avoid. Raising the API server replica count or the autoscaler maximum without a connection budget, which moves the failure to PostgreSQL.

Upgrades and changes​

Upgrade failure creating the fragment index​

  • Symptoms. An upgrade of the application release fails in the pre-upgrade hook, and the migration Job log shows Failed applying migrations: with duplicate key value violates unique constraint.

  • Diagnosis. The first statement shows whether the migration m20260827_140000_vector_embeddings is still pending, which an empty result confirms. The second, read-only statement reports each pipeline table with duplicate fragment keys; the fragment tables still carry the pipeline identifier as name while the migration is pending. A deployment with a custom database.schema replaces public:

    psql "<administrator-connection-url>" -c \
    "SELECT version FROM public.schema_migrations WHERE version = 'm20260827_140000_vector_embeddings';"
    psql "<administrator-connection-url>" <<'SQL'
    DO $$
    DECLARE p record; n bigint;
    BEGIN
    FOR p IN SELECT id FROM public.pipelines LOOP
    EXECUTE format('SELECT count(*) FROM (SELECT 1 FROM embeddings.%I GROUP BY document_id, version, order_id HAVING count(*) > 1) d', p.id::text) INTO n;
    IF n > 0 THEN RAISE NOTICE 'pipeline %: % duplicate fragment keys', p.id, n; END IF;
    END LOOP;
    END $$;
    SQL

    Expected result on an unaffected database: DO and no notice.

  • Cause. One migration of the upgrade creates a unique index on document, version and fragment order before the migration copies the existing fragments into the new table, so duplicate rows stop the copy. The migration runs in a transaction and is rolled back, and the failed hook leaves the previous version running.

  • Resolution. Keep the previous version running, collect the migration log and the output of the statement, and obtain a de-duplication procedure from the Foundation4 provider before the upgrade is repeated. Backup, restore and upgrades describes the backup before an upgrade.

  • Actions to avoid. Deleting fragment rows or dropping the index by hand. Running helm rollback after a later migration has succeeded, because an older image does not start against a newer database schema (inferred).

Pods using a previous value after a change​

  • Symptoms. After a renewed license, a new password, a new URL or a change of api-server.configs, the API server and the workers behave as before.

  • Diagnosis.

    kubectl get pods -n foundation4ai -l app.kubernetes.io/instance=foundation4ai
    helm history foundation4ai -n foundation4ai

    The AGE of the API server and worker pods is older than the last upgrade in the release history.

  • Cause. Foundation4 reads configuration once at startup. The charts write configuration and credentials to Secrets and a ConfigMap, and the Deployments carry no reference to the content of those objects, so an upgrade that changes only the Secrets or the ConfigMap does not restart the pods.

  • Resolution. Run the change procedure in Secrets and keys, which ends with the restart:

    kubectl rollout restart -n foundation4ai \
    deployment/foundation4ai-api-server deployment/foundation4ai-api-server-worker

    Expected result: both Deployments report restarted, and the new pods show the changed behavior.

  • Actions to avoid. Editing or deleting the Secret foundation4ai-api-server directly. The next upgrade replaces the Secret, and pods fail to start while the Secret is missing.