Troubleshooting
This page lists the problems that operators meet while they install, start, reach, load and upgrade Foundation4, in that order. Each entry names the observed symptom, the commands that confirm the cause, the cause, the resolution and the actions that make the problem worse or hide the cause. The commands use the namespace foundation4ai, the releases foundation4ai-core and foundation4ai, and the values file foundation4ai.values.yaml of Install on Kubernetes. Operator command reference describes each command and the output of a healthy deployment.
First checks
The following commands show the state of every pod, Job and volume claim, the latest events and the state of both releases:
kubectl get pods,jobs,pvc -n foundation4ai
kubectl get events -n foundation4ai --sort-by=.lastTimestamp | tail -20
helm list -n foundation4ai
Expected result on a healthy deployment: every pod is Running with all containers ready, the three installation Jobs show 1/1 completions, every volume claim is Bound, and both releases show deployed.
Installation
Pods or volume claims in pending status
-
Symptoms. After the core release installation, NATS JetStream or Prometheus server pods remain
Pending,kubectl get pvc -n foundation4aishows claims inPendingstatus, and Helm waits until the timeout. -
Diagnosis.
kubectl get storageclasskubectl describe pvc -n foundation4ai foundation4ai-core-nats-js-foundation4ai-core-nats-0kubectl describe pod -n foundation4ai foundation4ai-core-nats-0A list without a class marked
(default), or claim events that report that no storage class is set, point to storage. Pod events that reportInsufficient cpuorInsufficient memorypoint to node capacity instead. -
Cause. The NATS JetStream claims (3 of 10 GiB) and the Prometheus claim (8 GiB) name no storage class unless the values name one, so the cluster needs a default StorageClass.
-
Resolution. In the evaluation profile, mark a StorageClass as the default or install a local volume provisioner. A production installation names the class in
foundation4ai.values.yamlbefore the first installation, then runs the core release installation again:nats:config:jetstream:fileStore:pvc:storageClassName: <storage-class>prometheus:server:persistentVolume:storageClass: <storage-class>Missing node capacity is resolved with more nodes or larger nodes; Scaling and performance describes the resource requests of each component.
-
Actions to avoid. Deleting the NATS JetStream volume claims after documents have been submitted, which deletes the queued documents. Changing the storage class of an existing installation in the values, which the NATS JetStream StatefulSet does not apply to existing claims.
Image pull errors
-
Symptoms. Hook Job pods or application pods report
ErrImagePull,ImagePullBackOfforInvalidImageName, and the installation stops at the timeout. On a running deployment, new or restarted pods fail to pull while older pods keep running. -
Diagnosis.
kubectl describe pod -n foundation4ai <pod-name>kubectl get pod -n foundation4ai <pod-name> -o jsonpath='{.spec.imagePullSecrets}{"\n"}'The events show the image reference and the response of the registry, such as
not foundorunauthorized. The second command lists the pull secrets of the pod. -
Cause. A repository or tag value is missing or wrong, the registry requires credentials that no pull secret provides, or the registry token in the pull secret has expired. Some registries issue tokens that expire after a few hours. A pull secret listed only under a model or package image reaches the API server pods, but not the worker pods or the installation Jobs.
-
Resolution. Correct the repository and tag values; the charts provide no default tags. Create the pull secret and list the secret under
global.imagePullSecretsof the application release, which the API server, worker, Job and dashboard pods use. The core release takes pull secrets from the per-chart values listed in Charts, images and installation bundle. For short-lived registry tokens, refresh the pull secret on a schedule or pull from a mirror registry, as described on the same page. Then delete the failing pods, or repeat the failed installation step. -
Actions to avoid. Setting the pull policy to
Alwayswith short-lived tokens, which makes every pod restart depend on a valid token. Omitting the tags in the expectation of a default.
Image volume errors on older Kubernetes versions
-
Symptoms. With
api-server.modelsorapi-server.packagesset, the application installation fails with a validation error on a volume of the API server or worker Deployment, or the API server and worker pods remain inContainerCreating. -
Diagnosis.
kubectl versionkubectl get nodes -o widekubectl describe pod -n foundation4ai -l app.kubernetes.io/name=api-server-workerkubectl versionshows the server version, and theCONTAINER-RUNTIMEcolumn shows the runtime of each node. A Helm error that a volume must specify a volume type means that the Kubernetes API server dropped the image volume source. Pod events that report a failure to mount an image volume point to the kubelet or the container runtime. -
Cause. Model and package images use the Kubernetes image volume source, which needs support in the Kubernetes API server, the kubelet and the container runtime. The charts declare no minimum Kubernetes version.
-
Resolution. Install on a Kubernetes version and a container runtime that support image volumes, as listed in Deployment overview and requirements, or enable the
ImageVolumefeature gate on a version that provides the feature behind the gate. -
Actions to avoid. Removing
modelsfrom the values of a deployment whose pipelines use the models from the images. Those pipelines then fail to embed text.
Credentials missing after installation through a GitOps tool
-
Symptoms. A GitOps tool fails to render the charts, or the installation Jobs fail with an empty database URL or connection errors.
-
Diagnosis.
kubectl get secret foundation4ai-core -n foundation4ai \-o jsonpath='{.data.FOUNDATION4AI_DATABASE_URL}' | base64 -d | cut -d@ -f2; echoAn empty line means that the database URL is missing. The tool may instead report a template error in
secrets.yamlorsecret.yaml. -
Cause. Both charts read Secrets from the cluster while Helm renders the templates. Tools that render charts without cluster access, such as
helm templateand Argo CD, receive no Secret data. -
Resolution. Install and upgrade both releases with the Helm command line against the cluster, as described in Secrets and keys.
-
Actions to avoid. Storing credentials in a values file in a Git repository to replace the cluster lookup, which exposes every credential to the readers of the repository.
Database migration failure
-
Symptoms. The application installation or upgrade stops in the pre-install or pre-upgrade hook, and the pods of
job/foundation4ai-api-server-db-migrationend inError. -
Diagnosis.
kubectl logs -n foundation4ai job/foundation4ai-api-server-db-migrationkubectl get secret foundation4ai-core -n foundation4ai \-o jsonpath='{.data.FOUNDATION4AI_DATABASE_URL}' | base64 -d | cut -d@ -f2; echoError: "Failed connecting to database"carries no reason, because the command reports every connection failure with the same text. The second command prints the host, the database and the parameters of the URL without the credentials, or an empty line when the URL is missing. A PostgreSQL client started in the namespace, as described in Operator command reference, shows the reason of a connection failure.Error: "Failed applying migrations: <reason>"means that the connection succeeded and a migration failed; a reason withextension "vector" is not availableorpermission denied to create extension "vector"points to the pgvector extension. -
Cause. The URL is empty, because
postgres.enabledisfalseandPOSTGRES_URLis missing. A password contains characters that break the URL, because the charts insert passwords without encoding. The network, DNS, a network policy orpg_hba.confblocks the connection. The TLS settings do not match the server, such assslmode=verify-fullwith a server certificate from a private certificate authority thatdatabase.ca_certdoes not contain. The database lacks pgvector, or the database user may not create the extension. -
Resolution. Correct the value in
foundation4ai.secrets.env(hexadecimal passwords contain no URL characters), the TLS settings described in Security hardening, the network path, or the database prerequisites described in Database. Apply the Secret, upgrade the core release and repeat the application installation or upgrade. -
Actions to avoid. Deleting the bundled PostgreSQL pod to force a new connection: in the evaluation profile, a new pod starts with an empty database and a new system ID. Granting the Foundation4 database user superuser rights to create the extension, where an administrator can create the extension once.
Install waiting on the license check
-
Symptoms.
helm upgrade --installof the application release waits until the timeout and reports that the pre-install or pre-upgrade hook failed. The pod ofjob/foundation4ai-api-server-license-checkstaysRunningafter Helm stops. -
Diagnosis.
kubectl logs -n foundation4ai job/foundation4ai-api-server-license-checkThe log shows the system ID and the result of the check:
SystemID: <system ID>Checking for a valid license... InvalidResult Meaning OKThe license is valid; the hook completes MissingFOUNDATION4AI_APP_LICENSEis emptyInvalidThe license is malformed or truncated, or was issued for another system ID ExpiredThe license expiry date has passed -
Cause. After any result other than
OK, the check waits 365 days so that the system ID stays in the log. The hook therefore never completes, and Helm stops at the timeout. The Job has no deletion policy, so the waiting pod continues after Helm stops. A new database, including a logical restore, has a new system ID, so a license issued before the change reads asInvalid. -
Resolution. Send the system ID to the Foundation4 provider, as described in Licensing. Replace the value of
FOUNDATION4AI_APP_LICENSEinfoundation4ai.secrets.env, apply the Secret, upgrade the core release and delete the waiting Job:kubectl kustomize . | kubectl apply -f -helm upgrade foundation4ai-core ./charts/foundation4ai-core \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mkubectl delete job -n foundation4ai foundation4ai-api-server-license-checkExpected result:
secret/foundation4ai-secrets configured,STATUS: deployedforfoundation4ai-core, andjob.batch "foundation4ai-api-server-license-check" deleted. Then repeat the application installation or upgrade. A first installation that failed is removed withhelm uninstall foundation4ai -n foundation4aibefore the application release is installed again. The license check log then ends withChecking for a valid license... OK. -
Actions to avoid. Running
license checkthroughkubectl exec, which waits 365 days after a failure;license detailsprints the result and returns. Raising the Helm timeout, which only delays the failure. Recreating the database, which changes the system ID again.
Master key Job failure
-
Symptoms. The installation or upgrade stops after the license check, the pods of
job/foundation4ai-api-server-create-admin-api-keyend inError, and Helm reports that the hook failed. -
Diagnosis.
kubectl logs -n foundation4ai job/foundation4ai-api-server-create-admin-api-keyThe Job reports every connection failure as
Error: "Failed to create database connection: <reason>", and the reason names the component:Reason begins with Component ORM error:PostgreSQL Invalid license,Expired licenseLicense Redis error:Redis-compatible cache: address or password Nats error:NATS JetStream Decryption error: Error decoding secret keyApplication secret gRPC error:Address of the gRPC service in grpc.urlError: "Failed to parse config: <reason>"means that a value offoundation4ai.secrets.envcannot be parsed. A panic message that containsResult::unwrap(), with noError:line, means that the cache URL is malformed, for example because a password contains characters that break the URL. -
Cause. The master key Job connects to PostgreSQL, checks the license, and connects to the cache and NATS JetStream before the Job creates the master key, so every core service and a valid license must be available.
-
Resolution. Correct the component named in the log, then delete the failed Job and repeat the application installation or upgrade. Secrets and keys describes the values and the change procedure. The log of a successful run ends with
Admin API key created successfully. -
Actions to avoid. Generating a new master key or master secret for an existing installation. Creating a replacement administration key with
admin generate-admin-api-key, because the API server and the workers still authenticate with the configured master key at startup.
Prometheus server restarting after installation
-
Symptoms. After the core release installation, the pod of the Deployment
foundation4ai-core-prometheus-serverrestarts repeatedly. -
Diagnosis.
kubectl logs -n foundation4ai deploy/foundation4ai-core-prometheus-server -c prometheus-server --previouskubectl get configmap foundation4ai-core-prometheus-server -n foundation4ai -o yaml \| grep -c "job_name: kubernetes-pods"The log reports an error loading the configuration file with more than one scrape job named
kubernetes-pods, and the count is2. -
Cause. The core chart defines a
kubernetes-podsscrape job and removes the default job of the same name from the Prometheus chart. When the removal does not take effect, the configuration holds two jobs with one name, and Prometheus refuses the configuration. -
Resolution. Disable the default job in
foundation4ai.values.yamland upgrade the core release:prometheus:scrapeConfigs:kubernetes-pods:enabled: falsehelm upgrade foundation4ai-core ./charts/foundation4ai-core \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result:
STATUS: deployed, the Prometheus server pod isRunning, and the count of the diagnosis is1. -
Actions to avoid. Removing the scrape job of the core chart, which is the job that collects the Foundation4 metrics.
Startup
API server or worker containers restarting
-
Symptoms. After an installation, an upgrade or a restart, the
serverorworkercontainer restarts repeatedly and the pod showsCrashLoopBackOff. -
Diagnosis.
kubectl logs -n foundation4ai deploy/foundation4ai-api-server -c server --previouskubectl logs -n foundation4ai deploy/foundation4ai-api-server-worker -c worker --previousA license failure produces a warning with the system ID, followed by the error that stops the process (timestamps omitted):
WARN foundation4ai_lib::connection: Failed to validate license for system id: <system ID>Error: "Failed to initialize server: Failed to create database connection: Invalid license"The worker reports the same failures after
Failed to connect to Foundation4.ai:instead ofFailed to initialize server: Failed to create database connection:. The text after that prefix names the component:Log text Cause Invalid licenseorExpired licenseThe license does not match the system ID in the warning, or has expired ORM error:PostgreSQL connection or migration Redis error:Redis-compatible cache Nats error:NATS JetStream Decryption error: Error decoding secret keyApplication secret format Failed to initialize Foundation4aiCore with master key(API server),Failed to initialize Foundation4.ai:(worker)Master key Failed to parse config:Configuration file or value Failed to bind to addressListen address or port of the API server Migration file of version '<version>' is missingThe image is older than the database schema, for example after helm rollback(inferred) -
Cause. At startup, the API server and each worker connect to PostgreSQL, apply pending migrations, check the license, connect to the cache and NATS JetStream, then authenticate with the master key. A failure of any step stops the process, and Kubernetes restarts the container. The API server reports every failure of the connection steps, including a license failure, as
Failed to create database connection. -
Resolution. Correct the component named in the log: request a license for the system ID in the warning, as described in Licensing; restore PostgreSQL, the cache or NATS JetStream; restore the original application secret or master key values; or correct the configuration. A value in
foundation4ai.secrets.envchanges through the change procedure in Secrets and keys. An image older than the database is replaced by the current version, as described in Backup, restore and upgrades. -
Actions to avoid. Generating a new master secret for an existing installation. Deleting the bundled PostgreSQL pod, which in the evaluation profile creates an empty database with a new system ID.
Access
Login returning HTTP 401
-
Symptoms.
POST /loginreturns HTTP 401, and the dashboard sign-in fails. -
Diagnosis.
curl -sS -X POST "$FOUNDATION4_URL/login" \-H "x-api-key: $FOUNDATION4_API_KEY" \-H "x-api-key-secret: $FOUNDATION4_API_SECRET"kubectl exec -n foundation4ai deploy/foundation4ai-api-server -c server -- \./foundation4ai database check-connectionThe
messageof the error body names the cause, as listed in Errors: a missing header, a key identifier that is not a UUID, orInvalid API Key or Secret. The connection check printsOk.when the configuration of the API server reaches PostgreSQL, andError: "Failed connecting to database"otherwise. -
Cause.
Invalid API Key or Secretcovers an unknown key, a wrong secret, an inactive or expired key, and a database failure while the API server checks the key. A license problem does not produce HTTP 401: an invalid license stops the API server at startup, and an expired license returns HTTP 403Invalid license: Expiredon the operations that check the license. -
Resolution. Send the key identifier in
x-api-keyand the secret inx-api-key-secret, as described in Authenticate. The master key identifier and secret are the values ofFOUNDATION4AI_APP_MASTER_KEYandFOUNDATION4AI_APP_MASTER_SECRETinfoundation4ai.secrets.env. When a correct key fails for every client, restore the database connection of the API server. -
Actions to avoid. Replacing the master key values in
foundation4ai.secrets.envto regain access, which stops the API server and the workers at the next restart.
Ingress configured but unreachable
-
Symptoms. Requests to the host name time out, or the ingress controller returns HTTP 404 or 503. Through a Gateway,
GET /returns a redirect to/dashboard. -
Diagnosis.
kubectl get ingressclasskubectl describe ingress foundation4ai -n foundation4aikubectl get endpoints foundation4ai-api-server foundation4ai-dashboard -n foundation4aikubectl get httproute foundation4ai -n foundation4ai -o yamlkubectl describe ingressshows the class, the host and the backends of the Ingress. An endpoints entry without addresses means that no ready pod serves the Service. The status of the HTTPRoute lists each parent with anAcceptedcondition;Falsewith a reason such asNoMatchingParentmeans that the parent reference matches no Gateway listener. -
Cause.
ingress.classNamenames a class that no controller serves. The request host differs fromingress.hosts, whose chart default ischart-example.local. No API server pod is ready. The defaulthttpRoute.parentRefs(Gatewaygateway, listenerhttp) match no Gateway. A network policy blocks the controller. Through the chart's HTTPRoute, an exact request for/redirects to/dashboard, soGET /never reaches the API. -
Resolution. Set
classNameto an IngressClass from the list, set the host names iningress.hostsorhttpRoute.hostnamesand send requests to those names, and sethttpRoute.parentRefsto the Gateway and the listener. Allow the controller namespace in the network policies described in Security hardening. Check the API withPOST /loginrather thanGET /. -
Actions to avoid. Changing the API server Service to
LoadBalancerorNodePortto bypass the controller, which also bypasses the TLS of the ingress.
Dashboard sign-in failure through a port forward
- Symptoms. A port forward to the dashboard Service shows the dashboard page, but the sign-in fails with a correct key.
- Diagnosis. The network panel of the browser's developer tools shows the sign-in request
POST /loginsent to the forwarded local address, where the dashboard server answers instead of the API server. - Cause. The dashboard is served under
/dashboard/and calls the API on the origin of the dashboard page, so the dashboard works only where one host name serves both/dashboardand the API. The chart's Ingress and HTTPRoute provide that host name; a port forward to one Service does not. - Resolution. Open the dashboard through the Ingress or the HTTPRoute host name, as described in API and dashboard access.
- Actions to avoid. Exposing the dashboard or the API server through a
LoadBalancerorNodePortService to replace the port forward, which bypasses the TLS of the ingress.
Ingestion
Documents remaining in pending status
-
Symptoms. The pending count does not fall, or documents remain
pendinglong after submission. -
Diagnosis.
kubectl logs -n foundation4ai -l app.kubernetes.io/name=api-server-worker -c worker --tail=500 --prefix \| grep -E "Error processing documents|Failed processing document"kubectl get pods -n foundation4ai -l app.kubernetes.io/name=api-server-workerkubectl logs -n foundation4ai deploy/foundation4ai-api-server-worker -c grpc --tail=100kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- \nats consumer info DOCUMENTS document-processorA line with
Failed processing documentnames one document, the pipeline and the reason. A line withError processing documents for pipelinenames a failure of a whole pipeline group, such as an error of the gRPC service, of the embedding model orExpired license. The same group error every 10 seconds for one pipeline means that the group returns to the queue on each attempt. The gRPC service log shows the errors of Python text splitters and embedding providers, and the consumer information shows the jobs that wait and the jobs delivered more than once. -
Cause. Processing fails for the document, for the text splitter or embedding model of the pipeline, or for the license, or no worker runs. A document that the text splitter or embedding model rejects fails every document of the same pipeline in the batch, and a group failure returns every job of the batch to the queue. The worker retries every 10 seconds, the document stays
pendingwhile processing fails, and the queue removes a job 24 hours after submission, after which the document is not processed again. -
Resolution. Remove the cause shown in the log: restore the gRPC service or the model server, install a valid license as described in the next entry, or start the workers. Documents that are
pendingfor less than 24 hours are then processed without action. Documents that arependingmore than 24 hours after creation are submitted again, as described in Resubmission. -
Actions to avoid. Submitting documents again while the cause persists, which adds versions and jobs that fail in the same way. Purging or recreating the
DOCUMENTSstream, which removes the jobs of every queued document and leaves those documentspending.
Processing stopped after license expiry
-
Symptoms. The worker log shows
Error processing documents for pipeline <id>: Expired licenseevery 10 seconds, operations that check the license return HTTP 403Invalid license: Expired, and restarted API server or worker containers stop at startup withExpired license. -
Diagnosis.
kubectl logs -n foundation4ai -l app.kubernetes.io/name=api-server-worker -c worker --tail=200 --prefix \| grep -c "Expired license"LICENSE=$(grep '^FOUNDATION4AI_APP_LICENSE=' foundation4ai.secrets.env | cut -d= -f2-)kubectl exec -n foundation4ai deploy/foundation4ai-api-server -c server -- \./foundation4ai license details "$LICENSE"A count above zero and the output
Expiredconfirm the cause. -
Cause. The worker checks the license before each pipeline group, so an expired license fails every group. The processes read the license at startup, so a renewed license takes effect only after a restart, and a process that restarts with an expired license stops at startup. The queue removes jobs 24 hours after submission, so documents submitted more than 24 hours before the renewal remain
pending. -
Resolution. Obtain a renewed license, as described in Licensing, and apply the license with the change procedure in Secrets and keys, which ends with a restart of the API server and the workers. The license check log of the application upgrade shows
Checking for a valid license... OK. Then submit again the documents that are stillpendingmore than 24 hours after creation, as described in Resubmission. -
Actions to avoid. Upgrading a release for other changes before the renewal, which stops at the license check. Purging the
DOCUMENTSstream.
Workers timing out on NATS JetStream
-
Symptoms. Worker containers restart at startup with
Error: "Failed to connect to Foundation4.ai: Nats error: <reason>", where the reason reports a JetStream timeout or that JetStream is unavailable. The API server may report the same reason afterFailed to create database connection:, and document submission returns HTTP 500NATS error. -
Diagnosis.
kubectl get pods -n foundation4ai -l app.kubernetes.io/component=natskubectl logs -n foundation4ai foundation4ai-core-nats-0 -c nats --tail=50kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats account infokubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats stream info DOCUMENTSThe 3 NATS pods are expected
Runningwith every container ready.nats account infoshows whether JetStream is available to the account and how much storage is used, andnats stream infoshows the server that holds the stream. A NATS log that reports no JetStream leader, or storage limits reached, points to the NATS cluster. -
Cause. At startup, each worker and API server opens or creates the
documentsobject store and theDOCUMENTSstream. JetStream answers only while a majority of the 3 NATS servers runs, and refuses new data when the JetStream volume is full, at 10 GiB per server by default. A network policy that blocks port 4222 or 6222 has the same effect. -
Resolution. Restore the NATS pods, free or enlarge the JetStream storage described in Deployment overview and requirements, and allow ports 4222 and 6222 in the network policies. Then restart the workers:
kubectl rollout restart deployment/foundation4ai-api-server-worker -n foundation4aikubectl rollout status deployment/foundation4ai-api-server-worker -n foundation4aiExpected result:
successfully rolled out, and the worker log shows no further NATS error. -
Actions to avoid. Deleting the NATS JetStream volume claims or the
DOCUMENTSstream, which removes the queued jobs. Reducing NATS to one replica on an existing installation, which can remove the server that holds the stream.
Worker restarts on queue errors
-
Symptoms. Worker containers restart during processing, not at startup, and the pod shows no
OOMKilledreason. -
Diagnosis.
kubectl get pods -n foundation4ai -l app.kubernetes.io/name=api-server-workerkubectl logs -n foundation4ai <worker-pod-name> -c worker --previous | tail -5The last line of the previous container is one of the following:
Error: "Error retrieving messages from message queue.",Error: "Error processing message: Decryption error: Invalid message queue document"orError: "Message not found for document <id> version <version>". -
Cause. The worker ends the process when a read from the queue fails, when a job message cannot be decoded, or when a processing result names a document that the batch does not contain, and Kubernetes restarts the container. A transient NATS JetStream failure causes one restart per worker. A message that no worker can decode returns to the queue after the acknowledgement timeout and can restart each worker in turn until the queue removes the message 24 hours after submission (inferred).
-
Resolution. After
Error retrieving messages from message queue., check NATS JetStream as in the previous entry; the workers resume after the restart. When a decoding error repeats, collect the worker logs and the output ofnats stream info DOCUMENTSandnats consumer info DOCUMENTS document-processor, and contact the Foundation4 provider. -
Actions to avoid. Purging the
DOCUMENTSstream, which removes every queued job. Scaling the workers to 0 for more than 24 hours during the investigation, after which the queue has removed every waiting job.
Load
Workers restarted for exceeding memory
-
Symptoms. Worker pods restart during ingestion, and
kubectl describe podshowsLast State: TerminatedwithReason: OOMKilledand exit code 137, or the pod status isEvicted. -
Diagnosis.
kubectl get pods -n foundation4ai -l app.kubernetes.io/name=api-server-worker \-o jsonpath='{range .items[*]}{.metadata.name}{": "}{range .status.containerStatuses[*]}{.name}={.lastState.terminated.reason}{" "}{end}{"\n"}{end}'kubectl top pods -n foundation4ai --containersworker=OOMKilledorgrpc=OOMKillednames the container that exceeded the memory limit.kubectl top, which needs the metrics server, shows the memory of each container during a batch. -
Cause. Each worker takes up to 100 jobs per batch, and the memory of a batch grows with the size of the documents and with the embedding model. FastEmbed models run in the worker process, and Python embedding providers and text splitters run in the gRPC service. The charts set no resources, and one
api-server.resourcesvalue applies to theserver,workerandgrpccontainers, each with a separate limit. A worker that stops during a batch leaves the jobs unacknowledged, the queue delivers the jobs again after the acknowledgement timeout, and the same batch can exceed the limit again. -
Resolution. Set requests and a memory limit that covers the largest container, then upgrade the application release. The values below are an example, not a sizing recommendation; Scaling and performance describes how to size the containers:
api-server:resources:requests:cpu: "1"memory: 2Gilimits:memory: 4Gi -
Actions to avoid. Removing the memory limit on shared nodes, which lets one batch take memory from other pods. Adding worker replicas to reduce memory use, which does not change the size of a batch.
Low ingestion throughput
-
Symptoms. The pending count grows during steady submission, and the message count of the
DOCUMENTSstream rises. -
Diagnosis. Read the stream state, then forward the metrics port of a worker in a separate terminal and read the worker metrics:
kubectl exec -n foundation4ai deploy/foundation4ai-core-nats-box -- nats stream info DOCUMENTSkubectl port-forward -n foundation4ai deploy/foundation4ai-api-server-worker 9090:9090curl -s http://localhost:9090/metrics | grep foundation4ai_document_queueThe counters
foundation4ai_document_queue_processed,_success,_skipand_failedcount documents. The values_processed_latency,_text_splitter_latencyand_embedding_latencyare running totals in seconds. Two readings a few minutes apart give the documents per second of that worker and the share of the time spent in the text splitter and the embedding model. A rising_failedcounter means that the backlog comes from failures, as described in the entry on pending documents. The port forward reaches one worker pod. -
Cause. Each worker processes one batch at a time, and the pipeline groups of a batch one after another. Throughput is limited by the number of workers, by the CPU available to the worker and gRPC service containers, and by the time the embedding model needs per fragment.
-
Resolution. Add worker replicas with
api-server.replicaCount.worker, set CPU requests for the containers, or use an embedding model that runs on a dedicated model server, as described in Scaling and performance. -
Actions to avoid. Adding API server replicas, which do not process documents. Submitting more documents than the workers process in 24 hours, because the queue removes jobs after 24 hours.
HTTP 500 responses under load
-
Symptoms. Under load, requests return HTTP 500 after about 10 seconds with the message
Database connection error: Failed to acquire connection from pool: Connection pool timed out, or withFailed to acquire connection from pool: Connection pool timed out. -
Diagnosis.
psql "<administrator-connection-url>" -c "SHOW max_connections;"psql "<administrator-connection-url>" -c \"SELECT usename, state, count(*) FROM pg_stat_activity WHERE datname = '<database>' GROUP BY 1, 2;"kubectl get deployments,hpa -n foundation4aiThe connection counts per state show whether the API server pools are full, and the replica counts give the number of pools.
-
Cause. Each API server process holds at most
database.pool_sizeconnections, 3 in the chart defaults, and a request waits up to 10 seconds for a free connection. Bursts or long searches exhaust the pool. Each additional API server replica adds a pool, so replicas, including those that an autoscaler adds, can reachmax_connectionsof PostgreSQL, which then refuses new connections. Each worker holds 1 connection. -
Resolution. Raise the pool size in a
configsfile, keep the total belowmax_connectionsas described in Database, then upgrade the application release and restart the API server:api-server:configs:f4ai-10-database-pool.yaml: |database:pool_size: 10helm upgrade foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mkubectl rollout restart deployment/foundation4ai-api-server -n foundation4aiExpected result: Helm reports
STATUS: deployed, thendeployment.apps/foundation4ai-api-server restarted. -
Actions to avoid. Raising the API server replica count or the autoscaler maximum without a connection budget, which moves the failure to PostgreSQL.
Upgrades and changes
Upgrade failure creating the fragment index
-
Symptoms. An upgrade of the application release fails in the pre-upgrade hook, and the migration Job log shows
Failed applying migrations:withduplicate key value violates unique constraint. -
Diagnosis. The first statement shows whether the migration
m20260827_140000_vector_embeddingsis still pending, which an empty result confirms. The second, read-only statement reports each pipeline table with duplicate fragment keys; the fragment tables still carry the pipeline identifier as name while the migration is pending. A deployment with a customdatabase.schemareplacespublic:psql "<administrator-connection-url>" -c \"SELECT version FROM public.schema_migrations WHERE version = 'm20260827_140000_vector_embeddings';"psql "<administrator-connection-url>" <<'SQL'DO $$DECLARE p record; n bigint;BEGINFOR p IN SELECT id FROM public.pipelines LOOPEXECUTE format('SELECT count(*) FROM (SELECT 1 FROM embeddings.%I GROUP BY document_id, version, order_id HAVING count(*) > 1) d', p.id::text) INTO n;IF n > 0 THEN RAISE NOTICE 'pipeline %: % duplicate fragment keys', p.id, n; END IF;END LOOP;END $$;SQLExpected result on an unaffected database:
DOand no notice. -
Cause. One migration of the upgrade creates a unique index on document, version and fragment order before the migration copies the existing fragments into the new table, so duplicate rows stop the copy. The migration runs in a transaction and is rolled back, and the failed hook leaves the previous version running.
-
Resolution. Keep the previous version running, collect the migration log and the output of the statement, and obtain a de-duplication procedure from the Foundation4 provider before the upgrade is repeated. Backup, restore and upgrades describes the backup before an upgrade.
-
Actions to avoid. Deleting fragment rows or dropping the index by hand. Running
helm rollbackafter a later migration has succeeded, because an older image does not start against a newer database schema (inferred).
Pods using a previous value after a change
-
Symptoms. After a renewed license, a new password, a new URL or a change of
api-server.configs, the API server and the workers behave as before. -
Diagnosis.
kubectl get pods -n foundation4ai -l app.kubernetes.io/instance=foundation4aihelm history foundation4ai -n foundation4aiThe
AGEof the API server and worker pods is older than the last upgrade in the release history. -
Cause. Foundation4 reads configuration once at startup. The charts write configuration and credentials to Secrets and a ConfigMap, and the Deployments carry no reference to the content of those objects, so an upgrade that changes only the Secrets or the ConfigMap does not restart the pods.
-
Resolution. Run the change procedure in Secrets and keys, which ends with the restart:
kubectl rollout restart -n foundation4ai \deployment/foundation4ai-api-server deployment/foundation4ai-api-server-workerExpected result: both Deployments report
restarted, and the new pods show the changed behavior. -
Actions to avoid. Editing or deleting the Secret
foundation4ai-api-serverdirectly. The next upgrade replaces the Secret, and pods fail to start while the Secret is missing.