Skip to main content

Monitoring and logging

This page describes the metrics, queue state, health checks, logs and traces of a Foundation4 deployment, and the queries and alerts that an operator builds from these sources. Operators who run Foundation4 in production need this page, and operators who diagnose a problem use the commands. Troubleshooting lists the diagnoses by symptom.

Monitoring sources​

The core release installs a Prometheus server. The charts install no Alertmanager, no Grafana, no log collector and no OpenTelemetry collector; a production deployment connects Foundation4 to the monitoring stack of the organization.

SourceEndpointCollected by the bundled Prometheus
API server metrics/metrics on port 8000 of each API server podYes
Worker metricsPort 9090 of each worker podNo, until the worker scrape job of Worker scrape job is added
NATS JetStreamThe nats command-line tool, and an optional exporterNo
gRPC serviceNone; the Kubernetes liveness probe uses the gRPC health serviceNot applicable
Redis-compatible cache and PostgreSQLNo exporter installed by the chartsNo
LogsContainer output of every componentNot applicable
TracesOpenTelemetry export from the API serverNot applicable

The API server serves /metrics on the API port. Security hardening describes how to restrict the route at the ingress.

Prometheus scrape scope​

The core chart replaces the default jobs of the Prometheus chart with one job, kubernetes-pods. The job discovers pods in every namespace and keeps a pod only when the pod carries both annotations prometheus.io/app: foundation4ai and prometheus.io/scrape: "true". The job reads the path and port from the prometheus.io/path and prometheus.io/port annotations, scrapes every 10 seconds with a 2-second timeout, and ignores pods that are not running.

Each series carries the pod labels, with dots, slashes and hyphens replaced by underscores, and the labels namespace, pod and node. The label app_kubernetes_io_name separates the API server (api-server) from the workers (api-server-worker).

The API server and worker pods carry both annotations with port 8000. The worker serves metrics on port 9090, so the worker targets of kubernetes-pods report down and no worker metric reaches Prometheus. The core chart removes the default jobs through prometheus.scrapeConfigs: null, which the values file keeps, because a remaining default job named kubernetes-pods can stop Prometheus at startup.

Prometheus access and targets​

  1. Forward a local port to Prometheus in a separate terminal:

    kubectl -n foundation4ai port-forward svc/foundation4ai-core-prometheus-server 9090:80

    Expected result: Forwarding from 127.0.0.1:9090 -> 9090. The Prometheus user interface is at http://localhost:9090.

  2. List the active targets and the health of each:

    curl -s 'http://localhost:9090/api/v1/targets?state=active' | grep -o '"scrapeUrl":"[^"]*"\|"health":"[^"]*"'

    Expected result: one scrapeUrl ending in :8000/metrics with "health":"up" for each API server pod, and one for each worker pod with "health":"down" until the worker scrape job is added.

Worker scrape job​

The following values add a job that scrapes port 9090 of the worker pods. The job is a proposal that awaits the chart correction and has not been tested.

  1. Add the job to the values file, which the core release reads:

    prometheus:
    extraScrapeConfigs: |
    - job_name: foundation4ai-workers
    scrape_interval: 10s
    scrape_timeout: 2s
    kubernetes_sd_configs:
    - role: pod
    namespaces:
    names:
    - foundation4ai
    relabel_configs:
    - source_labels: [__meta_kubernetes_pod_label_app_kubernetes_io_name]
    action: keep
    regex: api-server-worker
    - source_labels: [__meta_kubernetes_pod_container_port_number]
    action: keep
    regex: "9090"
    - source_labels: [__meta_kubernetes_pod_phase]
    action: drop
    regex: Pending|Succeeded|Failed|Completed
    - action: labelmap
    regex: __meta_kubernetes_pod_label_(.+)
    - source_labels: [__meta_kubernetes_namespace]
    target_label: namespace
    - source_labels: [__meta_kubernetes_pod_name]
    target_label: pod

    Expected result: the values file holds prometheus.extraScrapeConfigs.

  2. Upgrade the core release:

    helm upgrade --install foundation4ai-core ./charts/foundation4ai-core \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed.

  3. Repeat step 2 of Prometheus access and targets.

    Expected result: one additional scrapeUrl ending in :9090/metrics with "health":"up" for each worker pod.

The worker targets of kubernetes-pods keep reporting down, so alerts on the up series use the foundation4ai-workers job for the workers.

API server metrics​

Metrics is the complete metric catalog, with labels, label values and counting rules. The following tables summarize the metrics for the queries on this page.

MetricTypeLabelsMeaning
axum_http_requests_totalCountermethod, endpoint, statusHTTP requests handled. endpoint is the route template, such as /pipelines/{pipeline_id}/search.
axum_http_requests_duration_secondsHistogrammethod, endpoint, statusSeconds from the arrival of the request to the response head, in buckets from 0.005 to 10 seconds. The time spent sending a streamed response is not included.
axum_http_requests_pendingGaugemethod, endpointRequests in progress, including streamed responses that are still sending
foundation4ai_document_queue_addedCounterNoneDocument create requests that pass the API key check and the parsing of the body, counted before validation, including requests that then fail validation, the permission check or the license check

The requests to /metrics are counted too.

Worker metrics​

The worker exports the following metrics on port 9090. Every metric carries a label description with a fixed text, and the output has no help lines. The counter names have no _total suffix.

MetricTypeMeaning
foundation4ai_document_queue_processedCounterDocuments in every processed pipeline group, counted on every attempt, including retries
foundation4ai_document_queue_successCounterDocuments processed successfully, and documents deleted before processing
foundation4ai_document_queue_skipCounterJobs whose document was no longer pending, such as a job delivered again after the document was processed, and documents with empty text
foundation4ai_document_queue_failedCounterDocuments whose processing failed, counted on every attempt, including every document of a failed pipeline group
foundation4ai_document_queue_processed_latencyGaugeSeconds spent processing pipeline groups, accumulated since the worker started
foundation4ai_document_queue_text_splitter_latencyGaugeSeconds spent splitting text, accumulated over successful groups
foundation4ai_document_queue_embedding_latencyGaugeSeconds spent computing embeddings, accumulated over successful groups

The three latency metrics are declared as gauges but only increase. Every value restarts at 0 when the worker process restarts.

Accumulating values​

Every Foundation4 metric except axum_http_requests_pending accumulates from process start, so a raw value depends on the uptime of the pod. Dashboards and alerts use the following functions:

  • Rates. rate(<metric>[5m]) gives the change per second over 5 minutes and treats a drop to 0 as a restart.
  • Totals over a period. increase(<metric>[1h]) gives the change over 1 hour.
  • Deployment totals. sum(rate(<metric>[5m])) adds the rates of all pods.
  • Ratios. A latency rate divided by a document rate gives seconds per document.

rate and increase apply to the worker latency gauges in the same way as to counters, because the gauges only increase between restarts. Prometheus can add an informational note to such queries, because the metric names do not end in _total.

PanelQuery
Documents submitted per minutesum(rate(foundation4ai_document_queue_added[5m])) * 60
Documents processed per minutesum(rate(foundation4ai_document_queue_success[5m])) * 60
Failed attempts per minutesum(rate(foundation4ai_document_queue_failed[5m])) * 60
Skipped jobs per minutesum(rate(foundation4ai_document_queue_skip[5m])) * 60
Processing seconds per documentsum(rate(foundation4ai_document_queue_processed_latency[5m])) / sum(rate(foundation4ai_document_queue_processed[5m]))
Embedding share of processing timesum(rate(foundation4ai_document_queue_embedding_latency[5m])) / sum(rate(foundation4ai_document_queue_processed_latency[5m]))
Splitting share of processing timesum(rate(foundation4ai_document_queue_text_splitter_latency[5m])) / sum(rate(foundation4ai_document_queue_processed_latency[5m]))
API requests per second by statussum by (status) (rate(axum_http_requests_total[5m]))
API server error ratiosum(rate(axum_http_requests_total{status=~"5.."}[5m])) / sum(rate(axum_http_requests_total[5m]))
API latency, 95th percentile by routehistogram_quantile(0.95, sum by (le, endpoint) (rate(axum_http_requests_duration_seconds_bucket[5m])))
API requests in progresssum(axum_http_requests_pending)
Targets upup{job=~"kubernetes-pods|foundation4ai-workers"}

The panels that use worker metrics need the worker scrape job. The queue depth comes from NATS JetStream, as described in Queue state. No metric reports the database connection pool; the API server error ratio and the pool warning described in Log messages show pool pressure.

The thresholds depend on the workload and are set after a period of measurement:

AlertCondition
Processing failuressum(increase(foundation4ai_document_queue_failed[15m])) > <failures>
Processing stoppedsum(rate(foundation4ai_document_queue_added[30m])) > 0 and sum(rate(foundation4ai_document_queue_success[30m])) == 0
Worker target downup{job="foundation4ai-workers"} == 0 for 5 minutes
API server target downup{job="kubernetes-pods", app_kubernetes_io_name="api-server"} == 0 for 5 minutes
API server errorsThe error ratio query of Recommended panels above <ratio> for 10 minutes
Oldest queued jobAge of the first message of the stream DOCUMENTS above <hours>, well below 24 hours, from Queue state

The bundled Prometheus evaluates rules from prometheus.serverFiles.alerting_rules.yml, but the core chart disables Alertmanager, so the bundled installation delivers no alert. A production deployment evaluates these conditions in the monitoring stack of the organization, which scrapes the same targets or receives the series from the bundled Prometheus.

Queue state​

The utility Deployment foundation4ai-core-nats-box carries the nats command-line tool. The stream DOCUMENTS holds the processing jobs, and the consumer document-processor delivers the jobs to the workers.

kubectl -n foundation4ai exec deploy/foundation4ai-core-nats-box -- \
nats --server nats://foundation4ai-core-nats:4222 stream info DOCUMENTS
kubectl -n foundation4ai exec deploy/foundation4ai-core-nats-box -- \
nats --server nats://foundation4ai-core-nats:4222 consumer info DOCUMENTS document-processor

Expected result: the stream information shows the number of messages and the sequence and time of the first message. The consumer information shows delivery and acknowledgement state. The field names below are those of the nats tool and can differ between versions of the tool:

FieldReading
Stream MessagesJobs in the queue, including jobs that a worker is processing. This value is the queue depth.
Stream First Sequence timeSubmission time of the oldest job. A job older than 24 hours is removed.
Consumer Unprocessed MessagesJobs not yet delivered to a worker. A value that rises during a load means that the workers process slower than the client application submits.
Consumer Outstanding AcksJobs delivered to a worker and not yet acknowledged
Consumer Redelivered MessagesJobs delivered more than once, after a failure or an expired acknowledgement wait
Consumer Last Ack timeTime of the last acknowledgement. An old time while jobs remain means that no worker completes jobs.
Consumer Waiting PullsFetch requests from workers. A value of 0 while workers should be running means that no worker is connected.

The NATS chart includes an exporter that publishes NATS server and JetStream metrics, including stream and consumer state. The core scrape job keeps the NATS pods only with the Foundation4 annotations, which the following untested values add:

nats:
promExporter:
enabled: true
podTemplate:
merge:
metadata:
annotations:
prometheus.io/app: foundation4ai
prometheus.io/scrape: "true"
prometheus.io/port: "7777"

The metric names depend on the exporter version; the exporter lists the names at http://<nats pod>:7777/metrics.

Health checks​

/healthz on the API server returns status 200 with every component reported as ok, without checking any component. The API server answers only after startup, which connects to PostgreSQL, the Redis-compatible cache and NATS JetStream and validates the license and the master key. A response from /healthz therefore means that the process started and serves HTTP, and the chart uses /healthz for the liveness and readiness probes of the server container.

The following checks test the components:

CheckCommandExpected result
Database connection from an API server podkubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- ./foundation4ai database check-connectionOk.
Authenticated API request that reads the databasecurl -s -o /dev/null -w "%{http_code}\n" -H "x-api-key: $FOUNDATION4_API_KEY" -H "x-api-key-secret: $FOUNDATION4_API_SECRET" "$FOUNDATION4_URL/pipelines?first=1"200
QueueThe stream command of Queue stateThe stream information, with no error
gRPC servicekubectl -n foundation4ai get podsAPI server and worker pods 2/2 ready; the liveness probe calls the gRPC health service
Worker metrics listenerkubectl -n foundation4ai port-forward deploy/foundation4ai-api-server-worker 9091:9090, then curl -s http://localhost:9091/healthOK
PrometheusWith the port forward of Prometheus access and targets, curl -s http://localhost:9090/-/readyA message that Prometheus is ready

An external availability check uses the authenticated request with a dedicated API key that holds read permission on one pipeline. Manage API keys and permissions describes such keys. The /statistics route is not a monitoring source in the current release: most of the fields report 0. Metrics describes each field.

Logs​

Every component writes logs to the container output only; no component writes log files. The retention of the logs follows the log collection of the cluster.

ComponentCommand
API serverkubectl -n foundation4ai logs deploy/foundation4ai-api-server -c server
gRPC service in the API server podskubectl -n foundation4ai logs deploy/foundation4ai-api-server -c grpc
Workers, all podskubectl -n foundation4ai logs -l app.kubernetes.io/name=api-server-worker -c worker --prefix --tail=500
Hook Jobskubectl -n foundation4ai logs job/foundation4ai-api-server-db-migration, and the same for foundation4ai-api-server-license-check and foundation4ai-api-server-create-admin-api-key
Dashboardkubectl -n foundation4ai logs deploy/foundation4ai-dashboard
NATS JetStreamkubectl -n foundation4ai logs foundation4ai-core-nats-0 -c nats

app.log_level sets the level of the API server and the workers, with the default warn; the gRPC service logs at INFO. Configuration reference describes the levels, the output formats and the procedure that changes the level. A production deployment keeps warn outside a limited diagnosis period, as described in Security hardening.

The API server returns an x-request-id header on the responses of the API routes. The exported traces carry the same value in the attribute request_id, and a log line that the API server writes while handling a request carries request_id=<value> in the prefix that names the HTTP request span. At warn, most requests produce no log line.

Log messages​

ComponentLevelMessageMeaning
WorkerINFOSuccessfully processed document <document> for pipeline <pipeline>Document processed
WorkerWARNFailed processing document <document> for pipeline <pipeline>: <reason>One document failed; the job returns after 10 seconds
WorkerERRORError processing documents for pipeline <pipeline>: <error>A pipeline group failed; every job of the batch returns after 10 seconds
WorkerWARNSkipped document <document> for pipeline <pipeline>: <status>The document was no longer pending, or the document text was empty and the document is marked success
WorkerWARNDocument <document> for pipeline <pipeline> was removed before processingThe document was deleted before processing
WorkerERRORFailed to acknowledge message: <error>The worker could not acknowledge a job to NATS JetStream
WorkerProcess exitError retrieving messages from message queue., Error processing message: <error>The worker process ends and Kubernetes restarts the container
API server and workerWARNFailed to validate license for system id: <system id>License rejected at startup; the process stops
API server and workerProcess exitFailed to parse config: <error>A configuration key is missing or invalid
API serverProcess exitFailed to create database connection: <error>Startup failed on the database, the license, the cache, NATS JetStream or the application secret
WorkerProcess exitFailed to connect to Foundation4.ai: <error>The same startup failure in a worker
API serverProcess exitFailed to initialize Foundation4aiCore with master keyThe master key check failed, as described in Secrets and keys
WorkerProcess exitFailed to initialize Foundation4.ai: <error>The same master key failure in a worker
API serverERRORFailed to refresh license object count cacheThe refresh of the license object counts, at startup and every 10 minutes, failed
API serverERRORFailed to refresh license document count cacheThe refresh of the document and fragment counts, at startup and every 3 hours, failed, as described in Licensing
API serverWARNno env var set or infered for OTEL_EXPORTER_OTLP_<SIGNAL>_PROTOCOL or OTEL_EXPORTER_OTLP_PROTOCOL; no <kind> exporter will be createdNo OpenTelemetry exporter is configured. Three lines appear at each start of an API server without the variables described in OpenTelemetry tracing: LOGS with log, METRICS with metric and TRACES with span
API serverWARNacquired connection, but time to acquire exceeded slow thresholdA request waited more than 2 seconds for a database connection (unverified)
gRPC serviceINFOgRPC server started on port 50051The gRPC service is ready
gRPC serviceINFOAdded package directory to site: <path>A package image was loaded

The lines at INFO from the API server and the workers appear only when app.log_level is info, debug or trace.

OpenTelemetry tracing​

The API server exports a span for each API request, with the request ID, and exports the log records through the OpenTelemetry Protocol (OTLP) when an OTLP variable is set. A request that carries trace context headers joins the trace of the caller. The workers, the gRPC service and the dashboard export no traces. The charts install no collector, so the collector runs in the cluster or elsewhere inside the network. Configuration reference lists the variables.

  1. Add the variables to the values file of the application release. Values that contain credentials, such as OTEL_EXPORTER_OTLP_HEADERS, go under api-server.secrets instead:

    api-server:
    configs:
    OTEL_EXPORTER_OTLP_ENDPOINT: "http://<collector host>:4317"
    OTEL_EXPORTER_OTLP_PROTOCOL: "grpc"
    OTEL_SERVICE_NAME: "foundation4ai-api-server"

    Expected result: the values file holds the three entries.

  2. Upgrade the application release and restart the API server:

    helm upgrade --install foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m
    kubectl -n foundation4ai rollout restart deploy/foundation4ai-api-server
    kubectl -n foundation4ai rollout status deploy/foundation4ai-api-server

    Expected result: Helm reports STATUS: deployed, and the status command ends with successfully rolled out.

  3. Confirm the variables in the server container:

    kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- env | grep '^OTEL_'

    Expected result: the three variables with the values from step 1.

  4. Send a request and read the request ID from the response headers:

    curl -s -D - -o /dev/null -X POST "$FOUNDATION4_URL/login" \
    -H "x-api-key: $FOUNDATION4_API_KEY" \
    -H "x-api-key-secret: $FOUNDATION4_API_SECRET" | grep -i '^x-request-id'

    Expected result: an x-request-id header. The trace backend of the collector shows a trace of service foundation4ai-api-server whose span carries the same value in request_id.