Monitoring and logging
This page describes the metrics, queue state, health checks, logs and traces of a Foundation4 deployment, and the queries and alerts that an operator builds from these sources. Operators who run Foundation4 in production need this page, and operators who diagnose a problem use the commands. Troubleshooting lists the diagnoses by symptom.
Monitoring sources
The core release installs a Prometheus server. The charts install no Alertmanager, no Grafana, no log collector and no OpenTelemetry collector; a production deployment connects Foundation4 to the monitoring stack of the organization.
| Source | Endpoint | Collected by the bundled Prometheus |
|---|---|---|
| API server metrics | /metrics on port 8000 of each API server pod | Yes |
| Worker metrics | Port 9090 of each worker pod | No, until the worker scrape job of Worker scrape job is added |
| NATS JetStream | The nats command-line tool, and an optional exporter | No |
| gRPC service | None; the Kubernetes liveness probe uses the gRPC health service | Not applicable |
| Redis-compatible cache and PostgreSQL | No exporter installed by the charts | No |
| Logs | Container output of every component | Not applicable |
| Traces | OpenTelemetry export from the API server | Not applicable |
The API server serves /metrics on the API port. Security hardening describes how to restrict the route at the ingress.
Prometheus scrape scope
The core chart replaces the default jobs of the Prometheus chart with one job, kubernetes-pods. The job discovers pods in every namespace and keeps a pod only when the pod carries both annotations prometheus.io/app: foundation4ai and prometheus.io/scrape: "true". The job reads the path and port from the prometheus.io/path and prometheus.io/port annotations, scrapes every 10 seconds with a 2-second timeout, and ignores pods that are not running.
Each series carries the pod labels, with dots, slashes and hyphens replaced by underscores, and the labels namespace, pod and node. The label app_kubernetes_io_name separates the API server (api-server) from the workers (api-server-worker).
The API server and worker pods carry both annotations with port 8000. The worker serves metrics on port 9090, so the worker targets of kubernetes-pods report down and no worker metric reaches Prometheus. The core chart removes the default jobs through prometheus.scrapeConfigs: null, which the values file keeps, because a remaining default job named kubernetes-pods can stop Prometheus at startup.
Prometheus access and targets
-
Forward a local port to Prometheus in a separate terminal:
kubectl -n foundation4ai port-forward svc/foundation4ai-core-prometheus-server 9090:80Expected result:
Forwarding from 127.0.0.1:9090 -> 9090. The Prometheus user interface is athttp://localhost:9090. -
List the active targets and the health of each:
curl -s 'http://localhost:9090/api/v1/targets?state=active' | grep -o '"scrapeUrl":"[^"]*"\|"health":"[^"]*"'Expected result: one
scrapeUrlending in:8000/metricswith"health":"up"for each API server pod, and one for each worker pod with"health":"down"until the worker scrape job is added.
Worker scrape job
The following values add a job that scrapes port 9090 of the worker pods. The job is a proposal that awaits the chart correction and has not been tested.
-
Add the job to the values file, which the core release reads:
prometheus:extraScrapeConfigs: |- job_name: foundation4ai-workersscrape_interval: 10sscrape_timeout: 2skubernetes_sd_configs:- role: podnamespaces:names:- foundation4airelabel_configs:- source_labels: [__meta_kubernetes_pod_label_app_kubernetes_io_name]action: keepregex: api-server-worker- source_labels: [__meta_kubernetes_pod_container_port_number]action: keepregex: "9090"- source_labels: [__meta_kubernetes_pod_phase]action: dropregex: Pending|Succeeded|Failed|Completed- action: labelmapregex: __meta_kubernetes_pod_label_(.+)- source_labels: [__meta_kubernetes_namespace]target_label: namespace- source_labels: [__meta_kubernetes_pod_name]target_label: podExpected result: the values file holds
prometheus.extraScrapeConfigs. -
Upgrade the core release:
helm upgrade --install foundation4ai-core ./charts/foundation4ai-core \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployed. -
Repeat step 2 of Prometheus access and targets.
Expected result: one additional
scrapeUrlending in:9090/metricswith"health":"up"for each worker pod.
The worker targets of kubernetes-pods keep reporting down, so alerts on the up series use the foundation4ai-workers job for the workers.
API server metrics
Metrics is the complete metric catalog, with labels, label values and counting rules. The following tables summarize the metrics for the queries on this page.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
axum_http_requests_total | Counter | method, endpoint, status | HTTP requests handled. endpoint is the route template, such as /pipelines/{pipeline_id}/search. |
axum_http_requests_duration_seconds | Histogram | method, endpoint, status | Seconds from the arrival of the request to the response head, in buckets from 0.005 to 10 seconds. The time spent sending a streamed response is not included. |
axum_http_requests_pending | Gauge | method, endpoint | Requests in progress, including streamed responses that are still sending |
foundation4ai_document_queue_added | Counter | None | Document create requests that pass the API key check and the parsing of the body, counted before validation, including requests that then fail validation, the permission check or the license check |
The requests to /metrics are counted too.
Worker metrics
The worker exports the following metrics on port 9090. Every metric carries a label description with a fixed text, and the output has no help lines. The counter names have no _total suffix.
| Metric | Type | Meaning |
|---|---|---|
foundation4ai_document_queue_processed | Counter | Documents in every processed pipeline group, counted on every attempt, including retries |
foundation4ai_document_queue_success | Counter | Documents processed successfully, and documents deleted before processing |
foundation4ai_document_queue_skip | Counter | Jobs whose document was no longer pending, such as a job delivered again after the document was processed, and documents with empty text |
foundation4ai_document_queue_failed | Counter | Documents whose processing failed, counted on every attempt, including every document of a failed pipeline group |
foundation4ai_document_queue_processed_latency | Gauge | Seconds spent processing pipeline groups, accumulated since the worker started |
foundation4ai_document_queue_text_splitter_latency | Gauge | Seconds spent splitting text, accumulated over successful groups |
foundation4ai_document_queue_embedding_latency | Gauge | Seconds spent computing embeddings, accumulated over successful groups |
The three latency metrics are declared as gauges but only increase. Every value restarts at 0 when the worker process restarts.
Accumulating values
Every Foundation4 metric except axum_http_requests_pending accumulates from process start, so a raw value depends on the uptime of the pod. Dashboards and alerts use the following functions:
- Rates.
rate(<metric>[5m])gives the change per second over 5 minutes and treats a drop to 0 as a restart. - Totals over a period.
increase(<metric>[1h])gives the change over 1 hour. - Deployment totals.
sum(rate(<metric>[5m]))adds the rates of all pods. - Ratios. A latency rate divided by a document rate gives seconds per document.
rate and increase apply to the worker latency gauges in the same way as to counters, because the gauges only increase between restarts. Prometheus can add an informational note to such queries, because the metric names do not end in _total.
Recommended panels
| Panel | Query |
|---|---|
| Documents submitted per minute | sum(rate(foundation4ai_document_queue_added[5m])) * 60 |
| Documents processed per minute | sum(rate(foundation4ai_document_queue_success[5m])) * 60 |
| Failed attempts per minute | sum(rate(foundation4ai_document_queue_failed[5m])) * 60 |
| Skipped jobs per minute | sum(rate(foundation4ai_document_queue_skip[5m])) * 60 |
| Processing seconds per document | sum(rate(foundation4ai_document_queue_processed_latency[5m])) / sum(rate(foundation4ai_document_queue_processed[5m])) |
| Embedding share of processing time | sum(rate(foundation4ai_document_queue_embedding_latency[5m])) / sum(rate(foundation4ai_document_queue_processed_latency[5m])) |
| Splitting share of processing time | sum(rate(foundation4ai_document_queue_text_splitter_latency[5m])) / sum(rate(foundation4ai_document_queue_processed_latency[5m])) |
| API requests per second by status | sum by (status) (rate(axum_http_requests_total[5m])) |
| API server error ratio | sum(rate(axum_http_requests_total{status=~"5.."}[5m])) / sum(rate(axum_http_requests_total[5m])) |
| API latency, 95th percentile by route | histogram_quantile(0.95, sum by (le, endpoint) (rate(axum_http_requests_duration_seconds_bucket[5m]))) |
| API requests in progress | sum(axum_http_requests_pending) |
| Targets up | up{job=~"kubernetes-pods|foundation4ai-workers"} |
The panels that use worker metrics need the worker scrape job. The queue depth comes from NATS JetStream, as described in Queue state. No metric reports the database connection pool; the API server error ratio and the pool warning described in Log messages show pool pressure.
Recommended alerts
The thresholds depend on the workload and are set after a period of measurement:
| Alert | Condition |
|---|---|
| Processing failures | sum(increase(foundation4ai_document_queue_failed[15m])) > <failures> |
| Processing stopped | sum(rate(foundation4ai_document_queue_added[30m])) > 0 and sum(rate(foundation4ai_document_queue_success[30m])) == 0 |
| Worker target down | up{job="foundation4ai-workers"} == 0 for 5 minutes |
| API server target down | up{job="kubernetes-pods", app_kubernetes_io_name="api-server"} == 0 for 5 minutes |
| API server errors | The error ratio query of Recommended panels above <ratio> for 10 minutes |
| Oldest queued job | Age of the first message of the stream DOCUMENTS above <hours>, well below 24 hours, from Queue state |
The bundled Prometheus evaluates rules from prometheus.serverFiles.alerting_rules.yml, but the core chart disables Alertmanager, so the bundled installation delivers no alert. A production deployment evaluates these conditions in the monitoring stack of the organization, which scrapes the same targets or receives the series from the bundled Prometheus.
Queue state
The utility Deployment foundation4ai-core-nats-box carries the nats command-line tool. The stream DOCUMENTS holds the processing jobs, and the consumer document-processor delivers the jobs to the workers.
kubectl -n foundation4ai exec deploy/foundation4ai-core-nats-box -- \
nats --server nats://foundation4ai-core-nats:4222 stream info DOCUMENTS
kubectl -n foundation4ai exec deploy/foundation4ai-core-nats-box -- \
nats --server nats://foundation4ai-core-nats:4222 consumer info DOCUMENTS document-processor
Expected result: the stream information shows the number of messages and the sequence and time of the first message. The consumer information shows delivery and acknowledgement state. The field names below are those of the nats tool and can differ between versions of the tool:
| Field | Reading |
|---|---|
Stream Messages | Jobs in the queue, including jobs that a worker is processing. This value is the queue depth. |
Stream First Sequence time | Submission time of the oldest job. A job older than 24 hours is removed. |
Consumer Unprocessed Messages | Jobs not yet delivered to a worker. A value that rises during a load means that the workers process slower than the client application submits. |
Consumer Outstanding Acks | Jobs delivered to a worker and not yet acknowledged |
Consumer Redelivered Messages | Jobs delivered more than once, after a failure or an expired acknowledgement wait |
Consumer Last Ack time | Time of the last acknowledgement. An old time while jobs remain means that no worker completes jobs. |
Consumer Waiting Pulls | Fetch requests from workers. A value of 0 while workers should be running means that no worker is connected. |
The NATS chart includes an exporter that publishes NATS server and JetStream metrics, including stream and consumer state. The core scrape job keeps the NATS pods only with the Foundation4 annotations, which the following untested values add:
nats:
promExporter:
enabled: true
podTemplate:
merge:
metadata:
annotations:
prometheus.io/app: foundation4ai
prometheus.io/scrape: "true"
prometheus.io/port: "7777"
The metric names depend on the exporter version; the exporter lists the names at http://<nats pod>:7777/metrics.
Health checks
/healthz on the API server returns status 200 with every component reported as ok, without checking any component. The API server answers only after startup, which connects to PostgreSQL, the Redis-compatible cache and NATS JetStream and validates the license and the master key. A response from /healthz therefore means that the process started and serves HTTP, and the chart uses /healthz for the liveness and readiness probes of the server container.
The following checks test the components:
| Check | Command | Expected result |
|---|---|---|
| Database connection from an API server pod | kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- ./foundation4ai database check-connection | Ok. |
| Authenticated API request that reads the database | curl -s -o /dev/null -w "%{http_code}\n" -H "x-api-key: $FOUNDATION4_API_KEY" -H "x-api-key-secret: $FOUNDATION4_API_SECRET" "$FOUNDATION4_URL/pipelines?first=1" | 200 |
| Queue | The stream command of Queue state | The stream information, with no error |
| gRPC service | kubectl -n foundation4ai get pods | API server and worker pods 2/2 ready; the liveness probe calls the gRPC health service |
| Worker metrics listener | kubectl -n foundation4ai port-forward deploy/foundation4ai-api-server-worker 9091:9090, then curl -s http://localhost:9091/health | OK |
| Prometheus | With the port forward of Prometheus access and targets, curl -s http://localhost:9090/-/ready | A message that Prometheus is ready |
An external availability check uses the authenticated request with a dedicated API key that holds read permission on one pipeline. Manage API keys and permissions describes such keys. The /statistics route is not a monitoring source in the current release: most of the fields report 0. Metrics describes each field.
Logs
Every component writes logs to the container output only; no component writes log files. The retention of the logs follows the log collection of the cluster.
| Component | Command |
|---|---|
| API server | kubectl -n foundation4ai logs deploy/foundation4ai-api-server -c server |
| gRPC service in the API server pods | kubectl -n foundation4ai logs deploy/foundation4ai-api-server -c grpc |
| Workers, all pods | kubectl -n foundation4ai logs -l app.kubernetes.io/name=api-server-worker -c worker --prefix --tail=500 |
| Hook Jobs | kubectl -n foundation4ai logs job/foundation4ai-api-server-db-migration, and the same for foundation4ai-api-server-license-check and foundation4ai-api-server-create-admin-api-key |
| Dashboard | kubectl -n foundation4ai logs deploy/foundation4ai-dashboard |
| NATS JetStream | kubectl -n foundation4ai logs foundation4ai-core-nats-0 -c nats |
app.log_level sets the level of the API server and the workers, with the default warn; the gRPC service logs at INFO. Configuration reference describes the levels, the output formats and the procedure that changes the level. A production deployment keeps warn outside a limited diagnosis period, as described in Security hardening.
The API server returns an x-request-id header on the responses of the API routes. The exported traces carry the same value in the attribute request_id, and a log line that the API server writes while handling a request carries request_id=<value> in the prefix that names the HTTP request span. At warn, most requests produce no log line.
Log messages
| Component | Level | Message | Meaning |
|---|---|---|---|
| Worker | INFO | Successfully processed document <document> for pipeline <pipeline> | Document processed |
| Worker | WARN | Failed processing document <document> for pipeline <pipeline>: <reason> | One document failed; the job returns after 10 seconds |
| Worker | ERROR | Error processing documents for pipeline <pipeline>: <error> | A pipeline group failed; every job of the batch returns after 10 seconds |
| Worker | WARN | Skipped document <document> for pipeline <pipeline>: <status> | The document was no longer pending, or the document text was empty and the document is marked success |
| Worker | WARN | Document <document> for pipeline <pipeline> was removed before processing | The document was deleted before processing |
| Worker | ERROR | Failed to acknowledge message: <error> | The worker could not acknowledge a job to NATS JetStream |
| Worker | Process exit | Error retrieving messages from message queue., Error processing message: <error> | The worker process ends and Kubernetes restarts the container |
| API server and worker | WARN | Failed to validate license for system id: <system id> | License rejected at startup; the process stops |
| API server and worker | Process exit | Failed to parse config: <error> | A configuration key is missing or invalid |
| API server | Process exit | Failed to create database connection: <error> | Startup failed on the database, the license, the cache, NATS JetStream or the application secret |
| Worker | Process exit | Failed to connect to Foundation4.ai: <error> | The same startup failure in a worker |
| API server | Process exit | Failed to initialize Foundation4aiCore with master key | The master key check failed, as described in Secrets and keys |
| Worker | Process exit | Failed to initialize Foundation4.ai: <error> | The same master key failure in a worker |
| API server | ERROR | Failed to refresh license object count cache | The refresh of the license object counts, at startup and every 10 minutes, failed |
| API server | ERROR | Failed to refresh license document count cache | The refresh of the document and fragment counts, at startup and every 3 hours, failed, as described in Licensing |
| API server | WARN | no env var set or infered for OTEL_EXPORTER_OTLP_<SIGNAL>_PROTOCOL or OTEL_EXPORTER_OTLP_PROTOCOL; no <kind> exporter will be created | No OpenTelemetry exporter is configured. Three lines appear at each start of an API server without the variables described in OpenTelemetry tracing: LOGS with log, METRICS with metric and TRACES with span |
| API server | WARN | acquired connection, but time to acquire exceeded slow threshold | A request waited more than 2 seconds for a database connection (unverified) |
| gRPC service | INFO | gRPC server started on port 50051 | The gRPC service is ready |
| gRPC service | INFO | Added package directory to site: <path> | A package image was loaded |
The lines at INFO from the API server and the workers appear only when app.log_level is info, debug or trace.
OpenTelemetry tracing
The API server exports a span for each API request, with the request ID, and exports the log records through the OpenTelemetry Protocol (OTLP) when an OTLP variable is set. A request that carries trace context headers joins the trace of the caller. The workers, the gRPC service and the dashboard export no traces. The charts install no collector, so the collector runs in the cluster or elsewhere inside the network. Configuration reference lists the variables.
-
Add the variables to the values file of the application release. Values that contain credentials, such as
OTEL_EXPORTER_OTLP_HEADERS, go underapi-server.secretsinstead:api-server:configs:OTEL_EXPORTER_OTLP_ENDPOINT: "http://<collector host>:4317"OTEL_EXPORTER_OTLP_PROTOCOL: "grpc"OTEL_SERVICE_NAME: "foundation4ai-api-server"Expected result: the values file holds the three entries.
-
Upgrade the application release and restart the API server:
helm upgrade --install foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mkubectl -n foundation4ai rollout restart deploy/foundation4ai-api-serverkubectl -n foundation4ai rollout status deploy/foundation4ai-api-serverExpected result: Helm reports
STATUS: deployed, and the status command ends withsuccessfully rolled out. -
Confirm the variables in the
servercontainer:kubectl -n foundation4ai exec deploy/foundation4ai-api-server -c server -- env | grep '^OTEL_'Expected result: the three variables with the values from step 1.
-
Send a request and read the request ID from the response headers:
curl -s -D - -o /dev/null -X POST "$FOUNDATION4_URL/login" \-H "x-api-key: $FOUNDATION4_API_KEY" \-H "x-api-key-secret: $FOUNDATION4_API_SECRET" | grep -i '^x-request-id'Expected result: an
x-request-idheader. The trace backend of the collector shows a trace of servicefoundation4ai-api-serverwhose span carries the same value inrequest_id.