Skip to main content

Observability

This page describes the signals that a Foundation4 deployment produces: request identifiers, logs, Prometheus metrics, OpenTelemetry traces and log records, agent execution traces, the health check and the statistics route. The page explains what each signal is for and how the signals relate for a single API request and for a document on the way through the queue. Integrators need this page to report and diagnose failed requests, and operators and security reviewers need this page to plan monitoring. Monitoring and logging describes the operator procedures, and Metrics lists every metric.

Signals​

SignalSourceRead fromPurpose
Request identifierAPI serverx-request-id response headerIdentifies one API request in a problem report, a trace and the API server logs
Trace contextAPI servertraceparent response headerConnects the request to the distributed trace of the client application
LogsEvery componentContainer outputStartup, license, connection and document processing events
Prometheus metricsAPI server and workers/metrics on port 8000 of the API server, port 9090 of each workerRequest rates, error ratios, latency and document processing counts
OpenTelemetry traces and log recordsAPI serverAn OpenTelemetry collectorOne span for each API request, with the request identifier
Agent execution traceAPI serverGET /tracing/{id}The fragments and messages behind one agent answer, for 60 minutes
Health checkAPI serverGET /healthzConfirms that the API server process started and serves HTTP
StatisticsAPI serverGET /statisticsNot a monitoring source in the current release

The gRPC service and the dashboard produce logs only. NATS JetStream, the queue between the API server and the workers, holds the queue state, which the nats command-line tool reads as described in Monitoring and logging.

Request identifiers​

The API server assigns a request identifier to every request that reaches an API operation, including the Model Context Protocol (MCP) endpoint /mcp and /statistics. When the request carries an x-request-id header, the API server keeps the value as sent; otherwise the API server generates a random universally unique identifier (UUID). The API server does not check a value that the client application sends, so a client application that sets the header uses a value that is unique for each request.

The identifier then follows the request through the API server:

  • Response. The API server returns the identifier in the x-request-id response header, on success and on error, including HTTP 401 responses.
  • Trace. The API server records the identifier in the request_id attribute of the span of the request, which the API server exports when OpenTelemetry export is enabled.
  • Log lines. A log line that the API server writes while handling the request carries request_id="<value>", with the value in quotation marks, in the line prefix that names the HTTP request span. The API server writes the log with terminal color codes, so a search for one request matches the identifier value alone, such as grep 3f1c9a52-7d4e-4b8a-9c61-2e5f0a7b8d13.

Responses from the root route /, the health check, /metrics, the API documentation routes and unknown paths carry no request identifier, as API conventions describes.

The request identifier stays inside the API server. The job that the API server places on the queue for a new document holds the document identifier, the version and the pipeline identifier only, so the worker logs name documents and pipelines, not requests. The gRPC service receives no request identifier either.

Trace context​

The API routes return a traceparent header in the W3C Trace Context format, with the trace identifier and the span identifier of the request span. A client application that sends a traceparent header makes the request span a child of the span that the header names, so the API server span joins the distributed trace of the client application. The API server returns the header whether or not OpenTelemetry export is enabled.

The following request sets a request identifier and prints the two headers. The shell variables are described in Authenticate:

curl -s -D - -o /dev/null "$FOUNDATION4_URL/pipelines/$PIPELINE_ID" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "x-request-id: 3f1c9a52-7d4e-4b8a-9c61-2e5f0a7b8d13" | grep -i -E '^(x-request-id|traceparent):'

The output shows the identifier as sent and the trace context of the request. The trace identifier and the span identifier differ on every request:

x-request-id: 3f1c9a52-7d4e-4b8a-9c61-2e5f0a7b8d13
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

Logs​

Every component writes logs to the container output, and no component writes log files. app.log_level sets the level of the API server and the workers, with the default warn; the gRPC service logs at INFO.

  • API server. At warn, the API server logs startup failures, license events and warnings from the libraries that the API server uses. The API server writes no log line for each request and no log line for a failed request: the error text of a failed request appears in the response body only.
  • Workers. At warn, a worker logs each document that fails processing, with the reason, each skipped document and each failure that affects all documents of one pipeline in a batch. At info, a worker also logs each document that the worker processes successfully. Each of these lines names the pipeline, and a line about one document also names the document.
  • gRPC service. The gRPC service logs startup and each package image that the service loads.

Configuration reference describes the levels and the output format of each process. Monitoring and logging lists the commands that read the logs and the meaning of each log message, and Security hardening describes the level of a production deployment.

Prometheus metrics​

Two processes export Prometheus metrics:

  • API server. /metrics on the API port 8000 reports HTTP requests by route, method and status, request durations, requests in progress and the number of document create requests.
  • Workers. Each worker serves metrics on port 9090 and reports the documents processed, completed, skipped and failed, and the seconds spent processing documents, splitting text and computing embeddings.

Each value covers one process since the process started, and no metric carries a pipeline, API key or document label. A deployment total is therefore the sum over the pods, and a rate over a period uses the Prometheus functions rate and increase. The metrics routes answer without an API key; Security hardening describes the restriction at the ingress.

The pod annotations of the charts direct the bundled Prometheus server to port 8000 of the worker pods, while the workers serve metrics on port 9090. Worker metrics therefore reach Prometheus only through the additional scrape job described in Worker scrape job. Metrics lists each metric with type, labels and counting rules.

OpenTelemetry traces and log records​

The API server exports traces and log records through the OpenTelemetry Protocol (OTLP) when the environment of the API server sets an OTLP endpoint or protocol variable, as listed in Configuration reference. Without these variables, the API server exports nothing. The workers, the gRPC service and the dashboard export no OpenTelemetry data, and the charts install no collector.

  • Request span. Each request to an API operation produces one span, named after the method and the route template, such as POST /pipelines/{pipeline_id}/search. The span records the method, the route template, the path, the response status code and the request identifier in request_id. A response with a 5xx status marks the span as an error.
  • Child spans. The spans below the request span follow app.log_level. At warn, a trace holds the request span only. At info, the trace adds spans for the route handler and for internal functions, and at debug and trace also spans for database statements. A production deployment keeps warn, as Security hardening describes.
  • Log records. The exported log records are the log lines of the API server at the same level. A record written during a request carries the trace identifier and the span identifier of the request (inferred).
  • Metrics. Metrics stay on the Prometheus endpoints. An OpenTelemetry collector receives no metrics from Foundation4 (inferred).

Monitoring and logging describes how to enable the export and how to find the span of a request.

Agent execution traces​

An agent execution trace is a record that the client application requests, reads through the API and uses to see the fragments and messages behind one agent answer. An agent execution trace is separate from the OpenTelemetry traces of the API server. Agents and prompt templates describes how an execution requests a trace.

  • Creation. An execution with tracing set to true stores the trace after the retrieval and the insertion of the values into the templates, and before the call to the LLM. A successful execution returns the trace identifier in the x-foundation4ai-tracing-id response header.
  • Retention. The Redis-compatible cache keeps the trace for 60 minutes after the execution. Reading the trace does not extend the period.
  • Access. Only the API key that ran the execution reads the trace, with GET /tracing/{id}. Another API key, and any request after 60 minutes, receives HTTP 404 with the message Trace with id <id> not found.
  • Storage. The cache holds the trace encrypted with the application secret of the deployment.

A trace contains three fields:

FieldContent
tracesAn object that maps the name of each retrieval placeholder to the fragments that the placeholder retrieved, in rank order, each with metadata, text and score
promptThe messages as sent to the LLM, each with role and template, after the input values and the fragment text were inserted. Messages with include set to false are absent.
inputsThe input values of the execution request, by query placeholder name

Trace interpretation​

The following trace comes from an execution of the support-answers agent of Agents and prompt templates in which the similarity search returned one fragment:

{
"traces": {
"context": [
{
"id": "0192f8c4-5d6e-7f80-9a1b-2c3d4e5f6a7b",
"classification": "internal",
"metadata": {"title": "Database credential rotation"},
"document_id": "0192f8c3-1a2b-7c3d-8e4f-5a6b7c8d9e0f",
"page_content": "Database credentials rotate every 30 days. An administrator starts an early rotation from the Credentials page of the operations console.",
"version": 1790740800000000,
"expired_at": null,
"score": 0.1834,
"order_id": 0
}
]
},
"prompt": [
{
"role": "system",
"template": "Answer using only the reference text. If the reference text does not contain the answer, say that the answer is not available."
},
{
"role": "user",
"template": "Reference text:\nDatabase credentials rotate every 30 days. An administrator starts an early rotation from the Credentials page of the operations console.\n\nQuestion: How do I rotate the database credentials?"
}
],
"inputs": {"question": "How do I rotate the database credentials?"}
}

A trace answers the following questions about an answer:

  • Retrieval result. An empty list under a placeholder means that the search found no fragment for the classifications and the filter of the request.
  • Match quality. The score of each fragment is the cosine distance from the target value, so a lower score is a closer match, as described in Search and retrieval.
  • Model input. The prompt messages show the exact text that the LLM received, so a missing fact in the answer can be traced to the retrieval or to the model.

The trace does not contain the answer of the LLM and does not contain the request identifier. A client application that keeps the x-request-id and the x-foundation4ai-tracing-id headers of an execution response can connect the agent execution trace with the OpenTelemetry trace of the same request.

Health check and statistics​

  • /healthz. The health check returns HTTP 200 with every component reported as ok, without checking PostgreSQL, NATS JetStream, the Redis-compatible cache or the gRPC service. The API server serves HTTP only after startup, which connects to PostgreSQL, the cache and NATS JetStream, so a response means that the process started and serves requests. Monitoring and logging lists checks that test each component.
  • /statistics. The statistics route returns document processing and agent execution figures for a period, read from Prometheus, and the number of jobs in the queue. In the current release, only the number of jobs reflects the deployment, and the other fields report 0 or an empty list. Metrics describes each field.

Signals of one request​

The following signals describe one API request:

SignalLocationLink to the request
Response headersClient applicationx-request-id, traceparent, and x-foundation4ai-tracing-id for an agent execution with tracing
Request spanTrace backend of the collectorThe request_id attribute equals the x-request-id value, and the trace identifier equals the second field of traceparent
API server log linesContainer output of the API serverrequest_id="<value>" in the line prefix. At warn, most requests produce no line.
HTTP metricsPrometheusCounted under the route template, the method and the status. No metric identifies a single request.
Agent execution traceGET /tracing/{id}The identifier from x-foundation4ai-tracing-id

A report of a failed request therefore quotes the status, the complete response body and the x-request-id value, as Errors describes. The response body holds the error text, and the request span shows the route, the status and the duration of the request.

Signals of a document​

A document passes through the API server, the queue and a worker. The request identifier ends at the queue; from the queue on, the document identifier and the pipeline identifier link the signals.

StageComponentSignalsIdentifier
SubmissionAPI serverRequest identifier, request span and foundation4ai_document_queue_added. The response holds the document id, the version and the status pending.Request identifier
QueueNATS JetStreamThe job in the stream DOCUMENTS: the queue depth and the age of the oldest job. The document_queue field of /statistics reports the queue depth.Document identifier and version
ProcessingWorkerfoundation4ai_document_queue_processed for each attempt, then success, skip or failed, and the processing, splitting and embedding seconds. A log line names the document, the pipeline and, on failure, the reason.Document identifier and pipeline identifier
ResultAPI serverThe document status, which the client application readsDocument identifier

A failed attempt returns the job to the queue, which redelivers the job after 10 seconds, so one document can increase the processed and failed counts several times. The worker metrics count documents for each worker process without document labels: the metrics show throughput and failure rates, and the worker log names the failing documents. Documents, versions and fragments describes the processing lifecycle, and Ingest documents reliably describes how a client application follows document status.