Skip to main content

Metrics

This page lists every metric that the Foundation4 processes export, with the exported name, the type, the labels, the emitting process, the event that changes the value and the meaning, and the fields of the /statistics route. Operators use this page to build dashboards and alerts, and integrators use this page to read processing counts during a load. Observability explains how the metrics relate to the other signals, and Monitoring and logging describes metric collection, queries and alerts.

Metric endpoints​

ProcessEndpointAccessContent
API serverGET /metrics on the API port, 8000 in the chartsNo API key. The route is served through the Service and the ingress like every API route.HTTP request metrics and foundation4ai_document_queue_added
WorkerPort 9090 of each worker pod, any path except /health; the charts use /metricsNo authentication. No Service exposes the port, so the port is reachable inside the cluster or through a port forward.Document processing metrics
Worker/health on port 9090Same as the metrics pathThe text OK
gRPC service and dashboardNoneNot applicableNot applicable

The worker listens on 0.0.0.0:9090, a fixed address that no configuration key changes. The listener starts before the worker connects to PostgreSQL and NATS JetStream, and serves the metrics for the life of the process. The liveness probe of the worker container in the charts reads /metrics on port 9090 every 30 seconds, so the probe tests the listener and not the processing of jobs.

The pod annotations of the charts name port 8000 for the API server and for the workers, so the bundled Prometheus server collects the worker metrics only through the scrape job described in Worker scrape job. Security hardening describes how to restrict /metrics and /statistics at the ingress.

Exposition format​

Both endpoints return the Prometheus text exposition format, with the content type text/plain; charset=utf-8 from the API server and text/plain from the workers. The following rules apply to every metric on this page:

  • Names. The exporter writes each name exactly as the code defines the name and adds no suffix. The foundation4ai_* counters therefore have no _total suffix, and axum_http_requests_total ends in _total because the suffix is part of the defined name.
  • Metadata. Each metric has a # TYPE line and no # HELP line.
  • Lifetime. Every value except axum_http_requests_pending accumulates from the start of the process, and every value returns to 0 when the process restarts. A series remains in the output for the life of the process.
  • First appearance. The foundation4ai_* series appear with the value 0 when the process starts. An HTTP series appears after the first request with the label values of the series.
  • Histograms. A histogram exports _bucket series with the label le, and _sum and _count series.

API server metrics​

The API server exports the following metrics on /metrics:

MetricTypeLabelsChangeMeaning
axum_http_requests_totalCountermethod, status, endpointIncreases by 1 when the API server returns the response head of a requestRequests handled, including requests rejected for a missing or invalid API key and requests to /metrics
axum_http_requests_duration_secondsHistogrammethod, status, endpointRecords one value when the API server returns the response headSeconds from the arrival of the request to the response head, in buckets of 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5 and 10 seconds. For a streamed response, the time spent sending the stream is not included.
axum_http_requests_pendingGaugemethod, endpointIncreases by 1 when a request arrives, and decreases by 1 when the response body is complete or the connection closesRequests in progress, including streamed responses that are still sending
foundation4ai_document_queue_addedCounterNoneSet to 0 at startup. Increases by 1 when a request to POST /documents or POST /pipelines/{pipeline_id}/documents passes the API key check and the parsing of the body, before any validation.Document create requests received since the API server started, including requests that then fail validation, the permission check or the license check

HTTP label values​

The HTTP labels take the following values:

LabelValues
methodThe HTTP method in upper case: GET, HEAD, POST, PUT, PATCH, DELETE, OPTIONS, CONNECT or TRACE. Any other method reports an empty value.
statusThe three-digit status code of the response head, such as 200, 401 or 500
endpointFor a request to an API operation, the route template as the endpoint reference shows the path, with the parameter names in braces, such as /pipelines/{pipeline_id}/search or /tracing/{execution_id}. The other routes report /, /healthz, /metrics, /statistics, /mcp, /openapi.json, /docs, /_docs and /_docs-swagger.

The following lines show the form of the API server output:

# TYPE axum_http_requests_total counter
axum_http_requests_total{method="POST",status="201",endpoint="/pipelines/{pipeline_id}/documents"} 1180
# TYPE foundation4ai_document_queue_added counter
foundation4ai_document_queue_added 1184

Worker metrics​

Each worker exports the following metrics on port 9090. A worker fetches up to 100 jobs at a time and processes the jobs of each pipeline as one group. Every worker metric carries the label description with the fixed value in the table.

MetricTypedescription valueChangeMeaning
foundation4ai_document_queue_processedCounterNumber of documents processedIncreases by the number of documents of a pipeline group each time the worker processes the group, before the outcome is knownDocuments processed, counted on every attempt, including retries and failed attempts
foundation4ai_document_queue_successCounterNumber of documents successfully processedIncreases by 1 for each document whose fragments the worker stored, and for each document deleted before processingDocuments completed
foundation4ai_document_queue_skipCounterNumber of documents skippedIncreases by 1 for each document that is no longer pending, and for each document with empty textJobs completed without processing
foundation4ai_document_queue_failedCounterNumber of documents failedIncreases by 1 for each document that fails, and by the number of documents of the group when the whole pipeline group failsFailed attempts
foundation4ai_document_queue_processed_latencyGaugeTotal latency of documents processedIncreases by the seconds of each attempt to process a pipeline group, successful or failedSeconds spent processing since the worker started
foundation4ai_document_queue_text_splitter_latencyGaugeTotal latency of text splitter computationIncreases by the seconds spent splitting the text of a pipeline group that the worker processed without a group failureSeconds spent splitting text
foundation4ai_document_queue_embedding_latencyGaugeTotal latency of embedding computationIncreases by the seconds spent computing the embeddings of a pipeline group, for every embedding model of the pipeline, when the worker processed the group without a group failureSeconds spent computing embeddings

Latency gauges​

The three latency metrics are gauges that only increase between restarts, so the Prometheus functions rate and increase apply to the latency metrics as to counters. Monitoring and logging describes the queries.

The following lines show the form of the worker output:

# TYPE foundation4ai_document_queue_processed counter
foundation4ai_document_queue_processed{description="Number of documents processed"} 1200
# TYPE foundation4ai_document_queue_processed_latency gauge
foundation4ai_document_queue_processed_latency{description="Total latency of documents processed"} 482.61

Processing outcomes​

Each document of a pipeline group ends in one of the following outcomes. Every outcome increases foundation4ai_document_queue_processed by 1 for the document.

OutcomeCounterWorker log lineJob
Fragments storedsuccessINFO Successfully processed document <document> for pipeline <pipeline>Acknowledged
Empty document textskipWARN Skipped document <document> for pipeline <pipeline>: pending. The document status becomes success.Acknowledged
Document no longer pending, such as a job delivered again after processingskipWARN Skipped document <document> for pipeline <pipeline>: <status>Acknowledged
Document deleted before processingsuccessWARN Document <document> for pipeline <pipeline> was removed before processingAcknowledged
Document failedfailedWARN Failed processing document <document> for pipeline <pipeline>: <reason>Returned to the queue and delivered again after 10 seconds
Pipeline group failedfailed, for every document of the groupERROR Error processing documents for pipeline <pipeline>: <error>Every job of the batch returned to the queue and delivered again after 10 seconds

A pipeline group fails as a whole when the worker cannot complete a step that covers every document of the group, such as the license check, the pipeline lookup, text splitting, the embedding computation or the database transaction. The INFO line appears only when app.log_level is info, debug or trace.

Labels added at collection​

The scrape job kubernetes-pods of the core chart adds the labels job, instance, namespace, pod and node and every pod label, with dots, slashes and hyphens replaced by underscores. The label app_kubernetes_io_name separates the API server (api-server) from the workers (api-server-worker). The worker pods carry fewer pod labels than the API server pod, so app_kubernetes_io_managed_by, app_kubernetes_io_version and helm_sh_chart appear on the API server series only. The job keeps the labels of the exported series when the names collide. Prometheus scrape scope describes the job.

Statistics route​

GET /statistics returns document processing and agent execution figures for a period, read from Prometheus, and the number of jobs in the queue. The route requires no API key, answers on the API port and is not used by the dashboard.

The query parameter interval is required and holds the length of the period, which ends at the time of the request. The value is a Prometheus duration: one or more of <n>y, <n>w, <n>d, <n>h, <n>m, <n>s and <n>ms, from the largest unit to the smallest, such as 15m, 1h30m or 24h. The value 0 is accepted as well.

Response fields​

The response contains the following fields. In the current release, only document_queue carries a value; the other fields report 0 or an empty list, so a period without activity and missing data look the same.

FieldTypeMeaning
documents_addedNumberDocuments submitted in the period
documents_processedNumberDocuments processed in the period
documents_processed_latencyNumberAverage processing time in seconds
documents_failedNumberFailed processing attempts in the period
agent_executionsNumberAgent executions in the period
agent_execution_latencyNumberAverage agent execution time in seconds
agent_execution_latenciesArray of pairs of a bucket bound and a countDistribution of agent execution times
document_queueIntegerNumber of messages in the NATS JetStream stream named by nats.stream, DOCUMENTS in the charts: the jobs waiting in the queue and the jobs that a worker is processing
document_queue_processed_latenciesArray of pairs of a bucket bound and a countDistribution of processing times

The following response is the form that the route returns in the current release:

{
"documents_added": 0.0,
"documents_processed": 0.0,
"documents_processed_latency": 0.0,
"documents_failed": 0.0,
"agent_executions": 0.0,
"agent_execution_latency": 0.0,
"agent_execution_latencies": [],
"document_queue": 42,
"document_queue_processed_latencies": []
}

Errors​

The route returns the following errors:

StatusBodyCause
400JSON error body with the message Invalid interval format: <value>. Must be a valid Prometheus duration (e.g., 15m, 1h, 24h).interval is not a Prometheus duration
400Plain textThe request has no interval parameter
500JSON error body with the message Failed to get stream info: <reason>The API server could not read the stream state from NATS JetStream

The queue figures of Queue state and the Prometheus queries of Recommended panels replace the route for monitoring.