Metrics
This page lists every metric that the Foundation4 processes export, with the exported name, the type, the labels, the emitting process, the event that changes the value and the meaning, and the fields of the /statistics route. Operators use this page to build dashboards and alerts, and integrators use this page to read processing counts during a load. Observability explains how the metrics relate to the other signals, and Monitoring and logging describes metric collection, queries and alerts.
Metric endpoints
| Process | Endpoint | Access | Content |
|---|---|---|---|
| API server | GET /metrics on the API port, 8000 in the charts | No API key. The route is served through the Service and the ingress like every API route. | HTTP request metrics and foundation4ai_document_queue_added |
| Worker | Port 9090 of each worker pod, any path except /health; the charts use /metrics | No authentication. No Service exposes the port, so the port is reachable inside the cluster or through a port forward. | Document processing metrics |
| Worker | /health on port 9090 | Same as the metrics path | The text OK |
| gRPC service and dashboard | None | Not applicable | Not applicable |
The worker listens on 0.0.0.0:9090, a fixed address that no configuration key changes. The listener starts before the worker connects to PostgreSQL and NATS JetStream, and serves the metrics for the life of the process. The liveness probe of the worker container in the charts reads /metrics on port 9090 every 30 seconds, so the probe tests the listener and not the processing of jobs.
The pod annotations of the charts name port 8000 for the API server and for the workers, so the bundled Prometheus server collects the worker metrics only through the scrape job described in Worker scrape job. Security hardening describes how to restrict /metrics and /statistics at the ingress.
Exposition format
Both endpoints return the Prometheus text exposition format, with the content type text/plain; charset=utf-8 from the API server and text/plain from the workers. The following rules apply to every metric on this page:
- Names. The exporter writes each name exactly as the code defines the name and adds no suffix. The
foundation4ai_*counters therefore have no_totalsuffix, andaxum_http_requests_totalends in_totalbecause the suffix is part of the defined name. - Metadata. Each metric has a
# TYPEline and no# HELPline. - Lifetime. Every value except
axum_http_requests_pendingaccumulates from the start of the process, and every value returns to 0 when the process restarts. A series remains in the output for the life of the process. - First appearance. The
foundation4ai_*series appear with the value 0 when the process starts. An HTTP series appears after the first request with the label values of the series. - Histograms. A histogram exports
_bucketseries with the labelle, and_sumand_countseries.
API server metrics
The API server exports the following metrics on /metrics:
| Metric | Type | Labels | Change | Meaning |
|---|---|---|---|---|
axum_http_requests_total | Counter | method, status, endpoint | Increases by 1 when the API server returns the response head of a request | Requests handled, including requests rejected for a missing or invalid API key and requests to /metrics |
axum_http_requests_duration_seconds | Histogram | method, status, endpoint | Records one value when the API server returns the response head | Seconds from the arrival of the request to the response head, in buckets of 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5 and 10 seconds. For a streamed response, the time spent sending the stream is not included. |
axum_http_requests_pending | Gauge | method, endpoint | Increases by 1 when a request arrives, and decreases by 1 when the response body is complete or the connection closes | Requests in progress, including streamed responses that are still sending |
foundation4ai_document_queue_added | Counter | None | Set to 0 at startup. Increases by 1 when a request to POST /documents or POST /pipelines/{pipeline_id}/documents passes the API key check and the parsing of the body, before any validation. | Document create requests received since the API server started, including requests that then fail validation, the permission check or the license check |
HTTP label values
The HTTP labels take the following values:
| Label | Values |
|---|---|
method | The HTTP method in upper case: GET, HEAD, POST, PUT, PATCH, DELETE, OPTIONS, CONNECT or TRACE. Any other method reports an empty value. |
status | The three-digit status code of the response head, such as 200, 401 or 500 |
endpoint | For a request to an API operation, the route template as the endpoint reference shows the path, with the parameter names in braces, such as /pipelines/{pipeline_id}/search or /tracing/{execution_id}. The other routes report /, /healthz, /metrics, /statistics, /mcp, /openapi.json, /docs, /_docs and /_docs-swagger. |
The following lines show the form of the API server output:
# TYPE axum_http_requests_total counter
axum_http_requests_total{method="POST",status="201",endpoint="/pipelines/{pipeline_id}/documents"} 1180
# TYPE foundation4ai_document_queue_added counter
foundation4ai_document_queue_added 1184
Worker metrics
Each worker exports the following metrics on port 9090. A worker fetches up to 100 jobs at a time and processes the jobs of each pipeline as one group. Every worker metric carries the label description with the fixed value in the table.
| Metric | Type | description value | Change | Meaning |
|---|---|---|---|---|
foundation4ai_document_queue_processed | Counter | Number of documents processed | Increases by the number of documents of a pipeline group each time the worker processes the group, before the outcome is known | Documents processed, counted on every attempt, including retries and failed attempts |
foundation4ai_document_queue_success | Counter | Number of documents successfully processed | Increases by 1 for each document whose fragments the worker stored, and for each document deleted before processing | Documents completed |
foundation4ai_document_queue_skip | Counter | Number of documents skipped | Increases by 1 for each document that is no longer pending, and for each document with empty text | Jobs completed without processing |
foundation4ai_document_queue_failed | Counter | Number of documents failed | Increases by 1 for each document that fails, and by the number of documents of the group when the whole pipeline group fails | Failed attempts |
foundation4ai_document_queue_processed_latency | Gauge | Total latency of documents processed | Increases by the seconds of each attempt to process a pipeline group, successful or failed | Seconds spent processing since the worker started |
foundation4ai_document_queue_text_splitter_latency | Gauge | Total latency of text splitter computation | Increases by the seconds spent splitting the text of a pipeline group that the worker processed without a group failure | Seconds spent splitting text |
foundation4ai_document_queue_embedding_latency | Gauge | Total latency of embedding computation | Increases by the seconds spent computing the embeddings of a pipeline group, for every embedding model of the pipeline, when the worker processed the group without a group failure | Seconds spent computing embeddings |
Latency gauges
The three latency metrics are gauges that only increase between restarts, so the Prometheus functions rate and increase apply to the latency metrics as to counters. Monitoring and logging describes the queries.
The following lines show the form of the worker output:
# TYPE foundation4ai_document_queue_processed counter
foundation4ai_document_queue_processed{description="Number of documents processed"} 1200
# TYPE foundation4ai_document_queue_processed_latency gauge
foundation4ai_document_queue_processed_latency{description="Total latency of documents processed"} 482.61
Processing outcomes
Each document of a pipeline group ends in one of the following outcomes. Every outcome increases foundation4ai_document_queue_processed by 1 for the document.
| Outcome | Counter | Worker log line | Job |
|---|---|---|---|
| Fragments stored | success | INFO Successfully processed document <document> for pipeline <pipeline> | Acknowledged |
| Empty document text | skip | WARN Skipped document <document> for pipeline <pipeline>: pending. The document status becomes success. | Acknowledged |
Document no longer pending, such as a job delivered again after processing | skip | WARN Skipped document <document> for pipeline <pipeline>: <status> | Acknowledged |
| Document deleted before processing | success | WARN Document <document> for pipeline <pipeline> was removed before processing | Acknowledged |
| Document failed | failed | WARN Failed processing document <document> for pipeline <pipeline>: <reason> | Returned to the queue and delivered again after 10 seconds |
| Pipeline group failed | failed, for every document of the group | ERROR Error processing documents for pipeline <pipeline>: <error> | Every job of the batch returned to the queue and delivered again after 10 seconds |
A pipeline group fails as a whole when the worker cannot complete a step that covers every document of the group, such as the license check, the pipeline lookup, text splitting, the embedding computation or the database transaction. The INFO line appears only when app.log_level is info, debug or trace.
Labels added at collection
The scrape job kubernetes-pods of the core chart adds the labels job, instance, namespace, pod and node and every pod label, with dots, slashes and hyphens replaced by underscores. The label app_kubernetes_io_name separates the API server (api-server) from the workers (api-server-worker). The worker pods carry fewer pod labels than the API server pod, so app_kubernetes_io_managed_by, app_kubernetes_io_version and helm_sh_chart appear on the API server series only. The job keeps the labels of the exported series when the names collide. Prometheus scrape scope describes the job.
Statistics route
GET /statistics returns document processing and agent execution figures for a period, read from Prometheus, and the number of jobs in the queue. The route requires no API key, answers on the API port and is not used by the dashboard.
The query parameter interval is required and holds the length of the period, which ends at the time of the request. The value is a Prometheus duration: one or more of <n>y, <n>w, <n>d, <n>h, <n>m, <n>s and <n>ms, from the largest unit to the smallest, such as 15m, 1h30m or 24h. The value 0 is accepted as well.
Response fields
The response contains the following fields. In the current release, only document_queue carries a value; the other fields report 0 or an empty list, so a period without activity and missing data look the same.
| Field | Type | Meaning |
|---|---|---|
documents_added | Number | Documents submitted in the period |
documents_processed | Number | Documents processed in the period |
documents_processed_latency | Number | Average processing time in seconds |
documents_failed | Number | Failed processing attempts in the period |
agent_executions | Number | Agent executions in the period |
agent_execution_latency | Number | Average agent execution time in seconds |
agent_execution_latencies | Array of pairs of a bucket bound and a count | Distribution of agent execution times |
document_queue | Integer | Number of messages in the NATS JetStream stream named by nats.stream, DOCUMENTS in the charts: the jobs waiting in the queue and the jobs that a worker is processing |
document_queue_processed_latencies | Array of pairs of a bucket bound and a count | Distribution of processing times |
The following response is the form that the route returns in the current release:
{
"documents_added": 0.0,
"documents_processed": 0.0,
"documents_processed_latency": 0.0,
"documents_failed": 0.0,
"agent_executions": 0.0,
"agent_execution_latency": 0.0,
"agent_execution_latencies": [],
"document_queue": 42,
"document_queue_processed_latencies": []
}
Errors
The route returns the following errors:
| Status | Body | Cause |
|---|---|---|
| 400 | JSON error body with the message Invalid interval format: <value>. Must be a valid Prometheus duration (e.g., 15m, 1h, 24h). | interval is not a Prometheus duration |
| 400 | Plain text | The request has no interval parameter |
| 500 | JSON error body with the message Failed to get stream info: <reason> | The API server could not read the stream state from NATS JetStream |
The queue figures of Queue state and the Prometheus queries of Recommended panels replace the route for monitoring.