Skip to main content

Scaling and performance

This page describes how each Foundation4 component scales, how the workers process documents, and how resource settings, database connections, autoscaling and the queue retention limit the capacity of a deployment. Operators who size a production deployment or plan a large document load need this page. Values reference lists every value that this page names.

Components and replicas​

ComponentDefault replicasScaled byScaling notes
API server1api-server.replicaCount.server or autoscalingStateless for the REST API. MCP sessions exist only on the replica that created them, so MCP clients need a single replica or session affinity. Each replica holds a database pool and loads each model that the replica uses.
Worker3api-server.replicaCount.worker or autoscalingEach replica processes one batch of jobs at a time with one database connection
gRPC serviceOne per API server pod and one per worker podThe pod countRuns as a sidecar container, a second container in each pod, with the resource settings of the pod's main container
Dashboard1dashboard.replicaCountServes the dashboard files; carries no processing load
NATS JetStream3nats.config.cluster.replicasHolds the queue and the submitted document text for 24 hours
Redis-compatible cache1Not scaled by the chartsHolds caches of API keys, pipeline, agent and text splitter records, and execution traces in memory
PostgreSQL with pgvectorExternal in productionOutside the chartsPrimary data store; the connection budget limits the Foundation4 replicas

The API server and the workers scale independently. API server replicas add capacity for requests: search, document submission and agent execution. Worker replicas add capacity for document processing.

API server capacity​

Every search in similarity or MMR mode embeds the query text before PostgreSQL runs the vector search. The embedding runs where the embedding model's provider runs:

  • FastEmbed. Inside the API server process. Inference runs on the threads that serve requests, and calls to one model are serialized within one process, so concurrent searches that use the same FastEmbed model wait for each other in each replica (inferred from the code).
  • gRPC service providers. In the gRPC service container of the same pod.
  • OpenAI-compatible embeddings. On the embeddings server, outside the pod.

For FastEmbed models, concurrent search capacity therefore grows with the number of API server replicas (inferred). Agents and direct queries add the latency of the registered model server, which the API server waits for.

Each API server process loads a model on first use and keeps every loaded model in memory until the process ends. Memory therefore grows with the number of distinct embedding models and text splitters that the deployment uses, and the first request after a restart or scale-up is slower while the model loads.

Worker processing​

Each worker process runs one loop:

  1. Fetch. The worker takes up to 100 jobs from NATS JetStream, waiting at most 250 milliseconds for the batch to fill.
  2. Group. The worker groups the jobs of the batch by pipeline and processes the groups one after another.
  3. Split. For each document of a group, the worker sends the text to the text splitter in the gRPC service.
  4. Embed. The worker requests the embeddings of all fragments of the group in one call per embedding model.
  5. Store. The worker writes the fragments and vectors to PostgreSQL in one transaction per group and acknowledges the jobs.

The batch size of 100 and the sequential processing are fixed in the code; no value changes the batch size or the order. Throughput therefore scales with the number of worker replicas.

The following behaviors reduce throughput when processing fails:

  • Failed groups. A splitting or embedding error in one document fails the whole group, and the worker returns every job of the batch to the queue, including jobs of other pipelines. The jobs return after 10 seconds.
  • Long batches. The queue redelivers a job that stays unacknowledged longer than the NATS JetStream acknowledgement wait, 30 seconds by default, so a batch that runs longer can be processed twice (inferred).
  • Queue errors. A queue read error or an undecodable job ends the worker process, and Kubernetes restarts the container.
  • Large embedding requests. Embedding requests to the gRPC service carry the vectors of a whole group, and the gRPC client keeps a 4 MiB message limit, which large groups of high-dimension vectors could exceed (unverified).

Ingest documents reliably describes the retries from the point of view of the client application.

Resource requests and limits​

The charts set no resource requests or limits for any container. The following keys set requests and limits:

ComponentKey
API server, workers and both gRPC service containersapi-server.resources
DashboardNone in practice, as described in Values reference; a namespace LimitRange applies defaults
NATS JetStreamnats.container.resources
Redis-compatible cacheredis.resources
Prometheusprometheus.server.resources
Bundled PostgreSQLpostgres.resources

One api-server.resources block applies to four containers: server and grpc in the API server pods, and worker and grpc in the worker pods. The block therefore covers the largest of the four:

  • FastEmbed models. Loaded in the server and worker containers.
  • GPT4All and Hugging Face models, and every text splitter. Loaded in the grpc containers.
  • Worker batches. The text, fragments and vectors of up to 100 documents, so worker memory depends on document size and embedding dimensions.

Because no reference numbers exist, the operator measures a representative load before setting the values.

Resource settings​

  1. Run a representative load, such as a first round of documents and the expected search traffic, and read the usage of each container. The command needs the Kubernetes metrics server:

    kubectl -n foundation4ai top pods --containers

    Expected result: CPU and memory usage for the server, worker, grpc and dashboard containers.

  2. Set the block from the highest usage of the four containers, with headroom, in the values file of the application release:

    api-server:
    resources:
    requests:
    cpu: <cpu request>
    memory: <memory request>
    limits:
    memory: <memory limit>

    Expected result: the values file holds one resources block under api-server.

  3. Upgrade the application release:

    helm upgrade --install foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed, and the API server and worker pods are replaced.

  4. Confirm the settings of the worker pod template:

    kubectl -n foundation4ai get deploy foundation4ai-api-server-worker \
    -o jsonpath='{range .spec.template.spec.containers[*]}{.name}{" "}{.resources}{"\n"}{end}'

    Expected result: the worker and grpc lines show the same requests and limits.

  5. After a new load, check for containers stopped for exceeding the memory limit:

    kubectl -n foundation4ai get pods \
    -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.status.containerStatuses[*].lastState.terminated.reason}{"\n"}{end}'

    Expected result: no line contains OOMKilled. A line with OOMKilled calls for a higher memory limit or smaller documents.

Database connections​

Each API server replica opens up to database.pool_size connections, 3 in the chart and 5 when unset, and each worker opens 1. The total at the maximum replica counts must stay below the connection limit of PostgreSQL. With autoscaling at the chart maximum of 100 replicas for both Deployments and the chart pool size, the deployment can open 400 connections. Database describes the connection budget and the HTTP 500 responses that pool exhaustion produces.

Autoscaling​

The chart creates a HorizontalPodAutoscaler for the API server when api-server.autoscaling.server.enabled is true, and one for the workers when api-server.autoscaling.worker.enabled is true. Each targets 80 percent CPU utilization by default and can add a memory target. While autoscaling is on, the Deployment carries no replica count.

The following conditions apply:

  • Requests. Utilization is measured against the requests, so autoscaling needs api-server.resources requests and the Kubernetes metrics server. Without requests, the autoscaler reports the target as unknown and does not scale.
  • Whole pod. The autoscaler computes utilization over all containers of the pod, including the gRPC service container.
  • Connection budget. The default maxReplicas of 100 exceeds the connection limit of most PostgreSQL servers, so maxReplicas is set from the budget in Database connections.
  • Queue depth. The worker autoscaler follows CPU, not the number of queued jobs. The charts include no autoscaling on queue depth.
  • Scale-down. The worker process stops on the interrupt signal but not on the termination signal that Kubernetes sends, so a stopping worker pod runs until the termination grace period ends, 30 seconds by default, and the jobs of the interrupted batch return to the queue after the acknowledgement wait (inferred). The jobs are delayed, not lost, while the jobs are younger than 24 hours.
  • Model loading. A new pod loads each model on first use, so the first requests on a new replica are slower.

Worker autoscaling​

  1. Set api-server.resources requests as described in Resource settings, then enable the worker autoscaler with a maximum that fits the connection budget:

    api-server:
    autoscaling:
    worker:
    enabled: true
    minReplicas: 3
    maxReplicas: <maximum workers>
    targetCPUUtilizationPercentage: 80

    Expected result: the values file holds the autoscaling.worker block.

  2. Upgrade the application release:

    helm upgrade --install foundation4ai ./charts/foundation4ai \
    -n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10m

    Expected result: Helm reports STATUS: deployed.

  3. Check the autoscaler:

    kubectl -n foundation4ai get hpa foundation4ai-api-server-worker

    Expected result: the TARGETS column shows a percentage against 80%, such as cpu: 12%/80%, and not <unknown>.

Queue retention and backfills​

NATS JetStream keeps each processing job, and the submitted document text, for 24 hours. The queue removes a job that no worker completed within 24 hours, and the document stays pending until the client application submits the document again. The following situations let jobs reach that age:

  • Backlog. A backfill submits documents faster than the workers process the documents.
  • Repeated failures. Processing fails on every attempt, for example because a model is missing.
  • Expired license. The workers stop processing documents while the license is expired, as described in Licensing.

A backfill therefore keeps the backlog at a size that the workers clear well within 24 hours. The client application submits in rounds and reads the pending count between rounds, as described in Pacing and batching. The operator measures the processing rate during the first rounds, as described in Throughput measurement, and divides the planned backlog by that rate.

The API server and the workers create the stream DOCUMENTS and the object store documents with one replica, so the queue lives on one of the 3 NATS servers (inferred from the code). The following command shows the stream configuration and state:

kubectl -n foundation4ai exec deploy/foundation4ai-core-nats-box -- \
nats --server nats://foundation4ai-core-nats:4222 stream info DOCUMENTS

Expected result: the configuration shows work-queue retention, a maximum age of 24 hours and the number of replicas, and the state shows the number of queued messages.

The processes create the stream with a fixed configuration at startup, so a stream edited by hand, such as one with more replicas, can prevent the API server and the workers from starting (inferred).

Throughput measurement​

The worker metrics measure processing once Prometheus collects the worker metrics, which needs the worker scrape job of Monitoring and logging. The following queries run in the Prometheus user interface during a load:

MeasureQuery
Documents processed per hoursum(rate(foundation4ai_document_queue_success[15m])) * 3600
Average processing seconds per documentsum(rate(foundation4ai_document_queue_processed_latency[15m])) / sum(rate(foundation4ai_document_queue_processed[15m]))
Share of processing time spent embeddingsum(rate(foundation4ai_document_queue_embedding_latency[15m])) / sum(rate(foundation4ai_document_queue_processed_latency[15m]))
Share of processing time spent splittingsum(rate(foundation4ai_document_queue_text_splitter_latency[15m])) / sum(rate(foundation4ai_document_queue_processed_latency[15m]))

The embedding and splitting shares show which step to address: a different embedding model or provider, a larger chunk size, or more workers. The measured rate for a given embedding model, text splitter and document size is the reference for pacing backfills and for the worker count.

Metadata indexes​

A metadata filter needs no index, and the indexes that POST /pipelines/{id}/indexes creates do not accelerate metadata filters, as described in Metadata indexes. Each index adds work to every fragment write of the pipeline, so a deployment creates metadata indexes only for a measured need.