Scaling and performance
This page describes how each Foundation4 component scales, how the workers process documents, and how resource settings, database connections, autoscaling and the queue retention limit the capacity of a deployment. Operators who size a production deployment or plan a large document load need this page. Values reference lists every value that this page names.
Components and replicas
| Component | Default replicas | Scaled by | Scaling notes |
|---|---|---|---|
| API server | 1 | api-server.replicaCount.server or autoscaling | Stateless for the REST API. MCP sessions exist only on the replica that created them, so MCP clients need a single replica or session affinity. Each replica holds a database pool and loads each model that the replica uses. |
| Worker | 3 | api-server.replicaCount.worker or autoscaling | Each replica processes one batch of jobs at a time with one database connection |
| gRPC service | One per API server pod and one per worker pod | The pod count | Runs as a sidecar container, a second container in each pod, with the resource settings of the pod's main container |
| Dashboard | 1 | dashboard.replicaCount | Serves the dashboard files; carries no processing load |
| NATS JetStream | 3 | nats.config.cluster.replicas | Holds the queue and the submitted document text for 24 hours |
| Redis-compatible cache | 1 | Not scaled by the charts | Holds caches of API keys, pipeline, agent and text splitter records, and execution traces in memory |
| PostgreSQL with pgvector | External in production | Outside the charts | Primary data store; the connection budget limits the Foundation4 replicas |
The API server and the workers scale independently. API server replicas add capacity for requests: search, document submission and agent execution. Worker replicas add capacity for document processing.
API server capacity
Every search in similarity or MMR mode embeds the query text before PostgreSQL runs the vector search. The embedding runs where the embedding model's provider runs:
- FastEmbed. Inside the API server process. Inference runs on the threads that serve requests, and calls to one model are serialized within one process, so concurrent searches that use the same FastEmbed model wait for each other in each replica (inferred from the code).
- gRPC service providers. In the gRPC service container of the same pod.
- OpenAI-compatible embeddings. On the embeddings server, outside the pod.
For FastEmbed models, concurrent search capacity therefore grows with the number of API server replicas (inferred). Agents and direct queries add the latency of the registered model server, which the API server waits for.
Each API server process loads a model on first use and keeps every loaded model in memory until the process ends. Memory therefore grows with the number of distinct embedding models and text splitters that the deployment uses, and the first request after a restart or scale-up is slower while the model loads.
Worker processing
Each worker process runs one loop:
- Fetch. The worker takes up to 100 jobs from NATS JetStream, waiting at most 250 milliseconds for the batch to fill.
- Group. The worker groups the jobs of the batch by pipeline and processes the groups one after another.
- Split. For each document of a group, the worker sends the text to the text splitter in the gRPC service.
- Embed. The worker requests the embeddings of all fragments of the group in one call per embedding model.
- Store. The worker writes the fragments and vectors to PostgreSQL in one transaction per group and acknowledges the jobs.
The batch size of 100 and the sequential processing are fixed in the code; no value changes the batch size or the order. Throughput therefore scales with the number of worker replicas.
The following behaviors reduce throughput when processing fails:
- Failed groups. A splitting or embedding error in one document fails the whole group, and the worker returns every job of the batch to the queue, including jobs of other pipelines. The jobs return after 10 seconds.
- Long batches. The queue redelivers a job that stays unacknowledged longer than the NATS JetStream acknowledgement wait, 30 seconds by default, so a batch that runs longer can be processed twice (inferred).
- Queue errors. A queue read error or an undecodable job ends the worker process, and Kubernetes restarts the container.
- Large embedding requests. Embedding requests to the gRPC service carry the vectors of a whole group, and the gRPC client keeps a 4 MiB message limit, which large groups of high-dimension vectors could exceed (unverified).
Ingest documents reliably describes the retries from the point of view of the client application.
Resource requests and limits
The charts set no resource requests or limits for any container. The following keys set requests and limits:
| Component | Key |
|---|---|
| API server, workers and both gRPC service containers | api-server.resources |
| Dashboard | None in practice, as described in Values reference; a namespace LimitRange applies defaults |
| NATS JetStream | nats.container.resources |
| Redis-compatible cache | redis.resources |
| Prometheus | prometheus.server.resources |
| Bundled PostgreSQL | postgres.resources |
One api-server.resources block applies to four containers: server and grpc in the API server pods, and worker and grpc in the worker pods. The block therefore covers the largest of the four:
- FastEmbed models. Loaded in the
serverandworkercontainers. - GPT4All and Hugging Face models, and every text splitter. Loaded in the
grpccontainers. - Worker batches. The text, fragments and vectors of up to 100 documents, so worker memory depends on document size and embedding dimensions.
Because no reference numbers exist, the operator measures a representative load before setting the values.
Resource settings
-
Run a representative load, such as a first round of documents and the expected search traffic, and read the usage of each container. The command needs the Kubernetes metrics server:
kubectl -n foundation4ai top pods --containersExpected result: CPU and memory usage for the
server,worker,grpcand dashboard containers. -
Set the block from the highest usage of the four containers, with headroom, in the values file of the application release:
api-server:resources:requests:cpu: <cpu request>memory: <memory request>limits:memory: <memory limit>Expected result: the values file holds one
resourcesblock underapi-server. -
Upgrade the application release:
helm upgrade --install foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployed, and the API server and worker pods are replaced. -
Confirm the settings of the worker pod template:
kubectl -n foundation4ai get deploy foundation4ai-api-server-worker \-o jsonpath='{range .spec.template.spec.containers[*]}{.name}{" "}{.resources}{"\n"}{end}'Expected result: the
workerandgrpclines show the same requests and limits. -
After a new load, check for containers stopped for exceeding the memory limit:
kubectl -n foundation4ai get pods \-o jsonpath='{range .items[*]}{.metadata.name}{" "}{.status.containerStatuses[*].lastState.terminated.reason}{"\n"}{end}'Expected result: no line contains
OOMKilled. A line withOOMKilledcalls for a higher memory limit or smaller documents.
Database connections
Each API server replica opens up to database.pool_size connections, 3 in the chart and 5 when unset, and each worker opens 1. The total at the maximum replica counts must stay below the connection limit of PostgreSQL. With autoscaling at the chart maximum of 100 replicas for both Deployments and the chart pool size, the deployment can open 400 connections. Database describes the connection budget and the HTTP 500 responses that pool exhaustion produces.
Autoscaling
The chart creates a HorizontalPodAutoscaler for the API server when api-server.autoscaling.server.enabled is true, and one for the workers when api-server.autoscaling.worker.enabled is true. Each targets 80 percent CPU utilization by default and can add a memory target. While autoscaling is on, the Deployment carries no replica count.
The following conditions apply:
- Requests. Utilization is measured against the requests, so autoscaling needs
api-server.resourcesrequests and the Kubernetes metrics server. Without requests, the autoscaler reports the target as unknown and does not scale. - Whole pod. The autoscaler computes utilization over all containers of the pod, including the gRPC service container.
- Connection budget. The default
maxReplicasof 100 exceeds the connection limit of most PostgreSQL servers, somaxReplicasis set from the budget in Database connections. - Queue depth. The worker autoscaler follows CPU, not the number of queued jobs. The charts include no autoscaling on queue depth.
- Scale-down. The worker process stops on the interrupt signal but not on the termination signal that Kubernetes sends, so a stopping worker pod runs until the termination grace period ends, 30 seconds by default, and the jobs of the interrupted batch return to the queue after the acknowledgement wait (inferred). The jobs are delayed, not lost, while the jobs are younger than 24 hours.
- Model loading. A new pod loads each model on first use, so the first requests on a new replica are slower.
Worker autoscaling
-
Set
api-server.resourcesrequests as described in Resource settings, then enable the worker autoscaler with a maximum that fits the connection budget:api-server:autoscaling:worker:enabled: trueminReplicas: 3maxReplicas: <maximum workers>targetCPUUtilizationPercentage: 80Expected result: the values file holds the
autoscaling.workerblock. -
Upgrade the application release:
helm upgrade --install foundation4ai ./charts/foundation4ai \-n foundation4ai -f foundation4ai.values.yaml --wait --timeout 10mExpected result: Helm reports
STATUS: deployed. -
Check the autoscaler:
kubectl -n foundation4ai get hpa foundation4ai-api-server-workerExpected result: the
TARGETScolumn shows a percentage against80%, such ascpu: 12%/80%, and not<unknown>.
Queue retention and backfills
NATS JetStream keeps each processing job, and the submitted document text, for 24 hours. The queue removes a job that no worker completed within 24 hours, and the document stays pending until the client application submits the document again. The following situations let jobs reach that age:
- Backlog. A backfill submits documents faster than the workers process the documents.
- Repeated failures. Processing fails on every attempt, for example because a model is missing.
- Expired license. The workers stop processing documents while the license is expired, as described in Licensing.
A backfill therefore keeps the backlog at a size that the workers clear well within 24 hours. The client application submits in rounds and reads the pending count between rounds, as described in Pacing and batching. The operator measures the processing rate during the first rounds, as described in Throughput measurement, and divides the planned backlog by that rate.
The API server and the workers create the stream DOCUMENTS and the object store documents with one replica, so the queue lives on one of the 3 NATS servers (inferred from the code). The following command shows the stream configuration and state:
kubectl -n foundation4ai exec deploy/foundation4ai-core-nats-box -- \
nats --server nats://foundation4ai-core-nats:4222 stream info DOCUMENTS
Expected result: the configuration shows work-queue retention, a maximum age of 24 hours and the number of replicas, and the state shows the number of queued messages.
The processes create the stream with a fixed configuration at startup, so a stream edited by hand, such as one with more replicas, can prevent the API server and the workers from starting (inferred).
Throughput measurement
The worker metrics measure processing once Prometheus collects the worker metrics, which needs the worker scrape job of Monitoring and logging. The following queries run in the Prometheus user interface during a load:
| Measure | Query |
|---|---|
| Documents processed per hour | sum(rate(foundation4ai_document_queue_success[15m])) * 3600 |
| Average processing seconds per document | sum(rate(foundation4ai_document_queue_processed_latency[15m])) / sum(rate(foundation4ai_document_queue_processed[15m])) |
| Share of processing time spent embedding | sum(rate(foundation4ai_document_queue_embedding_latency[15m])) / sum(rate(foundation4ai_document_queue_processed_latency[15m])) |
| Share of processing time spent splitting | sum(rate(foundation4ai_document_queue_text_splitter_latency[15m])) / sum(rate(foundation4ai_document_queue_processed_latency[15m])) |
The embedding and splitting shares show which step to address: a different embedding model or provider, a larger chunk size, or more workers. The measured rate for a given embedding model, text splitter and document size is the reference for pacing backfills and for the worker count.
Metadata indexes
A metadata filter needs no index, and the indexes that POST /pipelines/{id}/indexes creates do not accelerate metadata filters, as described in Metadata indexes. Each index adds work to every fragment write of the pipeline, so a deployment creates metadata indexes only for a measured need.