Architecture
Foundation4 consists of a small set of cooperating services backed by a single PostgreSQL database. This page describes each component, the ingestion and retrieval data paths, and the storage location and retention period of each category of data. Integrators can use this information to understand request behavior and latency. Operators can use this information to determine which components must be deployed, scaled, monitored and backed up.
System overview
The following diagram shows the runtime components of a Foundation4 deployment and the external systems that communicate with those components.
Solid borders indicate Foundation4 components. Dashed borders indicate systems outside Foundation4. Metrics and trace export to Prometheus and OpenTelemetry are omitted from the diagram.
Components
API server
The API server is a Rust service that receives all external requests. The API server exposes the following endpoints:
| Path | Purpose |
|---|---|
/ | REST API. Routes are served from the root path, with no /api or version prefix. |
/openapi.json | OpenAPI description of the REST API |
/docs | Interactive API documentation |
/mcp | Model Context Protocol (MCP) endpoint, using streamable HTTP |
/metrics | Prometheus metrics |
The API server executes search requests and agent runs directly. For ingestion requests, the API server validates the request, records the document in PostgreSQL, and places the document text on the work queue. The workers perform all further processing.
Workers
Workers run the same Rust binary as the API server, started in worker mode. Each worker retrieves ingestion jobs from the queue in batches of up to 100 documents, groups the jobs by pipeline, splits each document into fragments, computes embeddings, and writes the fragments, vectors and full-text index entries to PostgreSQL.
Ingestion is the most compute-intensive workload in Foundation4. The worker deployment therefore scales independently of the API server deployment. The default Helm installation runs three worker replicas.
gRPC service (sidecar)
The gRPC service is a Python service that hosts the parts of Foundation4 that are implemented in Python: all text splitters and the embedding providers that depend on Python libraries. The API server and the workers are written in Rust. When a request requires a text splitter or one of these embedding providers, the Rust process calls the gRPC service over gRPC.
The gRPC service is deployed as a sidecar container. A sidecar is a second container that runs inside the same Kubernetes pod as the main container and shares the pod's network namespace. Every API server pod and every worker pod contains one gRPC service container, which the Rust process reaches on the pod's loopback interface at port 50051. Because each pod contains a dedicated instance, the capacity of the gRPC service grows with the number of API server and worker replicas, and no separate deployment or service discovery is required.
The following table shows which process executes each component.
| Component | Executed by |
|---|---|
| Text splitters (all seven LangChain splitters) | gRPC service |
| GPT4All embedding models | gRPC service |
| Hugging Face sentence-transformers embedding models | gRPC service |
| Hugging Face inference endpoint embedding models | gRPC service |
| FastEmbed embedding models | API server or worker (Rust process) |
| OpenAI-compatible embedding endpoints | API server or worker (Rust process) |
| LLM calls | API server (Rust process) |
The gRPC service exists because the text-splitting library and several embedding libraries that Foundation4 depends on are available only as Python packages. Model weights and Python packages are supplied to the gRPC service container through container image volumes. For details, see Models, packages and air-gapped installs.
PostgreSQL with pgvector
PostgreSQL is the primary data store for Foundation4. All persistent application state, including configuration objects, documents, fragments, vectors and full-text indexes, is stored in PostgreSQL. The pgvector extension provides vector storage and nearest-neighbor indexing.
Foundation4 uses two schemas:
publiccontains shared tables for pipelines, documents, document versions, embedding models, text splitters, LLMs, agents, taxonomies, API keys and permissions.embeddingscontains a dedicated set of tables for each pipeline: encrypted fragment text, classifications, the document hierarchy, the full-text index, and one vector table for each embedding model. Vector tables use HNSW indexes for approximate nearest-neighbor search.
Foundation4 stores vectors in PostgreSQL rather than in a separate vector database. This design has two consequences. A deployment has a single database to secure, back up and restore. A single SQL query can combine vector distance ranking with metadata filters and classification constraints.
NATS JetStream
NATS JetStream provides the work queue between the API server and the workers. NATS JetStream also provides an object store that holds the raw text of each document until a worker has processed the document.
Redis-compatible cache
The Redis-compatible cache holds short-lived data only:
- Encrypted results of API key verification
- Frequently read objects, such as pipelines and agents
- Encrypted agent execution traces, retained for 60 minutes
The Helm installation deploys Valkey as the cache. The cache contains no data that is required for recovery.
Dashboard
The dashboard is a web application served at /dashboard on the same host as the API server. Operators authenticate to the dashboard with an API key and secret.
Observability
The API server and the workers publish Prometheus metrics. The API server can also export traces and logs through OpenTelemetry; the workers, the gRPC service and the dashboard export no OpenTelemetry data. The API server assigns an x-request-id value to every request to an API operation, returns the value in the response headers and records the value on the OpenTelemetry trace of the request. Observability explains the signals, and Monitoring and logging describes the operator procedures.
Ingestion path
Ingestion is asynchronous. The API server acknowledges each document immediately, and the workers process the document afterward. This separation prevents large ingestion loads from degrading search latency.
The ingestion path is built on the following design principles, each of which affects how client applications integrate with Foundation4.
- Completion is determined by polling. The create request returns a document identifier with the status
pending. The client application retrieves the document to determine when processing is complete and the document is searchable. Foundation4 does not currently provide webhooks or callbacks. - Failed jobs are retried automatically. If a worker cannot process a document, the worker returns the job to the queue, and the queue redelivers the job after 10 seconds. The document remains in the
pendingstate until processing succeeds. Worker logs and metrics record the cause of each failure. - Raw text is transient. The NATS object store holds the original text only until a worker processes the document. After processing, Foundation4 retains the fragments, with their text encrypted, together with their vectors and full-text index entries.
- Input is plain text. A document is submitted as a JSON body with a
contentsstring field. Content in other formats, such as PDF or Microsoft Word, must be converted to text before submission.
Retrieval path
Retrieval is synchronous. The API server processes each search or agent request to completion within the originating HTTP request.
- Search. For a request to
POST /pipelines/{id}/search, the API server embeds the query for vector search modes, or parses the query into search terms for full-text search. The API server then applies the requested classifications and metadata filters and returns ranked fragments with their scores and metadata. - Agent execution. For a request to
POST /agents/{id}/execute, the API server runs the searches defined by the agent's placeholders, inserts the results into the prompt template, calls the registered LLM, and streams the response as NDJSON, server-sent events or plain text. When tracing is enabled, the API server retains a record of the prompt and source fragments for 60 minutes. Client applications can use this record to display sources or to diagnose response quality.
The API server is the only component that communicates with an LLM. All outbound model traffic therefore originates from a single component, which simplifies network policy. In an air-gapped deployment, the LLM endpoint is another service on the internal network.
Data storage
The following table lists each category of data, the storage location, and the retention period.
| Data | Location | Retention |
|---|---|---|
| Pipelines, documents, versions, models, agents, API keys, permissions | PostgreSQL, public schema | Until deleted |
| Fragment text (encrypted), vectors, full-text index entries, metadata | PostgreSQL, embeddings schema, one table set per pipeline | Until the document is deleted |
| Raw document text awaiting processing | NATS object store | Until a worker processes the document, for at most 24 hours |
| API key verification results, cached objects | Redis-compatible cache | Minutes |
| Agent execution traces | Redis-compatible cache | 60 minutes |
| Embedding model weights and Python packages | Container images mounted into pods | For the life of the release |
The complete persistent state of a deployment consists of the PostgreSQL database, the application secret, and the identifier and secret of the master key. NATS JetStream holds only work in progress, and the cache is rebuilt automatically. For backup and recovery procedures, see Backup, restore and upgrades.