Skip to main content

Architecture

Foundation4 consists of a small set of cooperating services backed by a single PostgreSQL database. This page describes each component, the ingestion and retrieval data paths, and the storage location and retention period of each category of data. Integrators can use this information to understand request behavior and latency. Operators can use this information to determine which components must be deployed, scaled, monitored and backed up.

System overview​

The following diagram shows the runtime components of a Foundation4 deployment and the external systems that communicate with those components.

Solid borders indicate Foundation4 components. Dashed borders indicate systems outside Foundation4. Metrics and trace export to Prometheus and OpenTelemetry are omitted from the diagram.

Components​

API server​

The API server is a Rust service that receives all external requests. The API server exposes the following endpoints:

PathPurpose
/REST API. Routes are served from the root path, with no /api or version prefix.
/openapi.jsonOpenAPI description of the REST API
/docsInteractive API documentation
/mcpModel Context Protocol (MCP) endpoint, using streamable HTTP
/metricsPrometheus metrics

The API server executes search requests and agent runs directly. For ingestion requests, the API server validates the request, records the document in PostgreSQL, and places the document text on the work queue. The workers perform all further processing.

Workers​

Workers run the same Rust binary as the API server, started in worker mode. Each worker retrieves ingestion jobs from the queue in batches of up to 100 documents, groups the jobs by pipeline, splits each document into fragments, computes embeddings, and writes the fragments, vectors and full-text index entries to PostgreSQL.

Ingestion is the most compute-intensive workload in Foundation4. The worker deployment therefore scales independently of the API server deployment. The default Helm installation runs three worker replicas.

gRPC service (sidecar)​

The gRPC service is a Python service that hosts the parts of Foundation4 that are implemented in Python: all text splitters and the embedding providers that depend on Python libraries. The API server and the workers are written in Rust. When a request requires a text splitter or one of these embedding providers, the Rust process calls the gRPC service over gRPC.

The gRPC service is deployed as a sidecar container. A sidecar is a second container that runs inside the same Kubernetes pod as the main container and shares the pod's network namespace. Every API server pod and every worker pod contains one gRPC service container, which the Rust process reaches on the pod's loopback interface at port 50051. Because each pod contains a dedicated instance, the capacity of the gRPC service grows with the number of API server and worker replicas, and no separate deployment or service discovery is required.

The following table shows which process executes each component.

ComponentExecuted by
Text splitters (all seven LangChain splitters)gRPC service
GPT4All embedding modelsgRPC service
Hugging Face sentence-transformers embedding modelsgRPC service
Hugging Face inference endpoint embedding modelsgRPC service
FastEmbed embedding modelsAPI server or worker (Rust process)
OpenAI-compatible embedding endpointsAPI server or worker (Rust process)
LLM callsAPI server (Rust process)

The gRPC service exists because the text-splitting library and several embedding libraries that Foundation4 depends on are available only as Python packages. Model weights and Python packages are supplied to the gRPC service container through container image volumes. For details, see Models, packages and air-gapped installs.

PostgreSQL with pgvector​

PostgreSQL is the primary data store for Foundation4. All persistent application state, including configuration objects, documents, fragments, vectors and full-text indexes, is stored in PostgreSQL. The pgvector extension provides vector storage and nearest-neighbor indexing.

Foundation4 uses two schemas:

  • public contains shared tables for pipelines, documents, document versions, embedding models, text splitters, LLMs, agents, taxonomies, API keys and permissions.
  • embeddings contains a dedicated set of tables for each pipeline: encrypted fragment text, classifications, the document hierarchy, the full-text index, and one vector table for each embedding model. Vector tables use HNSW indexes for approximate nearest-neighbor search.

Foundation4 stores vectors in PostgreSQL rather than in a separate vector database. This design has two consequences. A deployment has a single database to secure, back up and restore. A single SQL query can combine vector distance ranking with metadata filters and classification constraints.

NATS JetStream​

NATS JetStream provides the work queue between the API server and the workers. NATS JetStream also provides an object store that holds the raw text of each document until a worker has processed the document.

Redis-compatible cache​

The Redis-compatible cache holds short-lived data only:

  • Encrypted results of API key verification
  • Frequently read objects, such as pipelines and agents
  • Encrypted agent execution traces, retained for 60 minutes

The Helm installation deploys Valkey as the cache. The cache contains no data that is required for recovery.

Dashboard​

The dashboard is a web application served at /dashboard on the same host as the API server. Operators authenticate to the dashboard with an API key and secret.

Observability​

The API server and the workers publish Prometheus metrics. The API server can also export traces and logs through OpenTelemetry; the workers, the gRPC service and the dashboard export no OpenTelemetry data. The API server assigns an x-request-id value to every request to an API operation, returns the value in the response headers and records the value on the OpenTelemetry trace of the request. Observability explains the signals, and Monitoring and logging describes the operator procedures.

Ingestion path​

Ingestion is asynchronous. The API server acknowledges each document immediately, and the workers process the document afterward. This separation prevents large ingestion loads from degrading search latency.

The ingestion path is built on the following design principles, each of which affects how client applications integrate with Foundation4.

  • Completion is determined by polling. The create request returns a document identifier with the status pending. The client application retrieves the document to determine when processing is complete and the document is searchable. Foundation4 does not currently provide webhooks or callbacks.
  • Failed jobs are retried automatically. If a worker cannot process a document, the worker returns the job to the queue, and the queue redelivers the job after 10 seconds. The document remains in the pending state until processing succeeds. Worker logs and metrics record the cause of each failure.
  • Raw text is transient. The NATS object store holds the original text only until a worker processes the document. After processing, Foundation4 retains the fragments, with their text encrypted, together with their vectors and full-text index entries.
  • Input is plain text. A document is submitted as a JSON body with a contents string field. Content in other formats, such as PDF or Microsoft Word, must be converted to text before submission.

Retrieval path​

Retrieval is synchronous. The API server processes each search or agent request to completion within the originating HTTP request.

  • Search. For a request to POST /pipelines/{id}/search, the API server embeds the query for vector search modes, or parses the query into search terms for full-text search. The API server then applies the requested classifications and metadata filters and returns ranked fragments with their scores and metadata.
  • Agent execution. For a request to POST /agents/{id}/execute, the API server runs the searches defined by the agent's placeholders, inserts the results into the prompt template, calls the registered LLM, and streams the response as NDJSON, server-sent events or plain text. When tracing is enabled, the API server retains a record of the prompt and source fragments for 60 minutes. Client applications can use this record to display sources or to diagnose response quality.

The API server is the only component that communicates with an LLM. All outbound model traffic therefore originates from a single component, which simplifies network policy. In an air-gapped deployment, the LLM endpoint is another service on the internal network.

Data storage​

The following table lists each category of data, the storage location, and the retention period.

DataLocationRetention
Pipelines, documents, versions, models, agents, API keys, permissionsPostgreSQL, public schemaUntil deleted
Fragment text (encrypted), vectors, full-text index entries, metadataPostgreSQL, embeddings schema, one table set per pipelineUntil the document is deleted
Raw document text awaiting processingNATS object storeUntil a worker processes the document, for at most 24 hours
API key verification results, cached objectsRedis-compatible cacheMinutes
Agent execution tracesRedis-compatible cache60 minutes
Embedding model weights and Python packagesContainer images mounted into podsFor the life of the release

The complete persistent state of a deployment consists of the PostgreSQL database, the application secret, and the identifier and secret of the master key. NATS JetStream holds only work in progress, and the cache is rebuilt automatically. For backup and recovery procedures, see Backup, restore and upgrades.