Data protection
This page describes how Foundation4 protects stored data and credentials, how data is protected in transit, how long each kind of data is retained, and what the operator of a deployment protects. Security reviewers and architects need this page to assess a deployment, and operators need this page to plan the controls around Foundation4. Security at a glance summarizes the mechanisms, and Secrets and keys and Security hardening in Deploy and operate hold the operator procedures.
Protection overview
Foundation4 encrypts selected data at the application layer, before the data reaches PostgreSQL or the Redis-compatible cache. Storage-level encryption, such as encrypted volumes or an encrypted database service, is provided by the platform and configured by the operator, independently of Foundation4.
| Data | Location | Application-layer protection | Key location |
|---|---|---|---|
| Fragment text | PostgreSQL, embeddings schema | ChaCha20-Poly1305, with one key for each classification | Same database, in the classifications table of the pipeline |
| LLM API keys | PostgreSQL, public schema | Fernet, keyed by the application secret | Kubernetes Secrets and the secrets file |
| API key verification results | Redis-compatible cache | Fernet, keyed by the application secret | Kubernetes Secrets and the secrets file |
| Agent execution traces | Redis-compatible cache | Fernet, keyed by the application secret | Kubernetes Secrets and the secrets file |
| API key secrets | PostgreSQL, public schema | Argon2id hash | No key |
| Classification keys, vectors, full-text index entries, metadata, identifiers and configuration | PostgreSQL | None | Not applicable |
| Frequently read objects, such as pipelines and agents | Redis-compatible cache | None | Not applicable |
The following diagram shows where the encrypted data and the keys are stored.
Each outer box is a storage location, and the dashed box is outside Foundation4. Each arrow points from a key to the data that the key encrypts and names the algorithm. Data without an incoming arrow has no application-layer encryption, except the API key secret hashes, which are one-way hashes.
Fragment text encryption
The worker encrypts the text of every fragment with ChaCha20-Poly1305, an authenticated encryption algorithm, before the worker writes the fragment to PostgreSQL.
- Keys. Each classification of a pipeline has a 256-bit key, generated from the operating system's random number generator when the classification is created, either with the pipeline or by a later change to the pipeline's classifications. Adding a classification that already exists keeps the existing key.
- Nonces. Each fragment is encrypted with a new random 96-bit nonce, which is stored next to the ciphertext.
- Storage. The keys are stored as raw bytes in the
encryption_keycolumn of the pipeline's classifications table,<pipeline id>:classifications, in theembeddingsschema. Foundation4 does not encrypt the keys with the application secret and does not use an external key management service. - Use. The worker encrypts each fragment with the key of the document's classification. The API server reads the key of each fragment's classification in the same database query that reads the fragment, and decrypts the text wherever a request uses the text: search results, fragment lists requested with
contents=true, agent retrieval and searches through the Model Context Protocol (MCP) server. - Integrity. The Poly1305 authentication tag detects altered ciphertext. A fragment that fails the check makes the request return HTTP 500 with the message
Decryption error. - Lifetime. A key exists as long as the classification exists. Removing a classification deletes the key together with every document and fragment that carries the classification, as Classifications describes. Foundation4 provides no operation that replaces the key of an existing classification.
Scope of fragment encryption
The keys and the encrypted text are stored in the same database. The encryption therefore protects fragment text in a copy of the fragment tables that excludes the classifications tables. The encryption does not protect fragment text from a party that can read both, such as a database role with read access to the embeddings schema, or a backup of the whole database.
The keys separate the ciphertext of each classification, but the keys do not decide which fragments a request can read. The API server reads the key of any fragment that a request returns. Which fragments a request returns depends on the permissions of the API key on the pipeline and on the classifications that the search request names, as Access control and Classifications describe.
Data without application-layer encryption
PostgreSQL stores the following data without application-layer encryption, because the database reads the data directly to answer searches and filters, or because the data identifies objects:
- Vectors. Nearest-neighbor search compares vectors inside PostgreSQL. Vectors are computed from the fragment text.
- Full-text index entries. For pipelines with full-text search, the worker sends the fragment text to PostgreSQL, which stores the normalized words of each fragment with their positions. Full-text queries match against these entries.
- Metadata. Filters compare metadata values inside PostgreSQL. Each document and each fragment carries a copy of the document's metadata.
- Identifiers and states. Document identifiers, external identifiers, versions, expiry times, processing statuses and classification names.
- Configuration. The names, descriptions and settings of pipelines, agents with their prompt templates, embedding models, text splitters, taxonomies, API keys and permissions, and LLM registrations apart from the API key.
Database access controls and storage-level encryption protect this data, as described in Operator responsibilities.
Application secret
The application secret is a Fernet key: 32 random bytes in URL-safe base64 encoding, read from the configuration key app.secret, which the installation sets from FOUNDATION4AI_APP_SECRET in the secrets file. Fernet is an authenticated symmetric encryption format. Fernet encrypts with the Advanced Encryption Standard (AES), using a 128-bit key in cipher block chaining (CBC) mode, and signs the result with a hash-based message authentication code (HMAC) that uses SHA-256. A value that is not a valid Fernet key stops the API server and the workers at startup with Error decoding secret key.
| Data | Encrypted when | Effect of a changed or lost application secret |
|---|---|---|
| LLM API keys | A create or update request sets api_key | Stored keys cannot be decrypted, and Foundation4 treats each LLM as having no key. Agent executions and direct queries that use the LLM return HTTP 400 Invalid parameter, with details.error set to API key is required for OpenAI provider, until the key is set again. |
| API key verification results | Each verification that reads the stored hash | Entries that cannot be decrypted are discarded and computed again |
| Agent execution traces | Each execution with tracing set to true | Traces written before the change return HTTP 500 Decryption error until the traces expire |
Fragment text and API key secrets do not depend on the application secret. Foundation4 has no operation that re-encrypts stored data under a new application secret, so a changed application secret has the same effect as a lost one: the API key of each LLM is set again with PATCH /llms/{id}, as LLMs describes. Secrets and keys describes the procedure for the application secret.
LLM API keys in responses
No endpoint returns the API key of an LLM. The following request reads an LLM registration and requires read permission on the LLM. The shell variables are described in Authenticate, and $LLM_ID holds the id of the registration.
curl "$FOUNDATION4_URL/llms/$LLM_ID" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"
The response has status 200 and contains the registration without the API key:
{
"id": "3f0c9b6e-2d4a-4c1e-9b7f-5a8d2e6c1f04",
"name": "internal-general",
"description": "General-purpose model on the internal model server",
"endpoint": "http://llm.internal:8000/v1",
"model": "example-model-8b-instruct"
}
API key secrets
Every authenticated request carries an API key identifier and secret. Foundation4 protects the secrets as follows:
- Generation.
POST /api-keysgenerates each secret from 36 random bytes, encoded in Base64 as 48 characters, and returns the secret once, in the response that creates the key. - Storage. Foundation4 stores an Argon2id hash of the secret, computed with a random salt. No endpoint returns the secret or the hash. A lost secret cannot be recovered, and the key is replaced as described in Access control.
- Verification. When a key is not in the cache, the API server checks the secret in the
x-api-key-secretheader against the stored hash. The API server keeps the result of each successful verification in the Redis-compatible cache, encrypted with the application secret, with an expiry of 15 minutes. - Master key. The identifier and secret of the master key come from the configuration keys
app.master.keyandapp.master.secret. The installation Job stores the Argon2id hash of the master key secret in PostgreSQL, and the API server and the workers verify the master key at startup. Secrets and keys describes where the configured values are stored.
Data in transit
Data travels between Foundation4, client applications, the data stores and model servers over the following connections:
- Client connections. The API server and the dashboard serve plain HTTP inside the cluster, and TLS for client applications and browsers terminates at the ingress controller or the Gateway. Every authenticated request carries an API key secret in the
x-api-key-secretheader, so every client path to Foundation4 uses TLS. Security hardening describes the configuration. - PostgreSQL. The
sslmodeparameter of the database URL sets encryption and certificate checks, anddatabase.ca_certsupplies a private certificate authority (CA) certificate. The connection carries the classification keys with each read of fragment text, the fragment text of pipelines with full-text search and the query text of full-text searches, so a production deployment usesverify-full, as Database describes. - Certificate checks. The default mode,
prefer, encrypts only when the server offers TLS. Neitherprefernorrequireverifies the server certificate, even whendatabase.ca_certis set; onlyverify-caandverify-fulldo. - Connections inside the namespace. The API server and the workers also connect to the Redis-compatible cache, to NATS JetStream, and to the gRPC service, which runs as a sidecar container, a second container in the same Kubernetes pod. The gRPC service receives document text for text splitting and for the embedding models that the gRPC service runs. Security hardening describes these connections and the network policies that restrict them.
- Model servers. Requests to embedding and LLM servers carry fragment text, queries and prompts. A request uses TLS when the endpoint URL starts with
https://, which Security hardening recommends for endpoints on shared networks.
Retention and deletion
The following table lists the retention of each kind of data.
| Data | Location | Retention |
|---|---|---|
| Raw document text and processing jobs | NATS JetStream | Until a worker processes the document, for at most 24 hours |
| Agent execution traces | Redis-compatible cache | 60 minutes, for executions with tracing set to true |
| API key verification results | Redis-compatible cache | Expiry of 15 minutes |
| Frequently read objects | Redis-compatible cache | Minutes |
| Document versions, fragments, vectors, full-text index entries and metadata | PostgreSQL | Until the document is deleted |
| Classification keys | PostgreSQL | Until the classification is removed or the pipeline is deleted |
The following rules apply:
- Raw text. The worker deletes the raw text of a document from the object store after the fragments of the document are stored. Raw text that no worker processes, such as the text of a document that fails processing or that is deleted while
pending, remains in the object store until the 24-hour limit. - Traces. A trace holds the input values, the messages sent to the model and the retrieved fragments with their decrypted text. Only the API key that ran the execution can read the trace, as Agents and prompt templates describes.
- Versions and expiry. An expiry, and a new version that expires the previous versions, mark versions as expired and keep the encrypted text of those versions in the database. Only deletion removes the text, as Documents, versions and fragments describes.
- Deletion. Deleting a document without the
as_ofparameter removes every version of the document, with the fragments, vectors and full-text index entries of each version.
Backups
A backup of the PostgreSQL database holds all stored content and configuration: the encrypted fragment text together with the classification keys, the vectors, full-text index entries, metadata and configuration, the API key secret hashes, and the LLM API keys encrypted with the application secret. A database backup therefore carries the same sensitivity as the database, and the fragment encryption provides no protection inside a backup of the whole database.
The LLM API keys in a backup can be decrypted only with the application secret that was in effect when the backup was taken. The application secret is therefore kept with each backup set but stored separately from the backup files, as Backup, restore and upgrades describes. NATS JetStream and the Redis-compatible cache are not part of the backup set.
Deleting a document changes no existing backup. Content deleted after a backup was taken remains in that backup, and in archived write-ahead log files of the database, until the backup retention of the organization removes those copies. The retention period of backups is therefore part of any deletion obligation, such as an erasure request.
Operator responsibilities
The operator protects the database, the secrets and the infrastructure around Foundation4:
| Asset | Reason | Controls |
|---|---|---|
| PostgreSQL access | The classification keys and the encrypted fragment text are in the same database | Database roles limited to Foundation4 and the database administrators, a dedicated database and network access limited to the Foundation4 pods, as Database and Security hardening describe |
| Database storage and backups | A backup holds all stored content, including the keys | Storage-level encryption of the database service, restricted access to backup files and a backup retention period that matches deletion obligations |
| Application secret and master key | The application secret decrypts LLM API keys, cache entries and traces, the master key administers the deployment, and a restore needs both values of the backup | Kubernetes Secrets encrypted at rest and readable only by operators, with copies in the organization's secret store apart from the backup files, as Security hardening describes |
| NATS JetStream and the cache | NATS JetStream holds submitted document text for up to 24 hours, and the cache holds cached objects and encrypted entries | Network policies that admit only the Foundation4 pods, and storage-level encryption of the NATS JetStream volumes |
| Transport | API key secrets, fragment text and prompts cross the network | TLS at the ingress or Gateway, sslmode=verify-full for PostgreSQL and https:// endpoints for model servers on shared networks |
| Logs and trace exports | Logs and OpenTelemetry traces are operational records of requests | Access to the log store and the trace backend limited to the operators of the installation, as Security hardening describes |
| API keys of client applications | A secret carries every permission of the key | One key for each client application, kept in the secret store of the application |