Skip to main content

Foundation4 documentation

The self-hosted retrieval platform for organizations that need search and AI answers over sensitive content, on infrastructure they control.

Starting points by role​

  • Integrators connect client applications to Foundation4 through the REST API. Start with Get started, then Concepts and the API reference.
  • Operators install and run Foundation4 on Kubernetes, including in air-gapped networks. Start with Deployment overview and requirements.
  • Administrators configure pipelines, language models, agents and API keys. Start with Concepts, then Guides.
  • Security reviewers assess how Foundation4 handles data and access. Start with Security at a glance.
  • Evaluators need to understand what Foundation4 does and whether Foundation4 fits their requirements. Continue with this page, then Architecture and Feature status.

What Foundation4 is​

Foundation4 is a data pipeline for natural-language content, purpose-built for retrieval-augmented generation (RAG), full-text search and hybrid search.

Client applications send text to Foundation4. Foundation4 splits the text into fragments, computes a vector for each fragment, indexes each fragment for full-text search, and stores the fragments in PostgreSQL with the fragment text encrypted. Client applications then retrieve the most relevant fragments by meaning, by exact terms, or both, and can pass those fragments to a large language model (LLM) to produce an answer grounded in the organization's own content.

A language model without retrieval knows only the content in the model's training data. RAG supplies relevant content at the moment a question is asked, so the answer reflects current documents, messages and policies rather than general knowledge. The quality of the answer depends on retrieval: finding the right passages in a large body of content, quickly, and within the scope that a given reader is permitted to see. Retrieval is the problem that Foundation4 is designed to solve.

Capabilities​

Foundation4 is API-first. Every capability is available through the REST API, and most integrations use nothing else.

  • Ingest. Client applications post plain text to a pipeline. The API server accepts each document immediately, and a pool of workers splits and embeds the document asynchronously, so ingestion load does not degrade search latency.
  • Retrieve. Each pipeline supports three search modes. Similarity search returns the fragments closest in meaning to a query. Maximal marginal relevance (MMR) search balances relevance against variety to avoid near-duplicate results. Full-text search returns fragments that contain the query terms, with the language of each fragment and each query detected automatically. Every search can be scoped by classification and filtered by metadata.
  • Generate. Agents combine retrieval with a language model. An agent is a prompt template with named placeholders. At run time, the API server fills the placeholders with search results, sends the prompt to a registered LLM, and streams the response to the client application.

A Model Context Protocol (MCP) endpoint allows AI assistants that support MCP to search pipelines directly. An administrative dashboard gives operators a web view of pipelines, documents and models.

Deployment​

Foundation4 runs entirely on customer infrastructure. Foundation4 installs on Kubernetes with Helm and stores all state in the organization's own PostgreSQL database. Foundation4 does not depend on a vendor cloud service and does not require internet access.

  • Air-gapped operation. Embedding models and their runtime packages ship as container images. Once those images are mirrored into a disconnected environment, the full stack runs without external connectivity, provided that the pipelines use embedding models and text splitters that work offline. The built-in FastEmbed models that are not in the API server image and the token text splitter download files from the internet on first use. Models, packages and air-gapped installs lists what works without internet access.
  • Model choice. Foundation4 works with any OpenAI-compatible language model that the deployment can reach: a commercial API, a model served in the organization's own data center, or a model on a single host in an isolated network.
  • Data residency. Content, embeddings, keys and logs remain in systems that the organization operates. Outbound connections go to the model endpoints that an administrator registers and, for the embedding models and text splitters that download files on first use, to the sources of those files.

Use cases​

  • Search across chat history. A chat platform sends each message to Foundation4 as the message is posted. When an end user searches, the platform runs a vector search and a full-text search limited to that user's rooms and merges the results, so matches reflect both meaning and exact wording.
  • Knowledge base assistants. An organization loads reference material into a pipeline. End users choose an assistant, a language model and a knowledge base, then ask questions in plain language. Foundation4 retrieves the relevant passages and the language model answers from those passages.

The Solutions section documents both use cases in detail.

Core concepts​

  • Pipeline. A container for one body of content and the rules for processing that content: the embedding model, the text splitter, the available classifications, the permitted metadata, and whether full-text search is enabled.
  • Document and fragment. A document is one piece of text sent to a pipeline. Foundation4 retains each version of a document and divides the document into fragments, the units that search returns.
  • Classification. An attribute that a reader must hold to see a document, following the attribute-based access control (ABAC) model. Each document carries exactly one classification, and each search request names the classifications that the caller holds. The organization defines what classifications represent, such as clearance levels, teams or tenants. Classifications can form a hierarchy, and each classification has a dedicated encryption key.
  • Metadata. Structured data attached to a document, such as a room, an author or a date. Search requests can filter on metadata.
  • Agent. A reusable prompt that brings search results into a language model call.
  • API key. The credential that identifies a caller and carries the caller's read, write and execute permissions.

The Concepts map shows how these objects relate.

Scope​

Foundation4 is the retrieval layer of an AI application. Several adjacent functions are outside the scope of Foundation4 by design.

  • Document parsing. Foundation4 accepts plain text. Content in PDF, Microsoft Word, HTML or other formats must be converted to text before ingestion.
  • Language models. Foundation4 does not host or train models. Foundation4 calls models that the organization provides.
  • User experience. User interfaces, conversation history and presentation belong to the client application.