Skip to main content

Documents, versions and fragments

A document is the unit of content that a client application submits to a pipeline. Foundation4 splits each document into fragments, which are the unit that search returns, and records each update to a document as a new version. This page describes the document lifecycle, versioning, expiry and deletion, point-in-time reads, and the structure of a fragment. Integrators need this page to design ingestion and to keep a pipeline synchronized with a source system.

Document content​

A document consists of plain text, one classification, and optional metadata. Foundation4 does not parse files: a client application that ingests PDF, Office or HTML files extracts the text before submission.

A document is submitted with POST /pipelines/{id}/documents, or with POST /documents and a pipeline_id field in the body. The request body contains the following fields.

FieldRequiredDescription
contentsYesThe plain text of the document
classificationYesOne classification defined in the pipeline
external_identifierNoThe identifier of the content in the client application's own system. Used to submit new versions of the same document.
metadataNoA JSON object. Validated against the pipeline's metadata schema.
text_splitter_idNoA text splitter to use instead of the pipeline's default text splitter
expire_older_versionsNoWhether the new version expires the previous versions. The default is true.

The API server validates the classification and the metadata before accepting the document. A classification that the pipeline does not define, or metadata that does not satisfy the schema, returns an error and creates no document.

Processing lifecycle​

Document processing is asynchronous. The create request returns HTTP 201 with the document identifier, the version and the status pending before any splitting or embedding takes place.

Each rectangle is a value of the document's status field. The circle marks the create request.

  1. Acceptance. The API server records the document with the status pending, stores the raw text in the NATS JetStream object store, and places a processing job on the queue. If the job cannot be placed on the queue, the document status becomes failed.
  2. Processing. A worker reads jobs in batches of up to 100 documents. For each document, the worker splits the text with the document's text splitter, computes a vector for each fragment when the pipeline has an embedding model, builds the full-text index entries when full-text search is enabled, encrypts the fragment text, and writes the fragments to PostgreSQL.
  3. Completion. The worker sets the status to success and deletes the raw text from the object store. A document with empty contents completes with the status success and no fragments.
  4. Retry. If processing fails, the document remains pending and the queue redelivers the job after 10 seconds.

The client application determines when processing is complete by retrieving the document with GET /pipelines/{id}/documents/{document_id} and reading the status field. A document is searchable only after the status is success.

Raw text retention​

Foundation4 retains the raw text of a document only until processing completes, and for at most 24 hours. After processing, the only stored copy of the text is the set of encrypted fragments. Foundation4 therefore cannot return a submitted document verbatim. Fragments can be retrieved in order, but text splitters that overlap adjacent fragments repeat text at each boundary. The client application's source system remains the authoritative copy of the original content.

Versions​

Every submission creates a version. The version identifier is the creation time of the version, expressed in microseconds since the Unix epoch, and is returned in the version field.

The external_identifier field determines whether a submission creates a new document or a new version of an existing document:

  • Without an external identifier. Each submission creates a new document with a new document identifier.
  • With a new external identifier. The submission creates a new document.
  • With an existing external identifier. The submission creates a new version of the document that has that external identifier in the same pipeline. The new version keeps the document identifier of the existing document.

A new version must carry the same classification as the existing document. A submission with a different classification returns an error, even when every earlier version has expired. Content that moves to a different classification requires deleting the document and submitting a new document.

When expire_older_versions is true, the new version replaces the previous version in search results. The previous version's fragments remain searchable until the new version is processed, so the document does not disappear from search during processing. When expire_older_versions is false, the previous versions remain current, and search returns fragments from every current version.

GET /documents/{document_id}/versions lists the current versions of a document, newest first. With include_expired=true, the list also includes the expired versions, each with its expiry time in expired_at. The list is empty for a document identifier that does not exist. GET /documents/{document_id} returns the latest version, or the version named in the version query parameter. This route returns the latest version even after the document has expired, with expired_at set, while GET /pipelines/{id}/documents/{document_id} returns 404 for an expired document unless the request sets include_expired=true.

Expiry and deletion​

Foundation4 provides two ways to remove content from search.

OperationRequestEffect
ExpireDELETE with expire=trueMarks every current version as expired. Search and document lists exclude expired versions. Version history remains readable.
DeleteDELETE without expireRemoves every version and every fragment of the document permanently

Both operations apply to a single document (/pipelines/{id}/documents/{document_id} or /documents/{document_id}) or to every document in a pipeline (/pipelines/{id}/documents). The optional as_of query parameter limits either operation to versions created at or before the given time.

Expiry suits content that has been withdrawn from use but must remain available for audit. Deletion suits content that must no longer exist in the deployment, such as content subject to a retention limit or an erasure request.

Point-in-time reads​

Document and fragment reads accept an as_of query parameter, expressed in microseconds since the Unix epoch. With as_of, Foundation4 returns each document as the document existed at that time: the latest version created at or before that time that had not yet expired. Point-in-time reads apply to the following endpoints:

  • GET /pipelines/{id}/documents
  • GET /pipelines/{id}/documents/{document_id}
  • GET /pipelines/{id}/documents/{document_id}/fragments
  • GET /documents/{document_id}/versions

A version returned for an earlier time shows its expiry time as it is now, so expired_at can be later than as_of. The fragment list of a document follows a different rule: it returns the fragments of the latest version created at or before as_of, or of the latest version when the request has no as_of, even when that version has expired. The expired_at field of each fragment shows the expiry.

Search always operates on current content. Point-in-time reads support audit, reproduction of an earlier answer, and comparison between versions.

Fragments​

A fragment is a section of a document produced by the text splitter. Search returns fragments, and each fragment carries the information that the client application needs to display, cite or re-check a result.

FieldDescription
idThe fragment identifier
document_idThe identifier of the source document
versionThe version of the source document, in microseconds since the Unix epoch
order_idThe position of the fragment within the document version, starting at 0
classificationThe classification of the source document
metadataA copy of the source document's metadata
page_contentThe decrypted fragment text, when the request includes it
expired_atThe expiry time, or null for a current fragment
scoreThe score for the query in search results, and null in fragment lists

Every fragment inherits the classification and metadata of the document, so classification scoping and metadata filters apply to fragments without additional configuration. The Search and retrieval page describes the meaning of score in each search mode.

GET /pipelines/{id}/documents/{document_id}/fragments lists the fragments of a document version in order. The fragment text is included when the request sets contents=true. GET /pipelines/{id}/fragments retrieves specific fragments by identifier, which allows a client application to resolve fragment identifiers stored with an earlier answer. The response is a JSON array that does not follow the order of the requested identifiers. It includes fragments of expired versions, with expired_at set, and omits identifiers whose fragments have been deleted. Fragments retrieved by identifier do not include the text, so a client application that needs the text lists the fragments of the document version with contents=true.