Ingest documents reliably
This guide describes how a client application submits documents to a pipeline, confirms that processing completed, finds documents that did not complete and submits those documents again. The guide also covers retries, idempotent updates, pacing, request size, text preparation, metadata design and license limits. Integrators need this guide to build an ingestion service, and operators need the troubleshooting entry to diagnose documents that remain unprocessed. Documents, versions and fragments describes the document model.
Requirements
The examples use the support-kb pipeline from First search, with the classifications public and internal and a metadata schema that declares the string field product. $PIPELINE_ID holds the pipeline identifier. The shell variables for the base URL and the credentials are described in Authenticate.
The API key of the ingestion service holds read and execute permission on the pipeline, execute permission on the pipeline's text splitter and a classification allow-list that includes the classification of each document. A service that also expires and deletes documents holds permission 7 on the pipeline. Permissions by operation lists the permissions of each document operation.
Document status
Document processing is asynchronous: the create request places a processing job in NATS JetStream, the queue between the API server and the workers, and returns before a worker splits and embeds the document. The status field of a document tells the client application what to do next.
| Status | Meaning | Client application action |
|---|---|---|
pending | The document is queued, or processing failed and the worker retries the document | Check the document again later. Submit the document again when the document is still pending 24 hours after creation. |
success | The fragments are stored and search returns the document | None |
failed | Foundation4 recorded the document but could not place the processing job on the queue | Submit the document again |
Processing lifecycle describes each stage of processing.
Retries and redelivery
The workers and the queue retry processing without action by the client application:
- Processing failure. When processing of a document fails, the worker returns the job to the queue, and the queue redelivers the job after 10 seconds. The queue sets no limit on the number of attempts.
- Queue age limit. The queue removes a job that is still waiting 24 hours after submission, and the object store removes the document's raw text after the same period. A document whose job is removed remains
pendingand is not processed again. - Worker stop. A job that a worker has taken but not acknowledged, for example because the worker pod stopped, returns to the queue after the acknowledgment wait of 30 seconds and is delivered to another worker.
- Repeated delivery. A job for a document that is no longer
pendingis acknowledged without processing, so a repeated delivery does not duplicate fragments. A job for a document that was deleted before processing is acknowledged without processing as well.
A client application therefore submits a document again in two cases only: the document status is failed, or the document is still pending 24 hours after creation. A document that is pending for a shorter time is still in the queue, and the worker processes the document once the cause of the failure is removed.
Document submission
Submit a document with POST /pipelines/{id}/documents and an external_identifier, the identifier of the content in the source system. The identifier makes the submission safe to repeat, as described in Idempotent submission.
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"classification": "public",
"external_identifier": "kb-account-lockout",
"metadata": {"product": "accounts"},
"contents": "Account lockout. After five failed sign-in attempts, an account is locked for 15 minutes. An administrator can unlock the account from the Users page."
}'
Expected result: status 201 and the new document with status set to pending:
{
"id": "01a0f18a-ca80-754e-ba2b-46e126f7cf9e",
"created_at": "2026-09-30T09:00:00.722235Z",
"updated_at": "2026-09-30T09:00:00.722235Z",
"expired_at": null,
"external_identifier": "kb-account-lockout",
"metadata": {"product": "accounts"},
"classification": "public",
"pipeline_id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"text_splitter_id": null,
"status": "pending",
"message": null,
"version": 1790758800722235
}
A status of failed in this response means that the processing job was not queued; submit the document again. Record the identifier and the version of the document:
export DOCUMENT_ID=<id from the response>
export VERSION=<version from the response>
A client application that writes to several pipelines can use POST /documents instead, with the pipeline identifier in the body. The request returns the same response:
curl -X POST "$FOUNDATION4_URL/documents" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"pipeline_id": "'"$PIPELINE_ID"'",
"classification": "public",
"external_identifier": "kb-account-lockout",
"metadata": {"product": "accounts"},
"contents": "Account lockout. After five failed sign-in attempts, an account is locked for 15 minutes. An administrator can unlock the account from the Users page."
}'
Processing check
Retrieve the submitted version with GET /documents/{document_id} and read the status field:
curl "$FOUNDATION4_URL/documents/$DOCUMENT_ID?version=$VERSION" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"
Expected result: status 200 and the document version. After processing, status is success:
{
"id": "01a0f18a-ca80-754e-ba2b-46e126f7cf9e",
"created_at": "2026-09-30T09:00:00.722235Z",
"updated_at": "2026-09-30T09:00:01.016043Z",
"expired_at": null,
"external_identifier": "kb-account-lockout",
"metadata": {"product": "accounts"},
"classification": "public",
"pipeline_id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"text_splitter_id": null,
"status": "success",
"message": null,
"version": 1790758800722235
}
The version parameter makes the check refer to the submitted version, even after a later submission creates a newer version or expires this version. Without the parameter, the request returns the latest version.
For a failing document, a check more often than every 10 seconds returns no new information, because the worker retries a failed document after 10 seconds. A client application that submits many documents checks the number of pending documents instead of each document, as described in Initial load.
Pending documents
The document list, GET /pipelines/{id}/documents, filters by status and by creation time. The following request lists the documents that are still pending more than 24 hours after creation, oldest first, for a check made at 2026-09-30T09:15:00Z. The $ of each operator suffix is escaped for the shell, as described in Filters and ordering.
curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?status=pending&created_at\$lt=2026-09-29T09:15:00Z&count=true&first=100&order_by=created_at&order_by=id" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"
Expected result: status 200 and one page of pending documents. page_info.count holds the total number of matching documents, and query_info repeats the filters that the API server applied:
{
"data": [
{
"id": "01a0e8e3-3700-7bb0-a499-3343833c03f8",
"created_at": "2026-09-28T16:40:00.945114Z",
"updated_at": "2026-09-28T16:40:00.945114Z",
"expired_at": null,
"external_identifier": "kb-sso-setup",
"metadata": {"product": "accounts"},
"classification": "internal",
"pipeline_id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"text_splitter_id": null,
"status": "pending",
"message": null,
"version": 1790613600945114
}
],
"page_info": {
"after": "WyIyMDI2LTA5LTI4VDE2OjQwOjAwLjk0NTExNFoiLCIwMWEwZThlMy0zNzAwLTdiYjAtYTQ5OS0zMzQzODMzYzAzZjgiXQ",
"before": "WyIyMDI2LTA5LTI4VDE2OjQwOjAwLjk0NTExNFoiLCIwMWEwZThlMy0zNzAwLTdiYjAtYTQ5OS0zMzQzODMzYzAzZjgiXQ",
"count": 1,
"has_next": false,
"has_prev": false,
"order_by": ["created_at", "id"]
},
"query_info": {
"created_at$lt": "2026-09-29T09:15:00Z",
"status": "pending"
}
}
Ordering by created_at and then id keeps cursor paging stable, as described in Pagination. When has_next is true, the next page is requested with after set to page_info.after. The list excludes expired versions, so a document that was submitted again leaves this list. Without the created_at filter, the same request returns every pending document, including documents that the workers have not yet processed.
Record the identifier of the pending document:
export STUCK_DOCUMENT_ID=<id of the pending document>
Resubmission
Submit the document again with the same external_identifier, the same classification and the text from the source system:
curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"classification": "internal",
"external_identifier": "kb-sso-setup",
"metadata": {"product": "accounts"},
"contents": "Single sign-on setup. Administrators configure single sign-on on the Security page by uploading the identity provider metadata file and mapping the email attribute."
}'
Expected result: status 201 and a new version of the same document. The id is unchanged, and version identifies the new version:
{
"id": "01a0e8e3-3700-7bb0-a499-3343833c03f8",
"created_at": "2026-09-30T09:20:00.606425Z",
"updated_at": "2026-09-30T09:20:00.606425Z",
"expired_at": null,
"external_identifier": "kb-sso-setup",
"metadata": {"product": "accounts"},
"classification": "internal",
"pipeline_id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"text_splitter_id": null,
"status": "pending",
"message": null,
"version": 1790760000606425
}
The new version expires the version that remained pending. The version history, GET /documents/{document_id}/versions, shows both versions, newest first, when the request includes expired versions:
curl "$FOUNDATION4_URL/documents/$STUCK_DOCUMENT_ID/versions?include_expired=true" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"
Expected result: status 200 and a JSON array. After processing, the new version has the status success, and the earlier version carries an expired_at time:
[
{
"id": "01a0e8e3-3700-7bb0-a499-3343833c03f8",
"created_at": "2026-09-30T09:20:00.606425Z",
"updated_at": "2026-09-30T09:20:00.786112Z",
"expired_at": null,
"external_identifier": "kb-sso-setup",
"metadata": {"product": "accounts"},
"classification": "internal",
"pipeline_id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"text_splitter_id": null,
"status": "success",
"message": null,
"version": 1790760000606425
},
{
"id": "01a0e8e3-3700-7bb0-a499-3343833c03f8",
"created_at": "2026-09-28T16:40:00.945114Z",
"updated_at": "2026-09-28T16:40:00.945114Z",
"expired_at": "2026-09-30T09:20:00.606425Z",
"external_identifier": "kb-sso-setup",
"metadata": {"product": "accounts"},
"classification": "internal",
"pipeline_id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"text_splitter_id": null,
"status": "pending",
"message": null,
"version": 1790613600945114
}
]
The job of the earlier version stays in the queue: the worker still processes it and, while it fails, retries it every 10 seconds until the queue removes it. When the earlier version itself causes the failure, for example because it names a text splitter that cannot run, delete that version once the new version is processed, with as_of set to the version of the earlier version. The as_of parameter limits the deletion to the versions created at or before that time, so the new version remains:
curl -X DELETE "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents/$STUCK_DOCUMENT_ID?as_of=<earlier version>" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"
Expected result: status 204. The worker then acknowledges the earlier job without processing it, and the other documents that shared a worker batch with it are processed.
A pending document without an external identifier cannot receive a new version. The client application deletes such a document with DELETE /pipelines/{id}/documents/{document_id} and submits the text as a new document, as described in Expiry and deletion.
Idempotent submission
The following rules make every create request safe to repeat:
- External identifier. Each document carries the source system's identifier in
external_identifier. A repeated submission with the same identifier in the same pipeline creates a new version of the same document instead of a second document, so a request whose response was lost can be sent again. - Older versions.
expire_older_versionskeeps the defaulttrue, so the latest version replaces the previous versions in search. Versions describes both settings. - Classification. A new version carries the classification of the existing document. A different classification returns HTTP 400 with the message
Classification mismatch. - Order of versions. The client application sends the versions of one document one at a time and waits for each response. When two requests for the same identifier arrive together, the API server records one after the other, and the version recorded last becomes the current version, whatever the order in which the client application sent the requests.
- Documents without an identifier. Each submission creates a new document. A client application that omits
external_identifierrecords each returnedid, and deletes the duplicate when a repeated request creates a second document.
Pacing and batching
Foundation4 has no bulk create request: each document is one create request. Each create request places one job in the queue, and each worker takes up to 100 jobs at a time. A deployment runs 3 workers by default.
- Rounds. The client application submits documents in rounds, such as 100 documents per round, and reads the pending count before the next round.
- Backlog limit. The client application keeps the pending count at a level that the workers clear well within 24 hours, because the queue removes jobs older than 24 hours. The decrease of the pending count per minute during the first rounds measures the throughput of the deployment.
- Parallel requests. Requests for different documents are independent and can run in parallel. Requests for the same external identifier follow the order rule of Idempotent submission.
- Worker capacity. Throughput depends on the text splitter, the embedding model and the number of workers. The operator scales the workers as described in Scaling and performance.
The Processing metrics of the API server and the workers show submitted, processed and failed documents during a load.
Request size
The API server accepts request bodies of up to 2 megabytes (2,097,152 bytes). The limit applies to the whole JSON body, including the metadata and the escape characters that JSON adds to quotes, backslashes and line breaks. The limit counts bytes, not characters: accented letters and characters of non-Latin scripts take two to four bytes each. A larger request returns HTTP 413 with the plain-text body Failed to buffer the request body: length limit exceeded, as listed in Responses without the error body, and creates no document.
A source document larger than the limit is divided into several documents. Each part carries an external identifier derived from the source identifier, such as kb-install-guide-part-1, and the metadata of the source document. When a revised source document produces fewer parts than before, the client application deletes the parts that no longer exist.
Text preparation
A document consists of plain text, as described in Document content. The client application prepares the text before submission:
- Extraction. The client application extracts the text of PDF, Office and HTML files. Foundation4 does not parse files.
- Boilerplate. The client application removes navigation, page headers, page footers and repeated notices, because such text produces fragments that match many queries.
- Titles. The document text begins with the title of the source, and the text of a long document keeps the section headings. A fragment carries only the text and the document metadata, so the title identifies the source in search results and in LLM prompts. Content preparation applies these rules to a knowledge base.
- Empty text. A document with empty contents completes with the status
successand no fragments. The client application skips sources whose extracted text is empty, such as scanned PDF files without a text layer.
Metadata design
The following rules apply to document metadata:
- Schema. The API server validates
metadataagainst the pipeline's metadata schema. A document that does not satisfy the schema returns HTTP 400 with the messageInvalid metadata, anddetailsdescribes each failure under the name of the failing field. A required field that is missing is reported under an empty key, for example{"": "\"region\" is a required property"}. Metadata schema describes the schema rules. - Top-level fields. Values that search requests filter on are top-level fields of
metadata, because filters reference top-level fields, as described in Filters. - Source reference. The metadata carries the values that a client application displays with a result, such as the title and the address of the source.
- Stable values. Every fragment carries a copy of the document metadata, and a change of metadata requires a new version. Values that change often, such as view counts or workflow states, remain in the source system.
License limits
Document creation requires a valid license. The create request returns HTTP 403 with the message Invalid license: Expired after the license expires, and with Invalid license: Exceeded("Documents") or Invalid license: Exceeded("Fragments") when the number of documents or fragments exceeds the license limit. HTTP 403 lists the license messages.
The API server refreshes the document and fragment counts every 3 hours. A limit therefore takes effect at the first refresh after the count passes the limit, and the refusal ends at the first refresh after the count falls within the limit.
Workers process documents only while the license has not expired. Documents that were not processed before the expiry remain pending. After the renewal, the client application submits again the documents that remain pending more than 24 hours after creation.
Error handling
The status code determines whether a failed create request is repeated:
- HTTP 400, 404, 413 and 422. The request is invalid, for example because of metadata outside the schema, a classification mismatch, an unknown pipeline, an unknown or excluded classification, a body over the size limit or a missing field. The client application corrects the request, because the same request fails again.
- HTTP 401 and 403. The key, the key's permissions or the license refuse the request. The client application stops submitting and reports the error, because repetition changes nothing until an administrator acts.
- HTTP 500. A component of the deployment, such as the database or the queue, failed. The client application repeats the request with an increasing delay; with an external identifier, the repetition creates at most a new version.
- Timeouts and connection errors. The document may or may not exist. The client application repeats the request with the same external identifier.
Errors lists every message with the cause.
Troubleshooting
Documents remaining in pending status
-
Symptoms. The pending count does not fall, or documents remain
pendinglong after submission. -
Diagnosis. The request in Pending documents lists the documents that are
pendingmore than 24 hours after creation. The worker log names the cause of each processing failure:kubectl logs -n foundation4ai -l app.kubernetes.io/name=api-server-worker -c worker --tail=500 --prefix \| grep -E "Error processing documents|Failed processing document"A line with
Failed processing documentnames one document, the pipeline and the reason. A line withError processing documents for pipelinenames a failure of a whole pipeline group, such as an error of the gRPC service, of the embedding model orExpired license. When the log shows neither line,kubectl get pods -n foundation4aishows whether the worker pods are running. -
Cause. Processing fails for the document, the license has expired or no worker is running. The worker retries a failing document every 10 seconds, and the queue removes the job of a document that remains
pendingfor 24 hours, so that document is not processed again. -
Resolution. Remove the cause shown in the log: restore the gRPC service or the embedding model, install a valid license, or start the workers. When one document causes the failure, delete that document, or that version as described in Resubmission; the other documents of the pipeline group are then processed. Documents that are
pendingfor less than 24 hours are then processed without action. Submit again, as described in Resubmission, each document that ispendingmore than 24 hours after creation, then confirm that the pending count falls. -
Actions to avoid. Submitting documents again while the cause persists, which adds versions and jobs that fail in the same way. Submitting again with
expire_older_versionsset tofalse, which keeps the pending version current. Purging or recreating theDOCUMENTSstream in NATS JetStream, which removes the jobs of every queued document and leaves those documentspending.