Skip to main content

Rocket.Chat message indexing

The Rocket.Chat indexing service keeps a Foundation4 pipeline current with the messages of a Rocket.Chat workspace, so that AI search can find the messages. This page describes what Foundation4 requires of the indexing service: the document for each message, the metadata that search depends on, edits and deletions, initial loading of existing messages, and the API key. Operators who run the indexing service and integrators who build a comparable service need this page. Rocket.Chat's documentation describes how to install and configure the indexing service. Rocket.Chat message search describes the pipeline and the search side.

Message documents​

The indexing service sends each message as one document. One document per message maps each search result to exactly one message. A long message is split into several fragments, and Rocket.Chat keeps the best-ranked fragment of each message.

Each document carries the following fields, described in Documents, versions and fragments:

FieldValue
contentsThe text of the message, as plain text
classificationThe classification of the message, described in Message classification
external_identifierThe message identifier in Rocket.Chat, which makes later edits new versions of the same document
metadataThe room, message identifier, sender and time of the message, described in Message metadata

The following request adds one message:

curl -X POST "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"classification": "user",
"external_identifier": "Xk3pL9qR2sT7vW1yZ",
"metadata": {
"room_id": "GENERAL",
"msg_id": "Xk3pL9qR2sT7vW1yZ",
"username": "alice",
"timestamp": "2026-09-12T10:00:00.000Z"
},
"contents": "The deploy to staging failed again at the database migration step."
}'

Expected result: status 201 and the new document with status set to pending. The document becomes searchable when the worker has processed the document and status is success.

Foundation4 ignores fields that the request does not define, so a misspelled field name is not reported. The indexing service checks the stored document in the response during setup.

Message metadata​

Rocket.Chat filters and resolves search results by four metadata fields. The pipeline schema declares the four fields as strings, as described in Rocket.Chat message search.

FieldContentRequired
room_idThe identifier of the room that holds the messageYes
msg_idThe identifier of the messageYes
usernameThe username of the sender, without @For the from: search operator
timestampThe time the message was sent, in the format described belowFor the after: and before: search operators and for Recency boost

The following rules apply to the metadata values:

  • Timestamp format. Foundation4 compares string values as text, so date filters are chronological only when every value has the same format. Rocket.Chat writes the dates of a search filter in Coordinated Universal Time (UTC) with milliseconds and the suffix Z, such as 2026-09-12T10:00:00.000Z. Every stored timestamp uses exactly this format.
  • Reserved names. Metadata does not use the field names similarity, score or distance. Rocket.Chat reads fields with those names as the search score of a result.
  • Top-level fields. Filters apply to top-level metadata fields only. Nested objects in metadata are stored, and cannot be filtered.

Message classification​

Each document carries exactly one classification, which must be defined in the pipeline. A reader sees the message in search results only when the classifications in the reader's search request include the message's classification. Rocket.Chat names user in every search request, as described in Classifications of readers, so a message with the classification user can be found by every reader who belongs to the message's room.

The classification of a message cannot change between versions. A new version with a different classification returns HTTP 400 with the message Classification mismatch. A message whose classification must change is deleted, as described in Deleted messages, and then added again with the new classification.

Edited messages​

The indexing service sends the edited text as a new document with the same external_identifier, the same classification and the current metadata. Foundation4 stores the edited text as a new version of the same document. expire_older_versions defaults to true, so the new version replaces the previous version in search results once the new version is processed. Documents, versions and fragments describes versions.

Deleted messages​

A message deleted in Rocket.Chat is never displayed in search results, because Rocket.Chat reads each result from the Rocket.Chat database. The message's fragments remain in the pipeline, occupy result positions in every search that matches the message, and keep the message text stored in Foundation4 until the document is deleted. The indexing service therefore deletes the document of each deleted message.

The delete request takes the Foundation4 document identifier. The indexing service finds the identifier with the external_identifier filter of the document list:

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?external_identifier=Xk3pL9qR2sT7vW1yZ" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and a list with one document in data. The id of the document is the value of $DOCUMENT_ID in the next request:

curl -X DELETE "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents/$DOCUMENT_ID" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 204. Foundation4 removes every version and every fragment of the document. A request with ?expire=true instead marks the document as expired and keeps the version history, which suits content that must remain auditable.

The following changes need no action in Foundation4:

  • Room membership. Rocket.Chat filters each search by the rooms the reader belongs to at the time of the search.
  • Global roles. Rocket.Chat names the reader's current global roles in each search request.

A deleted room is removed by deleting the documents of the room's messages. A complete re-index starts with DELETE /pipelines/{id}/documents, which deletes every document of the pipeline.

Initial load​

The indexing service loads existing messages through the same create request, one request per message. Foundation4 has no bulk create request.

The following limits govern an initial load:

  • Queue retention. Each create request places a processing job in NATS JetStream, the queue between the API server and the workers. The queue keeps a job for at most 24 hours. A job still waiting after 24 hours is removed without processing, and the document remains in pending status. An initial load is therefore paced so that the workers process the queue within 24 hours.
  • Worker throughput. Each worker takes up to 100 documents at a time from the queue. Throughput depends on the embedding model and on the number of workers, which the operator scales as described in Scaling and performance.

The number of documents waiting for processing is the total count of pending documents:

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?status=pending&count=true&first=1" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200, with the number of pending documents in page_info.count. The count falls toward 0 as the workers process the load.

Processing metrics​

The API server and the workers publish Prometheus metrics for document processing. Monitoring and logging lists every metric and the scrape configuration for the workers:

MetricSourceMeaning
foundation4ai_document_queue_addedAPI server, /metrics on port 8000Document create requests received since the API server started that pass the key check and the parsing of the body, including requests that then fail validation
foundation4ai_document_queue_processedWorker, port 9090Processing attempts, including retries
foundation4ai_document_queue_successWorker, port 9090Documents processed successfully
foundation4ai_document_queue_failedWorker, port 9090Failed processing attempts, including every document of the same pipeline in a failed batch

A failed count that rises during an initial load, or a pending count that does not fall, indicates a processing problem. The worker retries a failed document after 10 seconds, and the worker log names the failing document or pipeline and the cause.

Indexing API key​

The indexing service uses a dedicated API key, separate from the search key. The key holds all three permissions on the pipeline (permission 7), so that the service can add, find, expire and delete documents, and execute permission on the pipeline's text splitter (permission 1). The key's classification allow-list names the classifications that the service assigns to messages. Access control describes the permissions and the allow-list.