Skip to main content

Paginate and filter lists

This guide reads list operations page by page and narrows the results with query parameters: the cursor and offset pagination modes, the total count, the sort order, field filters, repeated parameters and the page_info and query_info objects of each response. The last section reads every page of a list safely in a client application. Integrators who synchronize objects or documents with another system, and administrators who script inventories of a deployment, need this guide. API conventions defines the parameters.

Pagination modes​

The query parameters of a request select one of two pagination modes:

  • Cursor mode. The default mode, with first for the page size and after for the position. Each page starts after the last object of the previous page, so objects added or removed during a read do not shift the later pages. Cursor mode suits complete reads and lists that change.
  • Offset mode. Selected by limit or offset. Each page starts at a position in the sorted list, so a client application can open page 5 directly. Objects added or removed during a read shift the later pages, which can skip or repeat objects.

A request that combines the parameters of both modes, or first with last, returns HTTP 400 with the message Pagination error. The default page size is 25 in both modes. Client applications request between 1 and 100 objects per page: in cursor mode, first=0 returns an empty page with has_next set to true and no cursor.

Example list​

The examples read the documents of the support-kb pipeline from First search, with the shell variables from that tutorial. The pipeline holds three documents, created two seconds apart in the order kb-password-reset, kb-invoice-export and kb-refund-policy. Document identifiers are time-ordered, so the order by id matches the order of creation.

First page in cursor mode​

Request the first page of two documents, the total count and an order by id and created_at:

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?first=2&count=true&order_by=id&order_by=created_at" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and the first two documents. The following response omits the document fields other than id, created_at and external_identifier:

{
"data": [
{
"id": "01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f",
"created_at": "2026-09-29T12:05:00.418273Z",
"external_identifier": "kb-password-reset"
},
{
"id": "01a0ed0d-d75f-7d5e-8f60-718293a4b5c6",
"created_at": "2026-09-29T12:05:02.431906Z",
"external_identifier": "kb-invoice-export"
}
],
"page_info": {
"after": "WyIwMWEwZWQwZC1kNzVmLTdkNWUtOGY2MC03MTgyOTNhNGI1YzYiLCIyMDI2LTA5LTI5VDEyOjA1OjAyLjQzMTkwNloiXQ",
"before": "WyIwMWEwZWQwZC1jZjgyLTdjM2QtOGU0Zi01YTZiN2M4ZDllMGYiLCIyMDI2LTA5LTI5VDEyOjA1OjAwLjQxODI3M1oiXQ",
"count": 3,
"has_next": true,
"has_prev": false,
"order_by": ["id", "created_at"]
},
"query_info": {}
}

count is the total number of documents that match the request, 3, not the number on the page. has_next is true, so another page follows. A document added with expire_older_versions set to false keeps several current versions, which share the id and appear as separate entries of the document list. The order by id and then created_at therefore gives every entry a unique position.

Next page​

Set a variable from page_info.after and repeat the request with the same page size and order:

export AFTER=<page_info.after from the response>

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?first=2&order_by=id&order_by=created_at&after=$AFTER" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and the third document. The following response omits the document fields other than id, created_at and external_identifier:

{
"data": [
{
"id": "01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b",
"created_at": "2026-09-29T12:05:04.447128Z",
"external_identifier": "kb-refund-policy"
}
],
"page_info": {
"after": "WyIwMWEwZWQwZC1kZjNmLTdmODAtOWExYi0yYzNkNGU1ZjZhN2IiLCIyMDI2LTA5LTI5VDEyOjA1OjA0LjQ0NzEyOFoiXQ",
"before": "WyIwMWEwZWQwZC1kZjNmLTdmODAtOWExYi0yYzNkNGU1ZjZhN2IiLCIyMDI2LTA5LTI5VDEyOjA1OjA0LjQ0NzEyOFoiXQ",
"has_next": false,
"has_prev": true,
"order_by": ["id", "created_at"]
},
"query_info": {}
}

has_next is false, so the list ends with this page. The request omits count, because each count runs an additional database query and the first page already reported the total. Cursors consist of URL-safe characters and need no encoding.

A cursor records the values of the order fields of one object, so a cursor is valid only with the order of the request that returned the cursor. A cursor with a different number of order fields returns HTTP 400 with the message Pagination error.

Offset mode​

Request the second page of two documents, in descending order of creation:

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?limit=2&offset=2&count=true&order_by=-created_at" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and the oldest document, the third in descending order. The following response omits the document fields other than id, created_at and external_identifier:

{
"data": [
{
"id": "01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f",
"created_at": "2026-09-29T12:05:00.418273Z",
"external_identifier": "kb-password-reset"
}
],
"page_info": {
"count": 3,
"limit": 2,
"offset": 2,
"order_by": ["-created_at"]
},
"query_info": {}
}

Offset mode returns no has_next and no cursors. More objects follow while offset plus limit is less than count, and without count, while a page holds limit objects. The page at position n, counting from 1, starts at an offset of limit multiplied by (n minus 1).

Sort order​

order_by sets the order of a list, with the following rules:

  • Direction. A field name alone, such as order_by=created_at, sorts ascending, and a - prefix, such as order_by=-created_at, sorts descending.
  • Plus prefix. A + prefix also sorts ascending, but a + in a query string decodes to a space. The prefix is therefore sent encoded, as order_by=%2Bcreated_at, or omitted. An unencoded + returns HTTP 400 with the message Invalid order_by parameter.
  • Several fields. Each order_by parameter adds a sort key, in the order of the parameters. order_by=classification&order_by=-created_at sorts by classification and, within each classification, by descending creation time.
  • Default order. A request without order_by sorts by id ascending. A request with order_by sorts by the requested fields only, so a list read in cursor mode ends the order with a unique field, such as id.
  • Sortable fields. Each list operation accepts a fixed set of fields, listed on the endpoint page. List documents accepts id, created_at, updated_at, external_identifier and classification.

A field that the operation cannot sort by returns HTTP 400:

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?order_by=status" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 400 and the rejected value in details:

{
"error": {
"message": "Invalid order_by parameter",
"details": ["Unknown value `status`"]
}
}

Field filters​

List operations accept filters named after the fields of the object, with an optional operator suffix. The following request lists the pipelines whose name starts with support and contains kb. The $ of each operator suffix is escaped as \$ inside double quotes:

curl "$FOUNDATION4_URL/pipelines?name\$startswith=support&name\$contains=kb" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and the support-kb pipeline. The following response omits the pipeline fields other than id and name:

{
"data": [
{
"id": "c4a7e2d1-6b3f-4f8e-a2d9-7e1b5c3f9a06",
"name": "support-kb"
}
],
"page_info": {
"after": "WyJjNGE3ZTJkMS02YjNmLTRmOGUtYTJkOS03ZTFiNWMzZjlhMDYiXQ",
"before": "WyJjNGE3ZTJkMS02YjNmLTRmOGUtYTJkOS03ZTFiNWMzZjlhMDYiXQ",
"has_next": false,
"has_prev": false,
"order_by": null
},
"query_info": {
"name$contains": ["kb"],
"name$startswith": "support"
}
}

The operators apply as follows:

  • Equality. field=value matches the value exactly.
  • $contains. Matches values that contain the text, case-sensitive.
  • $startswith. Matches values that start with the text, case-sensitive.
  • $gt, $ge, $lt and $le. Compare the field with the value: $gt is greater than, $ge greater than or equal to, $lt less than and $le less than or equal to. The comparison operators apply to timestamps and, on some lists, to names.

Several filters in one request must all match. The endpoint page of each list operation names the fields and operators that the operation accepts.

In the current release, _ and % in a $contains or $startswith value act as wildcards: _ matches any single character and % matches any sequence of characters, so name$contains=_ matches every name. A client application that filters on text containing these characters checks the returned values.

Timestamp filters​

Timestamps in filters use RFC 3339 format. A timestamp with an offset contains a +, which curl -G with --data-urlencode encodes. The following request lists the documents that were created at or after 14:05:01 at an offset of 2 hours from Coordinated Universal Time (UTC) and whose external identifier contains kb-:

curl -G "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents" \
--data-urlencode 'created_at$ge=2026-09-29T14:05:01+02:00' \
--data-urlencode 'external_identifier$contains=kb-' \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and the documents kb-invoice-export and kb-refund-policy. The following response omits data and page_info:

{
"query_info": {
"created_at$ge": "2026-09-29T12:05:01Z",
"external_identifier$contains": ["kb-"]
}
}

query_info repeats each recognized filter, with the timestamp converted to UTC. A timestamp written with the suffix Z, such as created_at$ge=2026-09-29T12:05:01Z, needs no encoding. A timestamp with an unencoded + returns HTTP 400 with a plain-text body.

Repeated parameters​

A query parameter can appear more than once only where the operation expects a list:

  • order_by. Each occurrence adds a sort key, as described in Sort order.
  • $contains filters. Each occurrence adds a condition, and every condition must match. name$contains=support&name$contains=kb matches names that contain both texts.
  • Other parameters. A repeated filter, such as status=pending&status=failed, returns HTTP 400 with a plain-text body that starts with Failed to deserialize query string. A list operation matches one value per filter, so a client application that needs several values sends one request for each value.

Unknown parameters​

The API server ignores query parameters that an operation does not define, including misspelled filter names. The following request misspells external_identifier:

curl "$FOUNDATION4_URL/pipelines/$PIPELINE_ID/documents?external_identifer\$contains=refund&count=true" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200 and all three documents, because no filter applies. The following response omits data and the page_info fields other than count:

{
"page_info": {
"count": 3
},
"query_info": {}
}

An empty query_info shows that the API server recognized no filter. The embedding provider list is an exception in the current release: its query_info lists every filter of the operation, with null or an empty list for each filter that the request did not set. A client application that builds filters dynamically compares query_info with the filters that the client application sent.

Page information​

page_info describes the page in the following fields:

  • after and before. Cursor mode only. The cursors of the last and the first object on the page. An empty page carries no cursors.
  • has_next. Cursor mode only. When paging forward with first, true when more objects follow the page in the requested order.
  • has_prev. Cursor mode only. A client application does not decide from has_prev whether earlier pages exist. A client application that returns to earlier pages keeps the before cursor of each page that the client application has read.
  • count. Present when the request sets count=true. The total number of objects that match the filters, independent of the page.
  • limit and offset. Offset mode only. The page size and position that the API server applied, including the defaults.
  • order_by. The requested order, with a - prefix for descending fields, or null for the default order by id.

Backward paging uses last and before: last=2&before=<page_info.before> returns the two objects that precede the first object of a page, in the requested order. A client application that pages backward stops at a page with fewer than last objects.

Complete reads​

A client application that reads every object of a list uses cursor mode and follows page_info.after until has_next is false. The following Python script uses only the standard library and prints the identifier of every document in the pipeline. The script reads the shell variables from Authenticate and PIPELINE_ID:

import json
import os
import urllib.parse
import urllib.request

BASE_URL = os.environ["FOUNDATION4_URL"]
HEADERS = {
"x-api-key": os.environ["FOUNDATION4_API_KEY"],
"x-api-key-secret": os.environ["FOUNDATION4_API_SECRET"],
}


def list_all(path, params):
"""Yield every object of a list operation, one page at a time."""
after = None
while True:
query = list(params)
if after:
query.append(("after", after))
url = f"{BASE_URL}{path}?{urllib.parse.urlencode(query)}"
request = urllib.request.Request(url, headers=HEADERS)
with urllib.request.urlopen(request) as response:
page = json.load(response)
yield from page["data"]
if not page["page_info"].get("has_next"):
return
after = page["page_info"]["after"]


pipeline_id = os.environ["PIPELINE_ID"]
params = [
("first", "100"),
("order_by", "id"),
("order_by", "created_at"),
("status", "success"),
]
for document in list_all(f"/pipelines/{pipeline_id}/documents", params):
print(document["id"], document["external_identifier"])

Expected result: one line for each processed document:

01a0ed0d-cf82-7c3d-8e4f-5a6b7c8d9e0f kb-password-reset
01a0ed0d-d75f-7d5e-8f60-718293a4b5c6 kb-invoice-export
01a0ed0d-df3f-7f80-9a1b-2c3d4e5f6a7b kb-refund-policy

A complete read follows these rules:

  • Unique order. The order ends with fields that together identify each object: id, the default, for most lists, and id followed by created_at for the document list. An order by a field whose values repeat, such as classification, adds these fields at the end.
  • Same parameters. Every page request repeats the filters, the order and the page size of the first request, and changes only after.
  • Stop condition. The read stops when has_next is false. The script does not compute the number of pages from count, because the list can change during the read.
  • Encoding. urllib.parse.urlencode encodes $, + and the characters of JSON filters, and a list of pairs repeats a parameter such as order_by.
  • Errors. urllib.request.urlopen raises an exception for every status of 400 or higher, so a failed page stops the read instead of ending the read early without notice.