Skip to main content

Call an LLM through Foundation4

This guide registers a model server as an LLM, grants a client application's key permission to use the LLM and sends prompts through the direct query endpoint and the OpenAI-compatible chat completions endpoint, including from the OpenAI Python library. Integrators need this guide to route a client application's model calls through Foundation4, and administrators need this guide to register model servers and grant access. LLMs describes LLM registrations, the response formats and the OpenAI-compatible endpoint.

Model server requirements​

The model server exposes an OpenAI-compatible API at an address that the API server reaches inside the deployment's network, such as http://llm.internal:8000/v1. The two endpoints in this guide call different APIs of the model server:

  • Direct query. POST /llms/{id}/query calls the Responses API of the model server and requires an API key in the registration. A model server without authentication accepts any non-empty key.
  • OpenAI-compatible endpoint. POST /openai/v1/chat/completions calls the chat completions API of the model server and sends the registered API key when the registration has one.

A model server that implements only chat completions therefore serves the OpenAI-compatible endpoint but not the direct query or agents. Model servers lists the requirements of each feature.

LLM registration​

Register the model server with POST /llms and an administration key that holds write permission on the llms object type. The shell variables are described in Authenticate.

curl -X POST "$FOUNDATION4_URL/llms" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"name": "internal-chat",
"description": "Chat model on the internal model server",
"endpoint": "http://llm.internal:8000/v1",
"model": "example-model-8b-instruct",
"api_key": "change-me"
}'

Replace endpoint, model and api_key with the values of the model server. endpoint is the base URL of the OpenAI-compatible API without a trailing slash.

Expected result: status 201 and the registration. The response never contains the API key:

{
"id": "7f1bda72-3e84-40df-8cf0-5375402c931b",
"name": "internal-chat",
"description": "Chat model on the internal model server",
"endpoint": "http://llm.internal:8000/v1",
"model": "example-model-8b-instruct"
}
export LLM_ID=<id from the response>

Registration does not contact the model server. LLM registration describes each field.

Client key permission​

The client application authenticates with a dedicated key, created as described in Manage API keys and permissions. Set the variables from that key:

export CLIENT_KEY_ID=<id of the client application's key>
export CLIENT_KEY_SECRET=<secret of the client application's key>

Both endpoints in this guide require execute permission on the LLM. Grant the permission with the administration key and POST /api-keys/{id}/permissions:

curl -X POST "$FOUNDATION4_URL/api-keys/$CLIENT_KEY_ID/permissions" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"permissions": [
{"object_type": "llms", "object_id": "'"$LLM_ID"'", "permission": 1}
]
}'

Expected result: status 200 and an empty body. The remaining requests authenticate with the client application's key.

Direct query​

POST /llms/{id}/query sends one prompt to the LLM as a single user message. "stream": false returns the complete answer and reports a model server failure as an error, so this setting tests a new registration, as in Model check:

curl -X POST "$FOUNDATION4_URL/llms/$LLM_ID/query" \
-H "x-api-key: $CLIENT_KEY_ID" \
-H "x-api-key-secret: $CLIENT_KEY_SECRET" \
-H "Content-Type: application/json" \
-d '{"query": "In one sentence, what does a text splitter do?", "stream": false, "temperature": 0.2}'

Expected result: status 200 and the answer as plain text:

A text splitter divides a document into fragments that are embedded and searched separately.

A request without the stream field returns newline-delimited JSON (NDJSON), one JSON object per line with the next part of the answer in content:

curl -X POST "$FOUNDATION4_URL/llms/$LLM_ID/query" \
-H "x-api-key: $CLIENT_KEY_ID" \
-H "x-api-key-secret: $CLIENT_KEY_SECRET" \
-H "Content-Type: application/json" \
-d '{"query": "In one sentence, what does a text splitter do?"}'

Expected result: status 200 and lines such as the following:

{"content":"A text splitter"}
{"content":" divides a document"}
{"content":" into fragments."}

"stream": "sse" returns server-sent events, with the same JSON object in the data field of each event:

curl -X POST "$FOUNDATION4_URL/llms/$LLM_ID/query" \
-H "x-api-key: $CLIENT_KEY_ID" \
-H "x-api-key-secret: $CLIENT_KEY_SECRET" \
-H "Content-Type: application/json" \
-d '{"query": "In one sentence, what does a text splitter do?", "stream": "sse"}'

Expected result: status 200 and events such as the following. Lines that contain only a colon are keep-alive comments, sent after each second without an event:

data: {"content":"A text splitter"}

data: {"content":" divides a document"}

data: {"content":" into fragments."}

Both streamed formats end when the answer is complete, and a line or event can carry an empty content value. Response formats describes the formats, and Errors during streamed responses describes what a client application observes when the model server fails during a streamed response.

OpenAI-compatible endpoint​

POST /openai/v1/chat/completions relays a chat completion request in the OpenAI format to the model server. The request carries the Foundation4 API key headers and the x-llm-id header with the LLM identifier:

curl -X POST "$FOUNDATION4_URL/openai/v1/chat/completions" \
-H "x-api-key: $CLIENT_KEY_ID" \
-H "x-api-key-secret: $CLIENT_KEY_SECRET" \
-H "x-llm-id: $LLM_ID" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What does a text splitter do?"}
],
"temperature": 0.2
}'

Expected result: the model server's status code, headers and body. Foundation4 inserts the registered model name when the body has no model field, so the answer comes from the registered model. The body of an OpenAI-compatible model server has the following form:

{
"id": "chatcmpl-5c1e9a7d",
"object": "chat.completion",
"created": 1790760300,
"model": "example-model-8b-instruct",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "A text splitter divides a document into fragments that are embedded and searched separately."
},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 24, "completion_tokens": 17, "total_tokens": 41}
}

A body with "stream": true returns the model server's stream as server-sent events:

curl -X POST "$FOUNDATION4_URL/openai/v1/chat/completions" \
-H "x-api-key: $CLIENT_KEY_ID" \
-H "x-api-key-secret: $CLIENT_KEY_SECRET" \
-H "x-llm-id: $LLM_ID" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "What does a text splitter do?"}
],
"stream": true
}'

Expected result: the model server's status code and the events exactly as the model server sends the events, with the content type text/event-stream. The following events are shortened to the first content event, the last content event and the end marker of an OpenAI-compatible model server:

data: {"id":"chatcmpl-5c1e9a7e","object":"chat.completion.chunk","created":1790760310,"model":"example-model-8b-instruct","choices":[{"index":0,"delta":{"role":"assistant","content":"A text splitter"},"finish_reason":null}]}

data: {"id":"chatcmpl-5c1e9a7e","object":"chat.completion.chunk","created":1790760310,"model":"example-model-8b-instruct","choices":[{"index":0,"delta":{"content":" separately."},"finish_reason":"stop"}]}

data: [DONE]

OpenAI Python library​

Client applications written for the OpenAI chat completions API use the OpenAI-compatible endpoint through the library's base URL and default headers. The following script uses the openai package and reads the exported shell variables:

import os

from openai import OpenAI

client = OpenAI(
base_url=os.environ["FOUNDATION4_URL"] + "/openai/v1",
api_key="unused",
default_headers={
"x-api-key": os.environ["CLIENT_KEY_ID"],
"x-api-key-secret": os.environ["CLIENT_KEY_SECRET"],
"x-llm-id": os.environ["LLM_ID"],
},
)

completion = client.chat.completions.create(
model="example-model-8b-instruct",
messages=[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What does a text splitter do?"},
],
temperature=0.2,
)
print(completion.choices[0].message.content)

Expected result: the answer text on standard output. The library requires an api_key value and sends the value as a bearer token. Foundation4 authenticates the request with the x-api-key and x-api-key-secret headers and does not forward the bearer token to the model server, so the placeholder unused serves. The library also requires the model argument, so the script passes the registered model name.

Endpoint scope​

The OpenAI-compatible endpoint is a relay to one registered LLM:

  • Model. A model field in the body takes precedence over the registered model name. Client applications omit the field, or send the registered name, so that the registration decides which model answers.
  • Model server key. Foundation4 sends the registered API key to the model server, so the model server's key stays out of client application configuration.
  • Other OpenAI paths. Chat completions is the only relayed API. Paths such as /openai/v1/models return HTTP 404.
  • Retrieval. The endpoint performs no retrieval and applies no prompt template. A request with any of the headers X-Agent-Id, X-Pipeline-Id or X-Pipeline-Classification returns an error. Answers grounded in a pipeline's content use agents, as described in Agents and prompt templates.

Errors​

The following errors are specific to LLM calls. Errors lists every message.

  • Missing LLM header. A request to the OpenAI-compatible endpoint without x-llm-id returns HTTP 400 with a plain-text body, and a value that is not a universally unique identifier (UUID) returns HTTP 400 with Invalid LLM ID: <reason>.
  • Unknown LLM. The direct query returns HTTP 404 with LLM with id <id> not found, and the OpenAI-compatible endpoint returns HTTP 404 with LLM not found. Both mean that the LLM does not exist or that the key lacks execute permission on the LLM.
  • Missing model server key. The direct query returns HTTP 400 with Invalid parameter and details.error set to API key is required for OpenAI provider when the registration has no API key.
  • License. The direct query returns HTTP 403 with Invalid license: <reason> when the license is not valid.
  • Model server failure. The direct query with "stream": false returns HTTP 500 with Stream error: <reason>. The OpenAI-compatible endpoint returns the model server's error status and body unchanged, and HTTP 500 when the model server cannot be reached.
  • Retrieval headers. The OpenAI-compatible endpoint returns HTTP 501 when a request carries all three retrieval headers, and HTTP 400 when a request carries one or two.