Agents and prompt templates
An agent is a stored prompt template that combines retrieval from a pipeline with a call to a registered LLM, which produces an answer grounded in the pipeline's content. This retrieval-augmented generation (RAG) pattern is the basis of question answering over a knowledge base. This page describes the prompt template, the placeholders that insert request values and retrieved fragments, the execution of an agent and the tools for testing an agent. Integrators need this page to design agents and to call agents from a client application.
Agent definition
An agent is created with POST /agents. The request contains the following fields.
| Field | Required | Description |
|---|---|---|
name | Yes | A name that is unique across the deployment |
description | No | Free text for administrators |
prompt | Yes | The prompt template: an ordered list of messages |
placeholders | Yes | The named values that the templates insert: request values and retrieval results |
metadata | No | Descriptive data for the client application. Each top-level value is a JSON object. Foundation4 stores metadata and does not use metadata during execution. |
An agent stores no pipeline and no LLM. Each execution names the pipeline and the LLM, so one agent can serve several pipelines and several models.
Prompt template
The prompt is a list of messages in the order that the model receives them. Each message has the following fields:
role.systemfor instructions to the model,userfor the request that the model answers, orassistantfor an example answer.template. The text of the message, with placeholders written as the placeholder name in braces, such as{question}.include. Whether the message is sent to the model. The default istrue. A message withfalseremains part of the stored agent but is never sent.
A prompt contains at least one system message and at least one user message, and the last message is a user message. Every name in braces in a template must be a placeholder of the agent. Foundation4 provides no built-in variables such as {context}; each name refers to a placeholder that the agent defines. Placeholder names consist of letters and digits and start with a letter. Foundation4 rejects an agent that breaks one of these rules with HTTP 400.
Placeholders
A placeholder is a named value that Foundation4 inserts into the templates at execution time. Each placeholder has a name, a type and, depending on the type, a target and params.
| Type | Value inserted | Fields |
|---|---|---|
query | A value that the client application supplies in each execution request, such as the end user's question | name, type |
similarity | The text of the fragments that a similarity search returns | name, type, target, params |
mmr | The text of the fragments that a maximal marginal relevance (MMR) search returns | name, type, target, params |
A retrieval placeholder (similarity or mmr) names a query placeholder in target. At execution, Foundation4 uses the value of that query placeholder as the search text, so the end user's question both appears in the prompt and selects the fragments. Several retrieval placeholders can target the same query placeholder or different query placeholders.
The params object of a retrieval placeholder accepts the parameters of the corresponding search mode, with the same defaults as described in Search and retrieval:
| Parameter | Placeholder types | Default | Description |
|---|---|---|---|
k | Similarity, MMR | 5 | The number of fragments to insert |
fetch_k | MMR | 10 | The number of candidates from which MMR selects |
lambda_mult | MMR | 0.5 | The balance between relevance and diversity, from 0 to 1 |
threshold | Similarity, MMR | None | The maximum cosine distance of a fragment from the search text |
The defaults apply when a placeholder omits params. A placeholder that includes params sets k, and for MMR also fetch_k, because a field omitted inside params takes the value 0 and Foundation4 rejects a placeholder with k of 0. Retrieval in agents always uses the cosine distance strategy and the pipeline's embedding model.
Foundation4 inserts the text of the retrieved fragments in rank order, separated by line breaks. The inserted text carries no fragment identifiers, scores or metadata. A template that needs the model to distinguish sources therefore labels the section that holds the placeholder, as in the example below.
Full-text search, hybrid search and conversation history are not available as placeholders.
The shell variables in the examples are described in Authenticate. The following request creates a question-answering agent with one query placeholder and one similarity placeholder:
curl -X POST "$FOUNDATION4_URL/agents" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"name": "support-answers",
"description": "Answers support questions from the product knowledge base",
"prompt": [
{
"role": "system",
"template": "Answer using only the reference text. If the reference text does not contain the answer, say that the answer is not available."
},
{
"role": "user",
"template": "Reference text:\n{context}\n\nQuestion: {question}"
}
],
"placeholders": [
{"name": "question", "type": "query"},
{"name": "context", "type": "similarity", "target": "question", "params": {"k": 5}}
]
}'
The response has status 201 and contains the stored agent, including the generated id, the include value of each message and every parameter of each retrieval placeholder.
Execution
POST /agents/{id}/execute runs an agent. The request carries two headers in addition to the API key headers:
x-pipeline-id. The pipeline that the retrieval placeholders search. One execution searches one pipeline.x-llm-id. The registered LLM that generates the answer. The LLMs page describes the requirements on the model server.
The request body contains the following fields.
| Field | Required | Description |
|---|---|---|
prompt | Yes | An object that maps the name of each query placeholder to the value for this execution. Every query placeholder must have a value. |
classification | Yes | The classifications that the reader holds, in any form that search accepts. The Classifications page describes the forms. |
filters | No | An object that maps the name of a retrieval placeholder to a metadata filter for that placeholder. The Metadata, filters and taxonomies page describes the filter language. |
stream | No | The response format. The default is newline-delimited JSON. |
tracing | No | Records a trace of the execution when true. The default is false. |
temperature | No | The sampling temperature. When omitted, the model server's default applies. |
Filters apply per placeholder, so one agent can retrieve from different subsets of a pipeline in one execution. For example, one placeholder can retrieve from product documentation and a second placeholder from release notes, each with a filter on a source metadata field.
The following request runs the agent created above and waits for the complete answer:
curl -X POST "$FOUNDATION4_URL/agents/$AGENT_ID/execute" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "x-pipeline-id: $PIPELINE_ID" \
-H "x-llm-id: $LLM_ID" \
-H "Content-Type: application/json" \
-d '{
"prompt": {"question": "How do I rotate the database credentials?"},
"classification": ["internal"],
"stream": false,
"tracing": true,
"temperature": 0.2
}'
The response has status 200 and contains the answer as plain text. The streamed formats and the handling of failures are the same as for direct queries, as described in Response formats on the LLMs page. A request that omits the value of a query placeholder returns HTTP 400, with the missing placeholder names in the error details.
Execution path
The API server runs the retrieval placeholders in the order of the placeholders list. Each retrieval applies the classifications of the request and the filter for that placeholder, as a search request does. The API server then inserts the input values and the fragment text into the templates, removes messages with include set to false, and sends the messages to the model server. The system messages become the instructions of the model, the other messages form the conversation in order, and the last message is the request that the model answers.
After the values are inserted, no text in the messages may have the form of a placeholder. An execution whose input values or retrieved fragments contain text in braces that has the form of a placeholder, such as {id}, returns HTTP 400.
Stateless execution
Agents keep no conversation state. Each execution is independent, and Foundation4 stores nothing about an execution except the optional trace. A client application that continues a conversation supplies the earlier turns in the input values of each execution, for example through a query placeholder that holds the conversation so far.
Testing and tracing
Three endpoints support the design and debugging of an agent:
- Prompt inspection.
GET /agents/{id}/promptreturns the stored messages, the placeholders andinput_variables, the names of the query placeholders that each execution must supply. - Retrieval dry run.
POST /agents/{id}/searchruns the retrieval placeholders without calling a model. The request takes thex-pipeline-idheader and theprompt,classificationandfiltersfields of an execution. The response maps each retrieval placeholder name to the fragments that the placeholder retrieves, with scores. - Execution trace. An execution with
tracingset totruereturns the trace identifier in thex-foundation4ai-tracing-idresponse header.GET /tracing/{id}returns the trace: the input values, the fragments retrieved for each placeholder and the messages exactly as sent to the model. The trace does not contain the answer.
A trace is retained for 60 minutes and can be read only with the API key that ran the execution. A trace holds the decrypted text of the retrieved fragments, encrypted with the application secret while stored.
Permissions
Creating an agent requires write permission on agents. Executing an agent requires execute permission on the agent, on the pipeline and on the LLM that the request names. The retrieval dry run requires execute permission on the agent and on the pipeline, and prompt inspection requires read permission on the agent. Access control describes permissions.