Skip to main content

First RAG answer

This tutorial registers an LLM, creates an agent that retrieves fragments from the pipeline of First search and runs the agent to produce an answer grounded in the pipeline's content. This pattern is retrieval-augmented generation (RAG). Integrators complete this tutorial to see the question-answering path end to end. The tutorial continues with the getting-started key, the shell variables and the support-kb pipeline from the earlier tutorials.

Model server​

Foundation4 does not host language models. An LLM registration points to a model server that the deployment can reach, on the same network or in an air-gapped environment. Agents call the model server through the OpenAI Responses API with an API key, so the model server implements the Responses API and the registration carries a key. LLMs describes the requirements.

LLM registration​

curl -X POST "$FOUNDATION4_URL/llms" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"name": "tutorial-llm",
"description": "Model for the getting-started tutorials",
"endpoint": "http://llm.internal:8000/v1",
"model": "example-model-8b-instruct",
"api_key": "change-me"
}'

Replace endpoint, model and api_key with the values of the model server. endpoint is the base URL of the OpenAI-compatible API, which ends in /v1 on most servers.

Expected result: status 201 and the registration. The response never contains the API key:

{
"id": "3c9e1f2a-6b7d-4e8f-9a0b-1c2d3e4f5a6b",
"name": "tutorial-llm",
"description": "Model for the getting-started tutorials",
"endpoint": "http://llm.internal:8000/v1",
"model": "example-model-8b-instruct"
}

Set the variable from the id field:

export LLM_ID=<id from the response>

Model check​

A direct query sends one prompt to the model server without retrieval. "stream": false returns the complete answer as plain text and reports a model server failure as an error, which makes the direct query a reliable check of the registration.

curl -X POST "$FOUNDATION4_URL/llms/$LLM_ID/query" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{"query": "Reply with the single word: ready", "stream": false}'

Expected result: status 200 and a plain-text answer such as ready. HTTP 500 with the message Stream error means that the model server rejected or failed the request.

Agent​

The agent in this tutorial has two placeholders. question is a query placeholder, which the client application fills in each execution. context is a similarity placeholder, which searches the pipeline with the value of question and inserts the text of the 3 closest fragments. Agents and prompt templates describes placeholders.

curl -X POST "$FOUNDATION4_URL/agents" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "Content-Type: application/json" \
-d '{
"name": "support-qa",
"description": "Answers support questions from the knowledge base",
"prompt": [
{
"role": "system",
"template": "Answer using only the reference text. If the reference text does not contain the answer, say that the answer is not available."
},
{
"role": "user",
"template": "Reference text:\n{context}\n\nQuestion: {question}"
}
],
"placeholders": [
{"name": "question", "type": "query"},
{"name": "context", "type": "similarity", "target": "question", "params": {"k": 3}}
]
}'

Expected result: status 201 and the stored agent, with an include value on each message and every parameter of the similarity placeholder. Set the variable from the id field:

export AGENT_ID=<id from the response>

Prompt inspection​

curl "$FOUNDATION4_URL/agents/$AGENT_ID/prompt" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET"

Expected result: status 200, with input_variables set to ["question"]: the values that each execution supplies.

Retrieval dry run​

The dry run performs the retrieval of every retrieval placeholder without calling the model server. The x-pipeline-id header names the pipeline to search.

curl -X POST "$FOUNDATION4_URL/agents/$AGENT_ID/search" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "x-pipeline-id: $PIPELINE_ID" \
-H "Content-Type: application/json" \
-d '{
"prompt": {"question": "How can finance export invoices?"},
"classification": "internal"
}'

Expected result: status 200 and an object with one key for each retrieval placeholder. The context key holds the fragments that the placeholder retrieves, in the format of a search result, with the invoice fragment first.

Execution​

An execution names the pipeline in x-pipeline-id and the LLM in x-llm-id. The request body supplies the value of each query placeholder and the reader's classification.

curl -X POST "$FOUNDATION4_URL/agents/$AGENT_ID/execute" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "x-pipeline-id: $PIPELINE_ID" \
-H "x-llm-id: $LLM_ID" \
-H "Content-Type: application/json" \
-d '{
"prompt": {"question": "How can finance export invoices?"},
"classification": "internal",
"stream": false
}'

Expected result: status 200 and the answer as plain text, based on the invoice fragment, such as the following:

Finance staff can export invoices as PDF or CSV files from the Billing page.

Repeat the request with "classification": "public". The retrieval then finds only the password reset fragment, and the answer states that the answer is not available.

Streamed execution​

Client applications usually stream the answer, so that the reader sees each part of the answer as the model generates the text. curl -N prints each part on arrival:

curl -N -X POST "$FOUNDATION4_URL/agents/$AGENT_ID/execute" \
-H "x-api-key: $FOUNDATION4_API_KEY" \
-H "x-api-key-secret: $FOUNDATION4_API_SECRET" \
-H "x-pipeline-id: $PIPELINE_ID" \
-H "x-llm-id: $LLM_ID" \
-H "Content-Type: application/json" \
-d '{
"prompt": {"question": "How can finance export invoices?"},
"classification": "internal"
}'

Expected result: status 200 and newline-delimited JSON, one object per line, each with the next part of the answer in content:

{"content":"Finance staff can export"}
{"content":" invoices as PDF or CSV files"}
{"content":" from the Billing page."}

Response formats describes the streamed formats, including server-sent events.