Relay a chat completion
POST/openai/v1/chat/completions
Relays an OpenAI chat completion request to the chat completions API of the LLM named in the x-llm-id header. Foundation4 forwards the body unchanged, inserts the registered model name when the body has no model field and sends the API key of the registration, when the registration has one, as a bearer token. Foundation4 returns the model server's status and body, including error responses, and performs no retrieval, prompt templating or tracing.
Permissions. Execute on the LLM.
LLMs describes the endpoint.
Request
Responses
- 200
- 400
- 401
- 404
- 500
- 501
- default
The model server's response, relayed unchanged. A non-streamed request receives the model server's body and headers, typically a chat completion object in JSON; a request with "stream": true receives the model server's server-sent events as text/event-stream.
The x-llm-id header is missing or is not a valid identifier (Invalid LLM ID: <reason>), or the request carries one or two of the headers X-Agent-Id, X-Pipeline-Id and X-Pipeline-Classification (Missing required headers: ...).
The authentication headers are missing or invalid, or the API key is inactive or expired. Errors lists the messages.
No LLM with this identifier exists, or the key lacks execute permission on the LLM (LLM not found).
Foundation4 could not reach the model server or read the model server's response. The message is the error text of the HTTP client.
The request carries all three headers X-Agent-Id, X-Pipeline-Id and X-Pipeline-Classification (F4ai enriched completions not implemented yet).
An error response of the model server, relayed with the model server's status and body.