Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

Chat Completions

POST 

/api/v1/workspaces/:workspace_name/deployments/v1/chat/completions

Run a deployment's active revision as a chat completion (OpenAI-style).

The chat completion endpoint sends a list of conversation messages to an AI model and returns the next assistant response. For multi-turn chats, clients maintain the history by including prior user and assistant messages in each request, along with the new user message. The run is stored in search history so you can inspect traces.

By default (stream: false), a single JSON response is returned once generation completes. Pass stream: true to instead receive the response streamed as SSE.

model identifies which deployed service to run, as {workspace_name}/{deployment_id}.

Example request:

{
"messages": [
{"role": "user", "content": "How does streaming work with Haystack Enterprise Platform?"}
],
"model": "my_workspace/3fa85f64-5717-4562-b3fc-2c963f66afa6"
}

Example non-streaming (stream: false, the default) response:

{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1730000000,
"model": "my_workspace/3fa85f64-5717-4562-b3fc-2c963f66afa6",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Streaming sends the response as SSE chunks."},
"finish_reason": "stop"
}
]
}

Extra OpenAI parameters​

You can send extra OpenAI fields besides messages, model, and stream, for example temperature, top_p, and other generation parameters. They are placed in a generation_kwargs dict and passed to the deployed service as a named input.

OpenAI-style tools and tool_choice are not supported and are ignored from the request.

Example request:

{
"messages": [{"role": "user", "content": "What's the weather in Berlin?"}],
"model": "my_workspace/3fa85f64-5717-4562-b3fc-2c963f66afa6",
"temperature": 1.0,
"max_completion_tokens": 512
}

The Agent receives those values when the revision YAML inputs mapping names generation_kwargs:

inputs:
messages:
- agent.messages
generation_kwargs:
- agent.generation_kwargs

generation_kwargs is merged into the agent's chat generator at run time.

Tools​

Client-side OpenAI-style tools and tool_choice are not supported. If the deployed service uses an Agent with its own configured tools, those pipeline tools still run inside the service.

File references​

A message can reference a file already stored in this workspace via a file content part: {"type": "file", "file": {"file_id": "<uuid>"}}. file_id must be a UUID of a file in this workspace, not an OpenAI file id. You can also send inline file_data (a base64 data URL) or an image_url part with a base64 data URL. input_audio parts and HTTP image_urls are ignored; they do not fail the request.

Request​

Responses​

Streaming response from the deployed service (SSE) when stream is true, or a single JSON chat.completion object when stream is false (default).