Chat Completions
POST/api/v1/workspaces/:workspace_name/deployments/v1/chat/completions
Run a deployment's active revision as a chat completion (OpenAI-style).
The chat completion endpoint sends a list of conversation messages to an AI model and returns the next assistant response. For multi-turn chats, clients maintain the history by including prior user and assistant messages in each request, along with the new user message. The run is stored in search history so you can inspect traces.
By default (stream: false), a single JSON response is returned once generation completes.
Pass stream: true to instead receive the response streamed as SSE.
model identifies which deployed service to run, as {workspace_name}/{deployment_id}.
Example request:
{
"messages": [
{"role": "user", "content": "How does streaming work with Haystack Enterprise Platform?"}
],
"model": "my_workspace/3fa85f64-5717-4562-b3fc-2c963f66afa6"
}
Example non-streaming (stream: false, the default) response:
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1730000000,
"model": "my_workspace/3fa85f64-5717-4562-b3fc-2c963f66afa6",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Streaming sends the response as SSE chunks."},
"finish_reason": "stop"
}
]
}
Extra OpenAI parameters
You can send extra OpenAI fields besides messages, model, and stream, for example
temperature, top_p, and other generation parameters. They are placed in a
generation_kwargs dict and passed to the deployed service as a named input.
OpenAI-style tools and tool_choice are not supported and are ignored from the request.
Example request:
{
"messages": [{"role": "user", "content": "What's the weather in Berlin?"}],
"model": "my_workspace/3fa85f64-5717-4562-b3fc-2c963f66afa6",
"temperature": 1.0,
"max_completion_tokens": 512
}
The Agent receives those values when the revision YAML inputs mapping names
generation_kwargs:
inputs:
messages:
- agent.messages
generation_kwargs:
- agent.generation_kwargs
generation_kwargs is merged into the agent's chat generator at run time.
Tools
Client-side OpenAI-style tools and tool_choice are not supported. If the deployed service
uses an Agent with its own configured tools, those pipeline tools still run inside the service.
File references
A message can reference a file already stored in this workspace via a file content part:
{"type": "file", "file": {"file_id": "<uuid>"}}. file_id must be a UUID of a file in this
workspace, not an OpenAI file id. You can also send inline file_data (a base64 data URL)
or an image_url part with a base64 data URL. input_audio parts and HTTP image_urls
are ignored; they do not fail the request.
Request
Responses
- 200
- 400
- 404
- 409
- 422
- 503
Streaming response from the deployed service (SSE) when stream is true, or a single JSON chat.completion object when stream is false (default).
No user message in messages, the deployed service has no query input, the 'model' field's workspace doesn't match the URL workspace, or a file content part's file_id is not a UUID.
We couldn't find the deployed service. Check if the deployment ID is correct, make sure the deployed service exists in the workspace you specified, and try again.
The deployed service was in a conflicted state (for example, it wasn't deployed).
Validation Error
The deployed service is unavailable or the run failed before a completion was produced. Retry the request.
Was this page helpful?