Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

HetznerChatGenerator

Generate chat responses using models served by the Hetzner Inference API.

Key Features​

  • Connects to the Hetzner Inference API, which is compatible with the OpenAI Chat Completions format.
  • Supports open-weight models such as Qwen/Qwen3.6-35B-A3B-FP8.
  • Accepts multimodal inputs, including text and images in ChatMessage objects.
  • Supports streaming responses through a configurable callback.
  • Supports tool calling and structured output via JSON schema or Pydantic models.
  • Configurable structured output (response_format) for enforcing a specific response shape.

Configuration​

  1. Drag the HetznerChatGenerator component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set the model to the Hetzner model you want to use (for example, Qwen/Qwen3.6-35B-A3B-FP8). For the current model list, query the /v1/models endpoint of the Hetzner Inference API.
    2. Create a secret with your Hetzner API token and set it as api_key. Use HETZNER_API_KEY as the environment variable name. For instructions, see Create Secrets.
  4. Go to the Advanced tab to configure generation_kwargs, timeout, and max_retries.

Connections​

HetznerChatGenerator receives a list of ChatMessage objects, typically from PromptBuilder or ChatPromptBuilder. It outputs a list of reply ChatMessage objects you can connect to AnswerBuilder or other downstream components.

Source Code​

To check this component's source code, open chat_generator.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

HetznerChatGenerator:
type: haystack_integrations.components.generators.hetzner.chat.chat_generator.HetznerChatGenerator
init_parameters:
api_key:
type: env_var
env_vars:
- HETZNER_API_KEY
strict: false
model: Qwen/Qwen3.6-35B-A3B-FP8
generation_kwargs:
max_tokens: 1024
temperature: 0.7

Using the Component in a Pipeline​

# haystack-pipeline
components:
prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
required_variables: "*"
template:
- role: user
content: "Answer the following question: {{ question }}"

llm:
type: haystack_integrations.components.generators.hetzner.chat.chat_generator.HetznerChatGenerator
init_parameters:
api_key:
type: env_var
env_vars:
- HETZNER_API_KEY
strict: false
model: Qwen/Qwen3.6-35B-A3B-FP8
generation_kwargs:
max_tokens: 1024

answer_builder:
type: deepset_cloud_custom_nodes.augmenters.deepset_answer_builder.DeepsetAnswerBuilder
init_parameters:
reference_pattern: acm

connections:
- sender: prompt_builder.prompt
receiver: llm.messages
- sender: llm.replies
receiver: answer_builder.replies

max_runs_per_component: 100

metadata: {}

inputs:
query:
- answer_builder.query
- prompt_builder.question

outputs:
answers: answer_builder.answers

Parameters​

Inputs​

ParameterTypeDescription
messagesList[ChatMessage]A list of chat messages representing the conversation so far.

Outputs​

ParameterTypeDescription
repliesList[ChatMessage]A list of generated reply messages from the model.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_keySecretSecret.from_env_var("HETZNER_API_KEY")The Hetzner Inference API token.
modelstrQwen/Qwen3.6-35B-A3B-FP8The name of the model to use. Query /v1/models for the current list.
streaming_callbackOptional[Callable]NoneA callback function for streaming responses.
api_base_urlOptional[str]https://inference.hetzner.com/api/v1The Hetzner Inference API base URL.
generation_kwargsOptional[Dict[str, Any]]NoneAdditional generation parameters such as max_tokens, temperature, top_p, and response_format.
toolsOptional[ToolsType]NoneA list of tools the model can use.
timeoutOptional[float]NoneRequest timeout in seconds.
max_retriesOptional[int]NoneMaximum number of retries on API errors. Defaults to the OPENAI_MAX_RETRIES environment variable or five.
http_client_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments for the underlying httpx.Client or httpx.AsyncClient.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
messagesList[ChatMessage]A list of chat messages representing the conversation.
streaming_callbackOptional[Callable]NoneA callback function to override the init-time streaming callback.
generation_kwargsOptional[Dict[str, Any]]NoneAdditional generation parameters to override init-time values.
toolsOptional[List[Tool]]NoneTools to make available to the model.