Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ParallelChatGenerator

Generate chat responses grounded in live web research using the Parallel Responses API.

Key Features​

  • Uses the Parallel Responses API (POST /v1/responses), which is compatible with the OpenAI Responses format.
  • Grounds every answer in live web research with citations included automatically.
  • Supports three research tiers via the reasoning.effort parameter: low (~5–10s), medium (~15–20s, default), and high (~30–60s).
  • Accepts and returns messages in ChatMessage format.
  • Supports streaming responses through a configurable callback.
  • Defaults to a 120-second timeout to accommodate the high research tier.

Configuration​

  1. Drag the ParallelChatGenerator component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Create a secret with your Parallel API key and set it as api_key. Use PARALLEL_API_KEY as the environment variable name. For instructions, see Create Secrets.
    2. Optionally configure the research tier by setting generation_kwargs with reasoning: {effort: low}, reasoning: {effort: medium}, or reasoning: {effort: high}.
  4. Go to the Advanced tab to configure timeout and max_retries.

Connections​

ParallelChatGenerator receives a list of ChatMessage objects, typically from PromptBuilder or ChatPromptBuilder. It outputs a list of reply ChatMessage objects you can connect to AnswerBuilder or other downstream components.

Source Code​

To check this component's source code, open chat_generator.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

ParallelChatGenerator:
type: haystack_integrations.components.generators.parallel.chat.chat_generator.ParallelChatGenerator
init_parameters:
api_key:
type: env_var
env_vars:
- PARALLEL_API_KEY
strict: false
model: parallel
generation_kwargs:
reasoning:
effort: medium

Using the Component in a Pipeline​

# haystack-pipeline
components:
prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
required_variables: "*"
template:
- role: user
content: "Answer the following question with up-to-date information: {{ question }}"

llm:
type: haystack_integrations.components.generators.parallel.chat.chat_generator.ParallelChatGenerator
init_parameters:
api_key:
type: env_var
env_vars:
- PARALLEL_API_KEY
strict: false
model: parallel
generation_kwargs:
reasoning:
effort: medium

answer_builder:
type: deepset_cloud_custom_nodes.augmenters.deepset_answer_builder.DeepsetAnswerBuilder
init_parameters:
reference_pattern: acm

connections:
- sender: prompt_builder.prompt
receiver: llm.messages
- sender: llm.replies
receiver: answer_builder.replies

max_runs_per_component: 100

metadata: {}

inputs:
query:
- answer_builder.query
- prompt_builder.question

outputs:
answers: answer_builder.answers

Parameters​

Inputs​

ParameterTypeDescription
messagesList[ChatMessage]A list of chat messages representing the conversation so far.

Outputs​

ParameterTypeDescription
repliesList[ChatMessage]A list of generated reply messages grounded in live web research.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_keySecretSecret.from_env_var("PARALLEL_API_KEY")The Parallel API key.
modelstrparallelThe Parallel Responses API model. Currently only parallel is supported.
api_base_urlOptional[str]https://api.parallel.ai/v1The Parallel API base URL.
streaming_callbackOptional[Callable]NoneA callback function for streaming responses.
generation_kwargsOptional[Dict[str, Any]]NoneAdditional parameters sent to the Parallel Responses API, such as reasoning (for example, {"effort": "low"}) to select the research tier, or text for structured output. Note: tools, temperature, top_p, and similar parameters are accepted for SDK compatibility but silently ignored.
timeoutOptional[float]120.0Timeout in seconds. Defaults to 120 to accommodate the high research tier.
max_retriesOptional[int]3Maximum number of retries. Kept low because every retry runs a full research call.
extra_headersOptional[Dict[str, Any]]NoneAdditional HTTP headers to include in requests.
http_client_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments for the underlying httpx.Client or httpx.AsyncClient.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
messagesList[ChatMessage]A list of chat messages representing the conversation.
streaming_callbackOptional[Callable]NoneA callback function to override the init-time streaming callback.
generation_kwargsOptional[Dict[str, Any]]NoneAdditional generation parameters to override init-time values.