Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ContentGate

Control the flow of documents in a deep-research agent pipeline by forwarding documents to the summarizer only when the fetched page contains readable text.

Key Features​

  • Acts as a gating component: passes documents downstream only when they contain actual text content.
  • Returns an empty signal when the fetched page has no readable content (for example, unsupported file types or blank pages), allowing the pipeline to handle the case gracefully.
  • Designed to work with LinkContentFetcher and FileTypeRouter in the read_url tool used by deep-research agents.
  • Requires no credentials or configuration—connects directly to document lists from converters.
  • Prevents empty documents from reaching the LLM summarizer, reducing unnecessary API calls.

Configuration​

ContentGate requires no credentials or environment variables.

  1. Drag the ContentGate component onto the canvas from the Component Library.
  2. Connect its documents input to the output of a DocumentJoiner or document converter.
  3. Connect its documents output to a ChatPromptBuilder and its empty output to a fallback branch.

Connections​

ContentGate receives a documents list from an upstream converter or joiner. It outputs either documents (when the first document has non-empty content) or empty: True (when the page has no readable text). In the read_url pipeline tool, the documents output connects to a ChatPromptBuilder for summarization.

Source Code​

To check this component's source code, open tools.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

content_gate:
type: haystack_integrations.agent_pack.deep_research.tools.ContentGate
init_parameters: {}

Using the Component in a Pipeline​

This example shows how ContentGate fits into a fetch-convert-summarize pipeline used by the deep-research agent's read_url tool.

# haystack-pipeline

components:
fetcher:
type: haystack.components.fetchers.link_content.LinkContentFetcher
init_parameters:
raise_on_failure: true
retry_attempts: 2
timeout: 10

router:
type: haystack.components.routers.file_type_router.FileTypeRouter
init_parameters:
mime_types:
- text/html
- application/pdf

html:
type: haystack.components.converters.html.HTMLToDocument
init_parameters: {}

pdf:
type: haystack.components.converters.pypdf.PyPDFToDocument
init_parameters: {}

joiner:
type: haystack.components.joiners.document_joiner.DocumentJoiner
init_parameters: {}

content_gate:
type: haystack_integrations.agent_pack.deep_research.tools.ContentGate
init_parameters: {}

prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
template:
- role: user
content: "Summarize the following page for the question: {{ question }}\n\n{% for doc in documents %}{{ doc.content }}\n{% endfor %}"
required_variables:
- question
- documents

summarizer:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini

connections:
- sender: fetcher.streams
receiver: router.sources
- sender: router.text/html
receiver: html.sources
- sender: router.application/pdf
receiver: pdf.sources
- sender: html.documents
receiver: joiner.documents
- sender: pdf.documents
receiver: joiner.documents
- sender: joiner.documents
receiver: content_gate.documents
- sender: content_gate.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: summarizer.messages

inputs:
query:
- fetcher.urls

outputs:
replies: summarizer.replies

max_runs_per_component: 100

metadata: {}

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]Documents produced by an upstream converter or joiner from the fetched page.

Outputs​

ParameterTypeDescription
documentsList[Document]The original documents, forwarded when the first document has non-empty text content.
emptyboolTrue when the page has no readable text content.

Init Parameters​

This component has no init parameters.