ContentGate
Control the flow of documents in a deep-research agent pipeline by forwarding documents to the summarizer only when the fetched page contains readable text.
Key Features
- Acts as a gating component: passes documents downstream only when they contain actual text content.
- Returns an
emptysignal when the fetched page has no readable content (for example, unsupported file types or blank pages), allowing the pipeline to handle the case gracefully. - Designed to work with
LinkContentFetcherandFileTypeRouterin theread_urltool used by deep-research agents. - Requires no credentials or configuration—connects directly to document lists from converters.
- Prevents empty documents from reaching the LLM summarizer, reducing unnecessary API calls.
Configuration
ContentGate requires no credentials or environment variables.
- Drag the
ContentGatecomponent onto the canvas from the Component Library. - Connect its
documentsinput to the output of aDocumentJoineror document converter. - Connect its
documentsoutput to aChatPromptBuilderand itsemptyoutput to a fallback branch.
Connections
ContentGate receives a documents list from an upstream converter or joiner. It outputs either documents (when the first document has non-empty content) or empty: True (when the page has no readable text). In the read_url pipeline tool, the documents output connects to a ChatPromptBuilder for summarization.
Source Code
To check this component's source code, open tools.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
content_gate:
type: haystack_integrations.agent_pack.deep_research.tools.ContentGate
init_parameters: {}
Using the Component in a Pipeline
This example shows how ContentGate fits into a fetch-convert-summarize pipeline used by the deep-research agent's read_url tool.
# haystack-pipeline
components:
fetcher:
type: haystack.components.fetchers.link_content.LinkContentFetcher
init_parameters:
raise_on_failure: true
retry_attempts: 2
timeout: 10
router:
type: haystack.components.routers.file_type_router.FileTypeRouter
init_parameters:
mime_types:
- text/html
- application/pdf
html:
type: haystack.components.converters.html.HTMLToDocument
init_parameters: {}
pdf:
type: haystack.components.converters.pypdf.PyPDFToDocument
init_parameters: {}
joiner:
type: haystack.components.joiners.document_joiner.DocumentJoiner
init_parameters: {}
content_gate:
type: haystack_integrations.agent_pack.deep_research.tools.ContentGate
init_parameters: {}
prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
template:
- role: user
content: "Summarize the following page for the question: {{ question }}\n\n{% for doc in documents %}{{ doc.content }}\n{% endfor %}"
required_variables:
- question
- documents
summarizer:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini
connections:
- sender: fetcher.streams
receiver: router.sources
- sender: router.text/html
receiver: html.sources
- sender: router.application/pdf
receiver: pdf.sources
- sender: html.documents
receiver: joiner.documents
- sender: pdf.documents
receiver: joiner.documents
- sender: joiner.documents
receiver: content_gate.documents
- sender: content_gate.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: summarizer.messages
inputs:
query:
- fetcher.urls
outputs:
replies: summarizer.replies
max_runs_per_component: 100
metadata: {}
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Documents produced by an upstream converter or joiner from the fetched page. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The original documents, forwarded when the first document has non-empty text content. |
empty | bool | True when the page has no readable text content. |
Init Parameters
This component has no init parameters.
Related Information
Was this page helpful?