Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

TavilyFetcher

Fetch and extract content from URLs as Haystack Documents using the Tavily Extract API.

Key Features​

  • Fetches and extracts content from specific URLs, including PDF files.
  • Supports basic and advanced extraction depths for trade-offs between speed, cost, and data richness.
  • Returns request-level metadata including response time, usage, and details about failed URLs.
  • Optionally includes extracted image URLs in Document metadata.
  • Supports both synchronous and asynchronous execution.
  • You need a Tavily API key from tavily.com.

Configuration​

  1. Sign up at tavily.com and get an API key.
  2. Set the TAVILY_API_KEY environment variable. For instructions, see Create Secrets.
  3. Configure TavilyFetcher in your pipeline YAML with the desired extract_depth.

Connections​

TavilyFetcher takes a list of URL strings as urls input and outputs a list of Document objects (documents) and a metadata dictionary (meta). The meta output includes response_time, usage, request_id, and failed_results.

Connect the pipeline's URL input to the urls input. Connect the documents output to a PromptBuilder, DocumentSplitter, or another downstream component.

Source Code​

To check this component's source code, open tavily_fetcher.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

tavily_fetcher:
type: haystack_integrations.components.fetchers.tavily.tavily_fetcher.TavilyFetcher
init_parameters:
api_key:
type: env_var
env_vars:
- TAVILY_API_KEY
strict: false
extract_depth: basic
include_images: false

Using the Component in a Pipeline​

# haystack-pipeline
components:
tavily_fetcher:
type: haystack_integrations.components.fetchers.tavily.tavily_fetcher.TavilyFetcher
init_parameters:
api_key:
type: env_var
env_vars:
- TAVILY_API_KEY
strict: false
extract_depth: basic
prompt_builder:
type: haystack.components.builders.prompt_builder.PromptBuilder
init_parameters:
template: |
Given the following fetched content:
{% for doc in documents %}{{ doc.content }}{% endfor %}
Answer: {{ query }}
llm:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini

connections:
- sender: tavily_fetcher.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: llm.messages

inputs:
urls:
- tavily_fetcher.urls
query:
- prompt_builder.query

outputs:
replies: llm.replies

max_runs_per_component: 100

metadata: {}

Parameters​

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_keySecretSecret.from_env_var("TAVILY_API_KEY")Tavily API key.
extract_depthstrbasicExtraction depth. basic is faster and lower cost; advanced returns more data including tables but has higher latency and cost.
include_imagesboolFalseIf True, includes extracted image URLs in each Document's metadata under the images key.
extract_paramsdict | NoneNoneAdditional parameters passed to the Tavily Extract API. Supported keys include format, include_favicon, query, and chunks_per_source. See the Tavily Extract API reference for details.