TavilyFetcher
Fetch and extract content from URLs as Haystack Documents using the Tavily Extract API.
Key Features
- Fetches and extracts content from specific URLs, including PDF files.
- Supports basic and advanced extraction depths for trade-offs between speed, cost, and data richness.
- Returns request-level metadata including response time, usage, and details about failed URLs.
- Optionally includes extracted image URLs in Document metadata.
- Supports both synchronous and asynchronous execution.
- You need a Tavily API key from tavily.com.
Configuration
- Sign up at tavily.com and get an API key.
- Set the
TAVILY_API_KEYenvironment variable. For instructions, see Create Secrets. - Configure
TavilyFetcherin your pipeline YAML with the desiredextract_depth.
Connections
TavilyFetcher takes a list of URL strings as urls input and outputs a list of Document objects (documents) and a metadata dictionary (meta). The meta output includes response_time, usage, request_id, and failed_results.
Connect the pipeline's URL input to the urls input. Connect the documents output to a PromptBuilder, DocumentSplitter, or another downstream component.
Source Code
To check this component's source code, open tavily_fetcher.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
tavily_fetcher:
type: haystack_integrations.components.fetchers.tavily.tavily_fetcher.TavilyFetcher
init_parameters:
api_key:
type: env_var
env_vars:
- TAVILY_API_KEY
strict: false
extract_depth: basic
include_images: false
Using the Component in a Pipeline
# haystack-pipeline
components:
tavily_fetcher:
type: haystack_integrations.components.fetchers.tavily.tavily_fetcher.TavilyFetcher
init_parameters:
api_key:
type: env_var
env_vars:
- TAVILY_API_KEY
strict: false
extract_depth: basic
prompt_builder:
type: haystack.components.builders.prompt_builder.PromptBuilder
init_parameters:
template: |
Given the following fetched content:
{% for doc in documents %}{{ doc.content }}{% endfor %}
Answer: {{ query }}
llm:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini
connections:
- sender: tavily_fetcher.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: llm.messages
inputs:
urls:
- tavily_fetcher.urls
query:
- prompt_builder.query
outputs:
replies: llm.replies
max_runs_per_component: 100
metadata: {}
Parameters
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | Secret | Secret.from_env_var("TAVILY_API_KEY") | Tavily API key. |
extract_depth | str | basic | Extraction depth. basic is faster and lower cost; advanced returns more data including tables but has higher latency and cost. |
include_images | bool | False | If True, includes extracted image URLs in each Document's metadata under the images key. |
extract_params | dict | None | None | Additional parameters passed to the Tavily Extract API. Supported keys include format, include_favicon, query, and chunks_per_source. See the Tavily Extract API reference for details. |
Related Information
Was this page helpful?