Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

FirecrawlCrawler

Crawl websites and return content as Haystack Documents using the Firecrawl API.

Key Features​

  • Crawls entire websites or documentation sites by following links from a starting URL.
  • Returns page content as Markdown formatted for LLM consumption.
  • Configurable page limit to control the crawl depth and credit usage.
  • Supports both synchronous and asynchronous execution.
  • You need a Firecrawl API key from firecrawl.dev.

Configuration​

  1. Sign up at firecrawl.dev and get an API key.
  2. Set the FIRECRAWL_API_KEY environment variable. For instructions, see Create Secrets.
  3. Configure FirecrawlCrawler in your pipeline YAML with the desired crawl parameters, especially a limit to control how many pages are crawled.

Connections​

FirecrawlCrawler takes a list of URL strings as urls input and outputs a list of Document objects (documents), one per crawled page.

Connect the pipeline's URL input to the urls input. Connect the documents output to DocumentSplitter, DocumentWriter, or another processing component.

Source Code​

To check this component's source code, open firecrawl_crawler.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

firecrawl_crawler:
type: haystack_integrations.components.fetchers.firecrawl.firecrawl_crawler.FirecrawlCrawler
init_parameters:
api_key:
type: env_var
env_vars:
- FIRECRAWL_API_KEY
strict: false
params:
limit: 10
scrape_options:
formats:
- markdown

Using the Component in a Pipeline​

# haystack-pipeline
components:
firecrawl_crawler:
type: haystack_integrations.components.fetchers.firecrawl.firecrawl_crawler.FirecrawlCrawler
init_parameters:
api_key:
type: env_var
env_vars:
- FIRECRAWL_API_KEY
strict: false
params:
limit: 10
scrape_options:
formats:
- markdown
splitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: word
split_length: 250
split_overlap: 30
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: firecrawl-demo
embedding_dim: 768
create_index: true

connections:
- sender: firecrawl_crawler.documents
receiver: splitter.documents
- sender: splitter.documents
receiver: writer.documents

inputs:
urls:
- firecrawl_crawler.urls

max_runs_per_component: 100

metadata: {}

Parameters​

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_keySecretSecret.from_env_var("FIRECRAWL_API_KEY")Firecrawl API key.
paramsdict | NoneNoneParameters for the crawl request. Defaults to {"limit": 1, "scrape_options": {"formats": ["markdown"]}}. Set limit to control how many pages are crawled per URL. Without a limit, Firecrawl may crawl all subpages and consume credits quickly. See the Firecrawl crawl API reference for all available options.