Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

MarkdownHeaderSplitter

Split Markdown documents at headings that start with #. Each chunk keeps the heading path that led to it, so retrieval can tell which section a passage came from. Use this component in indexes built from Markdown files.

Key Features​

  • Splits on Markdown headings and stores the heading hierarchy in metadata.
  • Lets you choose which heading levels create a new chunk. Deeper headings stay inside the parent chunk.
  • Ignores heading-like lines inside fenced code blocks.
  • Optionally runs a second split by word, passage, period, or line when a section is still too long.
  • Tracks page numbers from a page-break character, form feed by default.

Configuration​

  1. Drag the MarkdownHeaderSplitter component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    • Set Header Split Levels to the heading levels that start a new chunk. For example, [1, 2] splits on # and ## only. The default is all six levels.
    • Toggle Keep Headers to leave headings in the chunk text. Turn it off to move headings into metadata instead.
  4. Go to the Advanced tab when a section should be split again:
    • Set Secondary Split to word, passage, period, or line. Leave it empty to keep one chunk per heading section.
    • Set Split Length, Split Overlap, and Split Threshold for that second split.

Connections​

MarkdownHeaderSplitter accepts a list of Document objects and outputs a list of split Document objects. Heading metadata includes the parent headings for the chunk.

It typically receives documents from MarkdownToDocument or DocumentCleaner, and sends split documents to an embedder or DocumentWriter.

Source Code​

To check this component's source code, open markdown_header_splitter.py in the Haystack repository.

Usage Examples​

Basic Configuration​

MarkdownHeaderSplitter:
type: haystack.components.preprocessors.markdown_header_splitter.MarkdownHeaderSplitter
init_parameters:
keep_headers: true
header_split_levels:
- 1
- 2
secondary_split: word
split_length: 200
split_overlap: 0
split_threshold: 0
skip_empty_documents: true

Using the Component in an Index​

This index converts Markdown files, splits them at first- and second-level headings, then writes the chunks.

# haystack-pipeline
components:
MarkdownToDocument:
type: haystack.components.converters.markdown.MarkdownToDocument
init_parameters:
table_to_single_line: false
progress_bar: true
store_full_path: false
MarkdownHeaderSplitter:
type: haystack.components.preprocessors.markdown_header_splitter.MarkdownHeaderSplitter
init_parameters:
keep_headers: true
header_split_levels:
- 1
- 2
secondary_split: word
split_length: 200
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: markdown-sections
create_index: true
policy: NONE

connections:
- sender: MarkdownToDocument.documents
receiver: MarkdownHeaderSplitter.documents
- sender: MarkdownHeaderSplitter.documents
receiver: DocumentWriter.documents

max_runs_per_component: 100

inputs:
files:
- MarkdownToDocument.sources

metadata: {}

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]List of Markdown documents to split.

Outputs​

ParameterTypeDescription
documentsList[Document]List of documents split at headings. Each chunk carries heading metadata, a split ID, and a page number.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
page_break_characterstr\fCharacter that marks a page break. The default is a form feed.
keep_headersboolTrueWhen True, headings stay in the chunk text. When False, headings are stored in metadata.
header_split_levelsOptional[List[int]]NoneHeading levels from 1 to 6 that start a new chunk. When this is empty, the component splits on all six levels. Levels must be unique.
secondary_splitOptional[Literal['word', 'passage', 'period', 'line']]NoneOptional second split applied to each heading section.
split_lengthint200Maximum number of units in each chunk when a second split is set.
split_overlapint0Number of overlapping units between chunks when a second split is set.
split_thresholdint0Minimum number of units per chunk when a second split is set. Smaller chunks are merged with the previous chunk.
skip_empty_documentsboolTrueWhether to skip documents with empty content. Set this to False when a later component, such as LLMDocumentContentExtractor, can read text from non-text files.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
documentsList[Document]List of Markdown documents to split.