Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

EmbeddingBasedDocumentSplitter

Split documents where the topic changes. The component groups sentences, embeds each group, and starts a new chunk when the embedding distance between neighbors is high. Use it in indexes when fixed-size chunks cut across topics.

Key Features​

  • Splits text at semantic boundaries instead of a fixed word or character count.
  • Groups a configurable number of sentences before embedding them.
  • Treats cosine distances above a percentile threshold as split points.
  • Merges chunks shorter than a minimum length and splits chunks longer than a maximum length.
  • Adds source_id, split_id, split_idx_start, and page_number metadata to each chunk. Page numbers come from form feed characters in the original text.

Configuration​

  1. Drag the EmbeddingBasedDocumentSplitter component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab, set Document Embedder to a document embedder, such as SentenceTransformersDocumentEmbedder. The splitter uses this embedder to compare sentence groups.
  4. Go to the Advanced tab:
    • Set Sentences Per Group to the number of sentences embedded together. The default is three.
    • Set Percentile to the cosine-distance cutoff. Distances above this percentile become split points. The default is 0.95.
    • Set Min Length and Max Length in characters. Shorter chunks are merged. Longer chunks are split again.
    • Set Language for sentence tokenization. The default is en.

Connections​

EmbeddingBasedDocumentSplitter accepts a list of Document objects and outputs a list of split Document objects.

It typically receives documents from converters or DocumentCleaner, and sends split documents to a document writer. The embedder you set on this component is used only to find split points. If you also want the stored chunks to carry embeddings, add a separate document embedder after the splitter.

Source Code​

To check this component's source code, open embedding_based_document_splitter.py in the Haystack repository.

Usage Examples​

Basic Configuration​

EmbeddingBasedDocumentSplitter:
type: haystack.components.preprocessors.embedding_based_document_splitter.EmbeddingBasedDocumentSplitter
init_parameters:
document_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.sentence_transformers_document_embedder.SentenceTransformersDocumentEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
sentences_per_group: 3
percentile: 0.95
min_length: 50
max_length: 1000
language: en

Using the Component in an Index​

This index splits cleaned documents at topic changes, then writes the chunks to a document store.

# haystack-pipeline
components:
TextFileToDocument:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters:
encoding: utf-8
store_full_path: false
DocumentCleaner:
type: haystack.components.preprocessors.document_cleaner.DocumentCleaner
init_parameters:
remove_empty_lines: true
remove_extra_whitespaces: true
EmbeddingBasedDocumentSplitter:
type: haystack.components.preprocessors.embedding_based_document_splitter.EmbeddingBasedDocumentSplitter
init_parameters:
document_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.sentence_transformers_document_embedder.SentenceTransformersDocumentEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
sentences_per_group: 3
percentile: 0.95
min_length: 50
max_length: 1000
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: semantic-chunks
embedding_dim: 384
create_index: true
similarity: cosine
policy: NONE

connections:
- sender: TextFileToDocument.documents
receiver: DocumentCleaner.documents
- sender: DocumentCleaner.documents
receiver: EmbeddingBasedDocumentSplitter.documents
- sender: EmbeddingBasedDocumentSplitter.documents
receiver: DocumentWriter.documents

max_runs_per_component: 100

inputs:
files:
- TextFileToDocument.sources

metadata: {}

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]List of text documents to split. Each document needs content.

Outputs​

ParameterTypeDescription
documentsList[Document]List of split documents. Each chunk includes source_id, split_id, split_idx_start, and page_number metadata, plus the original document metadata.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_embedderDocumentEmbedderThe document embedder used to compare sentence groups.
sentences_per_groupint3Number of sentences to group together before embedding. Must be greater than zero.
percentilefloat0.95Cosine-distance percentile used as the split threshold. Distances above this percentile are split points. Must be between 0.0 and 1.0.
min_lengthint50Minimum chunk length in characters. Shorter chunks are merged with a neighbor.
max_lengthint1000Maximum chunk length in characters. Longer chunks are split again. Must be greater than min_length.
languagestrenLanguage used for sentence tokenization.
use_split_rulesboolTrueWhether to apply extra sentence-splitting rules from the sentence splitter.
extend_abbreviationsboolTrueWhether to extend NLTK abbreviations with a curated list. Supported languages are English (en) and German (de).

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
documentsList[Document]List of text documents to split.