EmbeddingBasedDocumentSplitter
Split documents where the topic changes. The component groups sentences, embeds each group, and starts a new chunk when the embedding distance between neighbors is high. Use it in indexes when fixed-size chunks cut across topics.
Key Features
- Splits text at semantic boundaries instead of a fixed word or character count.
- Groups a configurable number of sentences before embedding them.
- Treats cosine distances above a percentile threshold as split points.
- Merges chunks shorter than a minimum length and splits chunks longer than a maximum length.
- Adds
source_id,split_id,split_idx_start, andpage_numbermetadata to each chunk. Page numbers come from form feed characters in the original text.
Configuration
- Drag the
EmbeddingBasedDocumentSplittercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab, set Document Embedder to a document embedder, such as
SentenceTransformersDocumentEmbedder. The splitter uses this embedder to compare sentence groups. - Go to the Advanced tab:
- Set Sentences Per Group to the number of sentences embedded together. The default is three.
- Set Percentile to the cosine-distance cutoff. Distances above this percentile become split points. The default is
0.95. - Set Min Length and Max Length in characters. Shorter chunks are merged. Longer chunks are split again.
- Set Language for sentence tokenization. The default is
en.
Connections
EmbeddingBasedDocumentSplitter accepts a list of Document objects and outputs a list of split Document objects.
It typically receives documents from converters or DocumentCleaner, and sends split documents to a document writer. The embedder you set on this component is used only to find split points. If you also want the stored chunks to carry embeddings, add a separate document embedder after the splitter.
Source Code
To check this component's source code, open embedding_based_document_splitter.py in the Haystack repository.
Usage Examples
Basic Configuration
EmbeddingBasedDocumentSplitter:
type: haystack.components.preprocessors.embedding_based_document_splitter.EmbeddingBasedDocumentSplitter
init_parameters:
document_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.sentence_transformers_document_embedder.SentenceTransformersDocumentEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
sentences_per_group: 3
percentile: 0.95
min_length: 50
max_length: 1000
language: en
Using the Component in an Index
This index splits cleaned documents at topic changes, then writes the chunks to a document store.
# haystack-pipeline
components:
TextFileToDocument:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters:
encoding: utf-8
store_full_path: false
DocumentCleaner:
type: haystack.components.preprocessors.document_cleaner.DocumentCleaner
init_parameters:
remove_empty_lines: true
remove_extra_whitespaces: true
EmbeddingBasedDocumentSplitter:
type: haystack.components.preprocessors.embedding_based_document_splitter.EmbeddingBasedDocumentSplitter
init_parameters:
document_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.sentence_transformers_document_embedder.SentenceTransformersDocumentEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
sentences_per_group: 3
percentile: 0.95
min_length: 50
max_length: 1000
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: semantic-chunks
embedding_dim: 384
create_index: true
similarity: cosine
policy: NONE
connections:
- sender: TextFileToDocument.documents
receiver: DocumentCleaner.documents
- sender: DocumentCleaner.documents
receiver: EmbeddingBasedDocumentSplitter.documents
- sender: EmbeddingBasedDocumentSplitter.documents
receiver: DocumentWriter.documents
max_runs_per_component: 100
inputs:
files:
- TextFileToDocument.sources
metadata: {}
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | List of text documents to split. Each document needs content. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | List of split documents. Each chunk includes source_id, split_id, split_idx_start, and page_number metadata, plus the original document metadata. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_embedder | DocumentEmbedder | The document embedder used to compare sentence groups. | |
sentences_per_group | int | 3 | Number of sentences to group together before embedding. Must be greater than zero. |
percentile | float | 0.95 | Cosine-distance percentile used as the split threshold. Distances above this percentile are split points. Must be between 0.0 and 1.0. |
min_length | int | 50 | Minimum chunk length in characters. Shorter chunks are merged with a neighbor. |
max_length | int | 1000 | Maximum chunk length in characters. Longer chunks are split again. Must be greater than min_length. |
language | str | en | Language used for sentence tokenization. |
use_split_rules | bool | True | Whether to apply extra sentence-splitting rules from the sentence splitter. |
extend_abbreviations | bool | True | Whether to extend NLTK abbreviations with a curated list. Supported languages are English (en) and German (de). |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | List of text documents to split. |
Related Information
Was this page helpful?