ChineseDocumentSplitter
Split Chinese text Documents into chunks using the HanLP NLP library, which understands Chinese word boundaries.
Key Features
- Splits Chinese text by word, sentence, passage, page, line, period, or a custom splitting function.
- Uses HanLP for accurate Chinese word segmentation — essential since Chinese text has no spaces between words.
- Supports coarse and fine granularity word segmentation.
- Optionally respects sentence boundaries when splitting by word count.
- Preserves metadata such as page number, split index, and overlap information across splits.
- Loads HanLP models on first use and caches them locally.
Configuration
- Install the
hanlp-haystackintegration. - Configure
ChineseDocumentSplitterin your pipeline YAML with the desired split settings.
ChineseDocumentSplitter downloads HanLP models on first use. The models are cached locally after the initial download. Make sure your environment has internet access when the pipeline first runs.
Connections
ChineseDocumentSplitter takes a list of Document objects as documents input and outputs a list of split Document objects as documents.
Connect a document converter's documents output to the documents input. Connect the documents output to a document writer or embedding component.
Source Code
To check this component's source code, open chinese_document_splitter.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
chinese_splitter:
type: haystack_integrations.components.preprocessors.hanlp.chinese_document_splitter.ChineseDocumentSplitter
init_parameters:
split_by: word
split_length: 1000
split_overlap: 200
granularity: coarse
Using the Component in a Pipeline
# haystack-pipeline
components:
converter:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters: {}
chinese_splitter:
type: haystack_integrations.components.preprocessors.hanlp.chinese_document_splitter.ChineseDocumentSplitter
init_parameters:
split_by: word
split_length: 1000
split_overlap: 200
respect_sentence_boundary: true
granularity: coarse
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: chinese-docs
embedding_dim: 768
create_index: true
connections:
- sender: converter.documents
receiver: chinese_splitter.documents
- sender: chinese_splitter.documents
receiver: writer.documents
inputs:
files:
- converter.sources
max_runs_per_component: 100
metadata: {}
Parameters
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
split_by | str | word | The unit for splitting. One of word, sentence, passage, page, line, period, or function. |
split_length | int | 1000 | The maximum number of units in each split. Must be positive. |
split_overlap | int | 200 | The number of overlapping units between splits. Must be less than split_length. |
split_threshold | int | 0 | The minimum number of units per split. Splits below this threshold are merged into the previous split. |
respect_sentence_boundary | bool | False | If True and split_by is word, uses HanLP sentence detection so splits occur only between sentences. |
splitting_function | Callable | None | None | Custom splitting function required when split_by is function. Accepts a str and returns a list[str]. |
granularity | str | coarse | Chinese word segmentation granularity. Either coarse or fine. |
Related Information
Was this page helpful?