Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ChineseDocumentSplitter

Split Chinese text Documents into chunks using the HanLP NLP library, which understands Chinese word boundaries.

Key Features​

  • Splits Chinese text by word, sentence, passage, page, line, period, or a custom splitting function.
  • Uses HanLP for accurate Chinese word segmentation — essential since Chinese text has no spaces between words.
  • Supports coarse and fine granularity word segmentation.
  • Optionally respects sentence boundaries when splitting by word count.
  • Preserves metadata such as page number, split index, and overlap information across splits.
  • Loads HanLP models on first use and caches them locally.

Configuration​

  1. Install the hanlp-haystack integration.
  2. Configure ChineseDocumentSplitter in your pipeline YAML with the desired split settings.
HanLP Model Download

ChineseDocumentSplitter downloads HanLP models on first use. The models are cached locally after the initial download. Make sure your environment has internet access when the pipeline first runs.

Connections​

ChineseDocumentSplitter takes a list of Document objects as documents input and outputs a list of split Document objects as documents.

Connect a document converter's documents output to the documents input. Connect the documents output to a document writer or embedding component.

Source Code​

To check this component's source code, open chinese_document_splitter.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

chinese_splitter:
type: haystack_integrations.components.preprocessors.hanlp.chinese_document_splitter.ChineseDocumentSplitter
init_parameters:
split_by: word
split_length: 1000
split_overlap: 200
granularity: coarse

Using the Component in a Pipeline​

# haystack-pipeline
components:
converter:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters: {}
chinese_splitter:
type: haystack_integrations.components.preprocessors.hanlp.chinese_document_splitter.ChineseDocumentSplitter
init_parameters:
split_by: word
split_length: 1000
split_overlap: 200
respect_sentence_boundary: true
granularity: coarse
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: chinese-docs
embedding_dim: 768
create_index: true

connections:
- sender: converter.documents
receiver: chinese_splitter.documents
- sender: chinese_splitter.documents
receiver: writer.documents

inputs:
files:
- converter.sources

max_runs_per_component: 100

metadata: {}

Parameters​

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
split_bystrwordThe unit for splitting. One of word, sentence, passage, page, line, period, or function.
split_lengthint1000The maximum number of units in each split. Must be positive.
split_overlapint200The number of overlapping units between splits. Must be less than split_length.
split_thresholdint0The minimum number of units per split. Splits below this threshold are merged into the previous split.
respect_sentence_boundaryboolFalseIf True and split_by is word, uses HanLP sentence detection so splits occur only between sentences.
splitting_functionCallable | NoneNoneCustom splitting function required when split_by is function. Accepts a str and returns a list[str].
granularitystrcoarseChinese word segmentation granularity. Either coarse or fine.