Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

PythonCodeSplitter

Split Python source into chunks that follow the file's structure. Functions and methods stay whole, and each chunk reads from top to bottom like the original file. Use this component when you index code for search.

Key Features​

  • Parses each file with Python's syntax tree and keeps imports, functions, classes, and methods as units.
  • Merges those units toward a target size measured in effective lines. A long physical line counts as more than one effective line.
  • Leaves the primary split without overlap. Only an oversized function is split again, by line, and those pieces can overlap.
  • Can move function, method, and class docstrings out of the chunk text and into metadata.
  • Can prefix a chunk with the class signature when the chunk contains class members but not the class header.
  • Stores source_id, split_id, start_line, end_line, and unit_kinds on each chunk. Parent metadata, including file_name, is copied through.

The document content must be valid Python. The component raises a syntax error when the source does not parse.

Configuration​

  1. Drag the PythonCodeSplitter component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    • Set Min Effective Lines and Max Effective Lines. The splitter keeps adding the next unit while the chunk is below the minimum, and it aims for the maximum. Defaults are 20 and 100.
    • Set Expected Chars Per Line to decide how many characters count as one effective line. The default is 45.
  4. Go to the Advanced tab:
    • Set Oversized Factor. A function longer than this factor times the maximum is split by line. The default factor is three.
    • Toggle Strip Docstrings to move function, method, and class docstrings into meta.docstrings. The module docstring stays in the text.
    • Toggle Preserve Class Definition to repeat the class signature on chunks that contain members but not the class header.
    • Set Secondary Split Overlap and Secondary Split Length for the line split used on oversized functions.

Connections​

PythonCodeSplitter accepts a list of Document objects whose content is Python source. It outputs a list of split Document objects.

It typically receives documents from TextFileToDocument and sends chunks to an embedder or DocumentWriter. If you strip docstrings, include docstrings in the embedder's metadata fields so that text still affects retrieval.

Source Code​

To check this component's source code, open python_code_splitter.py in the Haystack repository.

Usage Examples​

Basic Configuration​

PythonCodeSplitter:
type: haystack.components.preprocessors.python_code_splitter.PythonCodeSplitter
init_parameters:
min_effective_lines: 20
max_effective_lines: 100
expected_chars_per_line: 45
oversized_factor: 3
strip_docstrings: false
preserve_class_definition: true
secondary_split_overlap: 5

Using the Component in an Index​

This index loads Python files as text and splits them by syntax before writing them to a document store.

# haystack-pipeline
components:
TextFileToDocument:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters:
encoding: utf-8
store_full_path: false
PythonCodeSplitter:
type: haystack.components.preprocessors.python_code_splitter.PythonCodeSplitter
init_parameters:
min_effective_lines: 20
max_effective_lines: 100
strip_docstrings: true
preserve_class_definition: true
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: python-source
create_index: true
policy: NONE

connections:
- sender: TextFileToDocument.documents
receiver: PythonCodeSplitter.documents
- sender: PythonCodeSplitter.documents
receiver: DocumentWriter.documents

max_runs_per_component: 100

inputs:
files:
- TextFileToDocument.sources

metadata: {}

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]List of documents whose content is Python source.

Outputs​

ParameterTypeDescription
documentsList[Document]List of syntax-aware chunks. Each chunk includes line numbers and the kinds of code units it contains.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
min_effective_linesint20Minimum effective lines per chunk. The splitter keeps merging the next unit while the chunk is below this size. Must be at least 1 and no greater than max_effective_lines.
max_effective_linesint100Target effective lines per chunk. Units are merged while that brings the total closer to this size.
expected_chars_per_lineint45Characters that count as one effective line. Longer lines count as more than one line.
oversized_factorint3A function longer than this factor times max_effective_lines is split by line, with overlap.
strip_docstringsboolFalseWhen True, function, method, and class docstrings move from the chunk text into meta.docstrings. The module docstring stays in the text.
preserve_class_definitionboolTrueWhen True, a chunk that contains class members but not the class header is prefixed with the class signature.
secondary_split_overlapint5Line overlap used only when an oversized function is split by line. The main split does not overlap.
secondary_split_lengthOptional[int]NoneLines per chunk for the oversized-function split. When empty, this uses max_effective_lines.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
documentsList[Document]List of documents whose content is Python source.