PythonCodeSplitter
Split Python source into chunks that follow the file's structure. Functions and methods stay whole, and each chunk reads from top to bottom like the original file. Use this component when you index code for search.
Key Features
- Parses each file with Python's syntax tree and keeps imports, functions, classes, and methods as units.
- Merges those units toward a target size measured in effective lines. A long physical line counts as more than one effective line.
- Leaves the primary split without overlap. Only an oversized function is split again, by line, and those pieces can overlap.
- Can move function, method, and class docstrings out of the chunk text and into metadata.
- Can prefix a chunk with the class signature when the chunk contains class members but not the class header.
- Stores
source_id,split_id,start_line,end_line, andunit_kindson each chunk. Parent metadata, includingfile_name, is copied through.
The document content must be valid Python. The component raises a syntax error when the source does not parse.
Configuration
- Drag the
PythonCodeSplittercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set Min Effective Lines and Max Effective Lines. The splitter keeps adding the next unit while the chunk is below the minimum, and it aims for the maximum. Defaults are 20 and 100.
- Set Expected Chars Per Line to decide how many characters count as one effective line. The default is 45.
- Go to the Advanced tab:
- Set Oversized Factor. A function longer than this factor times the maximum is split by line. The default factor is three.
- Toggle Strip Docstrings to move function, method, and class docstrings into
meta.docstrings. The module docstring stays in the text. - Toggle Preserve Class Definition to repeat the class signature on chunks that contain members but not the class header.
- Set Secondary Split Overlap and Secondary Split Length for the line split used on oversized functions.
Connections
PythonCodeSplitter accepts a list of Document objects whose content is Python source. It outputs a list of split Document objects.
It typically receives documents from TextFileToDocument and sends chunks to an embedder or DocumentWriter. If you strip docstrings, include docstrings in the embedder's metadata fields so that text still affects retrieval.
Source Code
To check this component's source code, open python_code_splitter.py in the Haystack repository.
Usage Examples
Basic Configuration
PythonCodeSplitter:
type: haystack.components.preprocessors.python_code_splitter.PythonCodeSplitter
init_parameters:
min_effective_lines: 20
max_effective_lines: 100
expected_chars_per_line: 45
oversized_factor: 3
strip_docstrings: false
preserve_class_definition: true
secondary_split_overlap: 5
Using the Component in an Index
This index loads Python files as text and splits them by syntax before writing them to a document store.
# haystack-pipeline
components:
TextFileToDocument:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters:
encoding: utf-8
store_full_path: false
PythonCodeSplitter:
type: haystack.components.preprocessors.python_code_splitter.PythonCodeSplitter
init_parameters:
min_effective_lines: 20
max_effective_lines: 100
strip_docstrings: true
preserve_class_definition: true
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: python-source
create_index: true
policy: NONE
connections:
- sender: TextFileToDocument.documents
receiver: PythonCodeSplitter.documents
- sender: PythonCodeSplitter.documents
receiver: DocumentWriter.documents
max_runs_per_component: 100
inputs:
files:
- TextFileToDocument.sources
metadata: {}
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | List of documents whose content is Python source. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | List of syntax-aware chunks. Each chunk includes line numbers and the kinds of code units it contains. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
min_effective_lines | int | 20 | Minimum effective lines per chunk. The splitter keeps merging the next unit while the chunk is below this size. Must be at least 1 and no greater than max_effective_lines. |
max_effective_lines | int | 100 | Target effective lines per chunk. Units are merged while that brings the total closer to this size. |
expected_chars_per_line | int | 45 | Characters that count as one effective line. Longer lines count as more than one line. |
oversized_factor | int | 3 | A function longer than this factor times max_effective_lines is split by line, with overlap. |
strip_docstrings | bool | False | When True, function, method, and class docstrings move from the chunk text into meta.docstrings. The module docstring stays in the text. |
preserve_class_definition | bool | True | When True, a chunk that contains class members but not the class header is prefixed with the class signature. |
secondary_split_overlap | int | 5 | Line overlap used only when an oversized function is split by line. The main split does not overlap. |
secondary_split_length | Optional[int] | None | Lines per chunk for the oversized-function split. When empty, this uses max_effective_lines. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | List of documents whose content is Python source. |
Related Information
Was this page helpful?