FastembedDocumentEmbedder
Compute embeddings for a list of documents using Fastembed embedding models. Use this component in indexing pipelines to prepare documents for embedding-based retrieval.
Key Features
- Uses CPU-optimized embedding models from the Fastembed library.
- Fast and lightweight — designed for production use without GPU requirements.
- Stores the computed embedding in each document's
embeddingfield. - Supports adding metadata fields to the text before embedding.
- Configurable batch size for efficient processing.
The embedding model you use to embed documents in your indexing pipeline must be the same as the embedding model you use to embed the query in your query pipeline.
This means the embedders for your indexing and query pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Configuration
- Drag the
FastembedDocumentEmbeddercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set the
modelto the Fastembed model you want to use (for example,BAAI/bge-small-en-v1.5). For a list of supported models, see the Fastembed documentation.
- Set the
- Go to the Advanced tab to configure
batch_size,meta_fields_to_embed,embedding_separator,prefix, andsuffix.
Connections
FastembedDocumentEmbedder receives a list of documents, typically from a document splitter or converter. It outputs the same documents with embeddings added to their embedding field, ready to be sent to a DocumentWriter.
Source Code
To check this component's source code, open fastembed_document_embedder.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
FastembedDocumentEmbedder:
type: haystack_integrations.components.embedders.fastembed.FastembedDocumentEmbedder
init_parameters:
model: BAAI/bge-small-en-v1.5
batch_size: 256
Using the Component in a Pipeline
# haystack-pipeline
components:
FastembedDocumentEmbedder:
type: haystack_integrations.components.embedders.fastembed.FastembedDocumentEmbedder
init_parameters:
model: BAAI/bge-small-en-v1.5
batch_size: 256
meta_fields_to_embed:
embedding_separator: "\n"
document_writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.qdrant.document_store.QdrantDocumentStore
init_parameters:
url: http://localhost:6333
index: documents
connections:
- sender: FastembedDocumentEmbedder.documents
receiver: document_writer.documents
max_runs_per_component: 100
metadata: {}
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents to embed. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The documents with their embedding field populated. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
model | str | BAAI/bge-small-en-v1.5 | The name of the Fastembed model to use. For a full list, see the Fastembed documentation. |
cache_dir | Optional[str] | None | The directory to cache downloaded models. |
threads | Optional[int] | None | The number of threads for model inference. |
prefix | str | "" | A string to add at the beginning of each document's text before embedding. |
suffix | str | "" | A string to add at the end of each document's text before embedding. |
batch_size | int | 256 | The number of documents to process in each batch. |
progress_bar | bool | True | Whether to show a progress bar during embedding. |
parallel | Optional[int] | None | The number of parallel processes for batch processing. |
local_files_only | bool | False | Whether to use only locally cached models. |
meta_fields_to_embed | Optional[List[str]] | None | A list of document metadata field names to include in the text before embedding. |
embedding_separator | str | "\n" | The separator used to join the document text and metadata fields. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
documents | List[Document] | A list of documents to embed. |
Related Information
Was this page helpful?