EdenAIDocumentEmbedder
Compute document embeddings using Eden AI's multi-provider embedding API. Use this component in indexing pipelines to embed documents before writing them to a document store.
Key Features
- Routes embedding requests through Eden AI to multiple providers—including OpenAI, Mistral, Cohere, and Google—using a single API key.
- Uses Eden AI's
provider/modelnaming convention, for exampleopenai/text-embedding-3-smallormistral/mistral-embed. - Stores the computed embedding in each document's
embeddingfield, making documents ready for semantic search and retrieval. - Supports EU data residency through Eden AI's infrastructure.
- Configurable batch size and progress bar for large document sets.
- Supports embedding additional metadata fields alongside document content.
The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.
This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Configuration
- Drag the
EdenAIDocumentEmbeddercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Create a secret with your Eden AI API key. Use
EDENAI_API_KEYas the secret key. For instructions, see Create Secrets. Get your API key from Eden AI. - Set the model field to the desired embedding model in
provider/modelformat, for exampleopenai/text-embedding-3-small.
- Create a secret with your Eden AI API key. Use
- Go to the Advanced tab to configure
batch_size,meta_fields_to_embed,timeout, andmax_retries.
Connections
EdenAIDocumentEmbedder receives a list of documents through its documents input. It outputs the same documents with their embeddings added through its documents output.
Connect it after converters (such as TextFileToDocument or HTMLToDocument) or after DocumentSplitter to embed chunks. Connect its documents output to DocumentWriter to store embedded documents in a document store.
Source Code
To check this component's source code, open document_embedder.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
EdenAIDocumentEmbedder:
type: haystack_integrations.components.embedders.edenai.document_embedder.EdenAIDocumentEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- EDENAI_API_KEY
strict: false
model: openai/text-embedding-3-small
api_base_url: https://api.edenai.run/v3
batch_size: 32
progress_bar: true
embedding_separator: "\n"
Using the Component in a Pipeline
This example shows an indexing pipeline that reads text files, splits them into chunks, embeds them using Eden AI, and writes them to an OpenSearch document store.
# haystack-pipeline
components:
TextFileToDocument:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters:
encoding: utf-8
store_full_path: false
DocumentSplitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: sentence
split_length: 100
split_overlap: 0
split_threshold: 0
splitting_function:
EdenAIDocumentEmbedder:
type: haystack_integrations.components.embedders.edenai.document_embedder.EdenAIDocumentEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- EDENAI_API_KEY
strict: false
model: openai/text-embedding-3-small
api_base_url: https://api.edenai.run/v3
batch_size: 32
progress_bar: true
meta_fields_to_embed:
embedding_separator: "\n"
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
policy: OVERWRITE
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: Standard-Index-English
max_chunk_bytes: 104857600
embedding_dim: 1536
return_embedding: false
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:
connections:
- sender: TextFileToDocument.documents
receiver: DocumentSplitter.documents
- sender: DocumentSplitter.documents
receiver: EdenAIDocumentEmbedder.documents
- sender: EdenAIDocumentEmbedder.documents
receiver: DocumentWriter.documents
max_runs_per_component: 100
metadata: {}
inputs:
files:
- TextFileToDocument.sources
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of Documents to embed. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Documents with their embedding field populated. |
meta | Dict[str, Any] | Metadata about the embedding request, including model name and usage statistics. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | Secret | Secret.from_env_var('EDENAI_API_KEY') | The Eden AI API key. |
model | str | openai/text-embedding-3-small | The Eden AI embedding model in provider/model format. For a full list, see the Eden AI models catalog. |
api_base_url | Optional[str] | https://api.edenai.run/v3 | The Eden AI API base URL. |
prefix | str | "" | A string to add to the beginning of each text. |
suffix | str | "" | A string to add to the end of each text. |
batch_size | int | 32 | Number of Documents to encode at once. |
progress_bar | bool | True | Whether to show a progress bar. Disable in production to keep logs clean. |
meta_fields_to_embed | Optional[List[str]] | None | Metadata fields to embed alongside the Document text. |
embedding_separator | str | \n | Separator used to concatenate metadata fields and Document text. |
timeout | Optional[float] | None | Timeout for API calls in seconds. Defaults to the OPENAI_TIMEOUT environment variable or 30 seconds. |
max_retries | Optional[int] | None | Maximum number of retries after an internal error. Defaults to the OPENAI_MAX_RETRIES environment variable or 5. |
http_client_kwargs | Optional[Dict[str, Any]] | None | Keyword arguments for a custom httpx.Client or httpx.AsyncClient. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of Documents to embed. |
Related Information
Was this page helpful?