GoogleGenAIMultimodalDocumentEmbedder
Compute embeddings for documents that contain images, PDFs, video, or audio files using Google AI multimodal embedding models.
Key Features
- Embeds non-textual documents (images, PDFs, video, and audio) using Google AI models such as
gemini-embedding-2. - Supports both the Gemini Developer API and Vertex AI.
- Reads file paths from document metadata and loads the files for embedding.
- Stores the computed embedding in each document's
embeddingfield, alongside the original document. - Supports batch processing with a configurable batch size for efficient embedding of large document sets.
- Configurable output dimensionality and other embedding settings via the
configparameter.
Configuration
- Drag the
GoogleGenAIMultimodalDocumentEmbeddercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Choose the API to use:
geminifor the Gemini Developer API orvertexfor Vertex AI. - For the Gemini Developer API, enter your Google API key (
GOOGLE_API_KEYorGEMINI_API_KEY). For Vertex AI, enter your GCP project ID and location. For instructions, see Create Secrets. - Set the
file_path_meta_fieldto the metadata field in your documents that holds the file path (default:file_path).
- Choose the API to use:
- Go to the Advanced tab to configure
batch_size,image_size,model, and other settings.
Connections
GoogleGenAIMultimodalDocumentEmbedder accepts a list of Document objects through its documents input. Each document must have a file path stored in the file_path metadata field (or the field you configure via file_path_meta_field). Supported file types are JPEG, PNG, PDF, MP4, QuickTime, MP3, WAV, and X-WAV.
It outputs a list of Document objects with embeddings stored in each document's embedding field, and a meta dictionary with model usage information.
Use this component in an indexing pipeline to embed multimodal documents before storing them. Connect its documents output to DocumentWriter.
Source Code
To check this component's source code, open multimodal_document_embedder.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
GoogleGenAIMultimodalDocumentEmbedder:
type: haystack_integrations.components.embedders.google_genai.multimodal_document_embedder.GoogleGenAIMultimodalDocumentEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- GOOGLE_API_KEY
- GEMINI_API_KEY
strict: false
api: gemini
model: gemini-embedding-2
file_path_meta_field: file_path
batch_size: 6
progress_bar: true
Using the Component in a Pipeline
This is an example of an indexing pipeline that embeds image documents and stores them in a document store.
# haystack-pipeline
components:
GoogleGenAIMultimodalDocumentEmbedder:
type: haystack_integrations.components.embedders.google_genai.multimodal_document_embedder.GoogleGenAIMultimodalDocumentEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- GOOGLE_API_KEY
- GEMINI_API_KEY
strict: false
api: gemini
vertex_ai_project:
vertex_ai_location:
model: gemini-embedding-2
file_path_meta_field: file_path
root_path:
image_size:
batch_size: 6
progress_bar: true
config:
timeout:
max_retries:
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: multimodal-embeddings
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: false
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:
policy: OVERWRITE
connections:
- sender: GoogleGenAIMultimodalDocumentEmbedder.documents
receiver: DocumentWriter.documents
max_runs_per_component: 100
metadata: {}
inputs:
documents:
- GoogleGenAIMultimodalDocumentEmbedder.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents to embed. Each document must have a file path stored in the metadata field specified by file_path_meta_field. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents with embeddings stored in their embedding field. |
meta | Dict[str, Any] | Information about the model used for embedding. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | Secret | Secret.from_env_var(['GOOGLE_API_KEY', 'GEMINI_API_KEY'], strict=False) | Google API key. Not needed when using Vertex AI with Application Default Credentials. |
api | Literal['gemini', 'vertex'] | gemini | Which API to use. Use gemini for the Gemini Developer API or vertex for Vertex AI. |
vertex_ai_project | Optional[str] | None | Google Cloud project ID for Vertex AI. Required when using Vertex AI with Application Default Credentials. |
vertex_ai_location | Optional[str] | None | Google Cloud location for Vertex AI (for example, us-central1, europe-west1). Required when using Vertex AI with Application Default Credentials. |
file_path_meta_field | str | file_path | The metadata field in the Document that contains the file path to embed. |
root_path | Optional[str] | None | The root directory path where document files are located. When set, file paths in document metadata are resolved relative to this path and cannot escape it. |
image_size | Optional[tuple[int, int]] | None | If provided, resizes images and PDF pages to fit within (width, height) while maintaining aspect ratio. Useful for reducing file size and processing time. |
model | str | gemini-embedding-2 | The name of the model to use for calculating embeddings. |
batch_size | int | 6 | Number of documents to embed at once. |
progress_bar | bool | True | If True, shows a progress bar when running. |
config | Optional[Dict[str, Any]] | None | A dictionary to configure embedding content options. For example, set {"output_dimensionality": 768} to reduce embedding dimensions. |
timeout | Optional[float] | None | Timeout in seconds for the underlying network requests. |
max_retries | Optional[int] | None | Maximum number of retries for the underlying network requests. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents to embed. |
Related Information
Was this page helpful?