Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

GoogleGenAIMultimodalDocumentEmbedder

Compute embeddings for documents that contain images, PDFs, video, or audio files using Google AI multimodal embedding models.

Key Features​

  • Embeds non-textual documents (images, PDFs, video, and audio) using Google AI models such as gemini-embedding-2.
  • Supports both the Gemini Developer API and Vertex AI.
  • Reads file paths from document metadata and loads the files for embedding.
  • Stores the computed embedding in each document's embedding field, alongside the original document.
  • Supports batch processing with a configurable batch size for efficient embedding of large document sets.
  • Configurable output dimensionality and other embedding settings via the config parameter.

Configuration​

  1. Drag the GoogleGenAIMultimodalDocumentEmbedder component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Choose the API to use: gemini for the Gemini Developer API or vertex for Vertex AI.
    2. For the Gemini Developer API, enter your Google API key (GOOGLE_API_KEY or GEMINI_API_KEY). For Vertex AI, enter your GCP project ID and location. For instructions, see Create Secrets.
    3. Set the file_path_meta_field to the metadata field in your documents that holds the file path (default: file_path).
  4. Go to the Advanced tab to configure batch_size, image_size, model, and other settings.

Connections​

GoogleGenAIMultimodalDocumentEmbedder accepts a list of Document objects through its documents input. Each document must have a file path stored in the file_path metadata field (or the field you configure via file_path_meta_field). Supported file types are JPEG, PNG, PDF, MP4, QuickTime, MP3, WAV, and X-WAV.

It outputs a list of Document objects with embeddings stored in each document's embedding field, and a meta dictionary with model usage information.

Use this component in an indexing pipeline to embed multimodal documents before storing them. Connect its documents output to DocumentWriter.

Source Code​

To check this component's source code, open multimodal_document_embedder.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

GoogleGenAIMultimodalDocumentEmbedder:
type: haystack_integrations.components.embedders.google_genai.multimodal_document_embedder.GoogleGenAIMultimodalDocumentEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- GOOGLE_API_KEY
- GEMINI_API_KEY
strict: false
api: gemini
model: gemini-embedding-2
file_path_meta_field: file_path
batch_size: 6
progress_bar: true

Using the Component in a Pipeline​

This is an example of an indexing pipeline that embeds image documents and stores them in a document store.

# haystack-pipeline
components:
GoogleGenAIMultimodalDocumentEmbedder:
type: haystack_integrations.components.embedders.google_genai.multimodal_document_embedder.GoogleGenAIMultimodalDocumentEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- GOOGLE_API_KEY
- GEMINI_API_KEY
strict: false
api: gemini
vertex_ai_project:
vertex_ai_location:
model: gemini-embedding-2
file_path_meta_field: file_path
root_path:
image_size:
batch_size: 6
progress_bar: true
config:
timeout:
max_retries:

DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: multimodal-embeddings
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: false
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:
policy: OVERWRITE

connections:
- sender: GoogleGenAIMultimodalDocumentEmbedder.documents
receiver: DocumentWriter.documents

max_runs_per_component: 100

metadata: {}

inputs:
documents:
- GoogleGenAIMultimodalDocumentEmbedder.documents

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]A list of documents to embed. Each document must have a file path stored in the metadata field specified by file_path_meta_field.

Outputs​

ParameterTypeDescription
documentsList[Document]A list of documents with embeddings stored in their embedding field.
metaDict[str, Any]Information about the model used for embedding.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_keySecretSecret.from_env_var(['GOOGLE_API_KEY', 'GEMINI_API_KEY'], strict=False)Google API key. Not needed when using Vertex AI with Application Default Credentials.
apiLiteral['gemini', 'vertex']geminiWhich API to use. Use gemini for the Gemini Developer API or vertex for Vertex AI.
vertex_ai_projectOptional[str]NoneGoogle Cloud project ID for Vertex AI. Required when using Vertex AI with Application Default Credentials.
vertex_ai_locationOptional[str]NoneGoogle Cloud location for Vertex AI (for example, us-central1, europe-west1). Required when using Vertex AI with Application Default Credentials.
file_path_meta_fieldstrfile_pathThe metadata field in the Document that contains the file path to embed.
root_pathOptional[str]NoneThe root directory path where document files are located. When set, file paths in document metadata are resolved relative to this path and cannot escape it.
image_sizeOptional[tuple[int, int]]NoneIf provided, resizes images and PDF pages to fit within (width, height) while maintaining aspect ratio. Useful for reducing file size and processing time.
modelstrgemini-embedding-2The name of the model to use for calculating embeddings.
batch_sizeint6Number of documents to embed at once.
progress_barboolTrueIf True, shows a progress bar when running.
configOptional[Dict[str, Any]]NoneA dictionary to configure embedding content options. For example, set {"output_dimensionality": 768} to reduce embedding dimensions.
timeoutOptional[float]NoneTimeout in seconds for the underlying network requests.
max_retriesOptional[int]NoneMaximum number of retries for the underlying network requests.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
documentsList[Document]A list of documents to embed.