Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

JinaDocumentImageEmbedder

Compute document embeddings based on images using Jina AI multimodal models. Use this component in indexing pipelines to embed image documents before writing them to a document store.

Key Features​

  • Uses Jina AI multimodal models (including jina-clip-v2 and jina-embeddings-v4) to embed documents containing images or PDFs.
  • Stores the computed embedding in each document's embedding field, enabling image-based semantic search and retrieval.
  • Reads image file paths from a configurable document metadata field (default: file_path).
  • Supports JPEG and PNG image files, as well as PDF pages converted to images automatically.
  • Configurable embedding dimensions and image resizing for efficient storage and processing.
  • Supports both synchronous and asynchronous operation via run() and run_async().

Configuration​

  1. Drag the JinaDocumentImageEmbedder component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Create a secret with your Jina API key. Use JINA_API_KEY as the secret key. For instructions, see Create Secrets. Get your API key from Jina AI.
    2. Set the model field. Supported models include jina-clip-v1, jina-clip-v2 (default), and jina-embeddings-v4. Note that embedding_dimension is only supported by jina-embeddings-v4.
    3. Set the file_path_meta_field to the document metadata key that holds the image file path (default: file_path).
  4. Go to the Advanced tab to configure root_path, embedding_dimension, image_size, and batch_size.

Connections​

JinaDocumentImageEmbedder receives a list of documents through its documents input. Each document must include an image file path in its metadata under the field specified by file_path_meta_field. It outputs the same documents with their embeddings added through its documents output.

Connect it after a converter that creates documents with image paths in their metadata. Connect its documents output to DocumentWriter to store embedded documents in a document store.

Source Code​

To check this component's source code, open document_image_embedder.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

JinaDocumentImageEmbedder:
type: haystack_integrations.components.embedders.jina.document_image_embedder.JinaDocumentImageEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- JINA_API_KEY
strict: false
model: jina-clip-v2
file_path_meta_field: file_path
batch_size: 5

Using the Component in a Pipeline​

This example shows an indexing pipeline that takes documents with image file paths in their metadata, embeds the images using Jina, and writes them to a document store.

# haystack-pipeline
components:
JinaDocumentImageEmbedder:
type: haystack_integrations.components.embedders.jina.document_image_embedder.JinaDocumentImageEmbedder
init_parameters:
api_key:
type: env_var
env_vars:
- JINA_API_KEY
strict: false
model: jina-clip-v2
file_path_meta_field: file_path
root_path:
embedding_dimension:
image_size:
batch_size: 5
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
policy: OVERWRITE
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: Image-Index
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: false
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:

connections:
- sender: JinaDocumentImageEmbedder.documents
receiver: DocumentWriter.documents

max_runs_per_component: 100

metadata: {}

inputs:
documents:
- JinaDocumentImageEmbedder.documents

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]A list of Documents to embed. Each document must have an image file path in its metadata under the field specified by file_path_meta_field.

Outputs​

ParameterTypeDescription
documentsList[Document]Documents with their embedding field populated and an embedding_source entry added to their metadata.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_keySecretSecret.from_env_var('JINA_API_KEY')The Jina API key.
modelstrjina-clip-v2The name of the Jina multimodal model. Supported models: jina-clip-v1, jina-clip-v2, jina-embeddings-v4. See the Jina documentation for the full list.
base_urlstrhttps://api.jina.ai/v1/embeddingsThe Jina API base URL.
file_path_meta_fieldstrfile_pathThe document metadata field that contains the file path to the image or PDF.
root_pathOptional[str]NoneThe root directory for resolving relative file paths. When None, file paths are treated as absolute.
embedding_dimensionOptional[int]NoneNumber of embedding dimensions to return. Only supported by jina-embeddings-v4.
image_sizeOptional[tuple[int, int]]NoneTarget dimensions (width, height) to resize images while maintaining aspect ratio. Reduces memory usage and processing time.
batch_sizeint5Number of images to send in each API request.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
documentsList[Document]A list of Documents to embed.