SupabasePgvectorEmbeddingRetriever
Retrieve documents from a SupabasePgvectorDocumentStore based on dense embedding similarity. This component is a thin wrapper around PgvectorEmbeddingRetriever adapted for use with Supabase.
The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.
This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Key Features
- Performs dense vector similarity search against a Supabase PostgreSQL database with pgvector.
- Supports
cosine_similarity,inner_product, andl2_distancevector functions. - Reads the connection string from the
SUPABASE_DB_URLenvironment variable by default. - Compatible with HNSW index search strategy for performance at scale.
- Configurable filter policy to merge or replace filters at query time.
Configuration
- Drag the
SupabasePgvectorEmbeddingRetrievercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- Configure the nested
SupabasePgvectorDocumentStore:- Set
connection_stringas a secret calledSUPABASE_DB_URL. The connection string format ispostgresql://postgres.[project-ref]:[password]@aws-0-[region].pooler.supabase.com:5432/postgres. For instructions, see Add Secrets. - Use session mode (port 5432) or a direct connection for best compatibility with pgvector.
- Set
- Connect a text embedder to provide
query_embeddingas input.
Connections
SupabasePgvectorEmbeddingRetriever receives a query_embedding (list of floats) from a text embedder such as SentenceTransformersTextEmbedder. It outputs a list of Document objects you can connect to PromptBuilder or other downstream components.
Source Code
To check this component's source code, open embedding_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
SupabasePgvectorEmbeddingRetriever:
type: haystack_integrations.components.retrievers.supabase.embedding_retriever.SupabasePgvectorEmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.supabase.document_store.SupabasePgvectorDocumentStore
init_parameters:
connection_string:
type: env_var
env_vars:
- SUPABASE_DB_URL
strict: false
embedding_dimension: 384
top_k: 10
Using the Component in a Pipeline
# haystack-pipeline
components:
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
document_store:
type: haystack_integrations.document_stores.supabase.document_store.SupabasePgvectorDocumentStore
init_parameters:
connection_string:
type: env_var
env_vars:
- SUPABASE_DB_URL
strict: false
embedding_dimension: 384
vector_function: cosine_similarity
retriever:
type: haystack_integrations.components.retrievers.supabase.embedding_retriever.SupabasePgvectorEmbeddingRetriever
init_parameters:
document_store: document_store
top_k: 5
vector_function: cosine_similarity
connections:
- sender: text_embedder.embedding
receiver: retriever.query_embedding
inputs:
query:
- text_embedder.text
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query_embedding | List[float] | The query embedding vector to search for similar documents. |
filters | Optional[Dict[str, Any]] | Filters to apply at query time. |
top_k | Optional[int] | Maximum number of documents to retrieve. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of the most similar documents from the document store. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | SupabasePgvectorDocumentStore | The Supabase pgvector document store to retrieve documents from. | |
filters | Optional[Dict[str, Any]] | None | Default filters applied to all searches. |
top_k | int | 10 | Maximum number of documents to return. |
vector_function | Optional[Literal["cosine_similarity", "inner_product", "l2_distance"]] | None | The similarity function to use. Defaults to the one set in the document store. When using HNSW search strategy, this should match the function used during index creation. |
filter_policy | FilterPolicy | FilterPolicy.REPLACE | How to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query_embedding | List[float] | The embedding vector to search with. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters to apply. |
top_k | Optional[int] | None | Maximum number of documents to retrieve. Overrides the init-time value. |
Related Information
Was this page helpful?