SupabasePgvectorKeywordRetriever
Retrieve documents from a SupabasePgvectorDocumentStore using PostgreSQL full-text search. This component is a thin wrapper around PgvectorKeywordRetriever adapted for use with Supabase.
Key Features
- Performs keyword-based full-text search using PostgreSQL's
ts_rank_cdfunction. - Ranks results based on term frequency, term proximity, and document section importance.
- Reads the connection string from the
SUPABASE_DB_URLenvironment variable by default. - Works without embeddings, making it suitable for pipelines that don't use dense retrieval.
- Configurable filter policy to merge or replace filters at query time.
Configuration
- Drag the
SupabasePgvectorKeywordRetrievercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- Configure the nested
SupabasePgvectorDocumentStore:- Set
connection_stringas a secret calledSUPABASE_DB_URL. The connection string format ispostgresql://postgres.[project-ref]:[password]@aws-0-[region].pooler.supabase.com:5432/postgres. For instructions, see Add Secrets. - Use session mode (port 5432) or a direct connection for best compatibility.
- Set
Connections
SupabasePgvectorKeywordRetriever receives a query string at runtime. It outputs a list of Document objects that you can connect to a PromptBuilder or other downstream components.
Source Code
To check this component's source code, open keyword_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
SupabasePgvectorKeywordRetriever:
type: haystack_integrations.components.retrievers.supabase.keyword_retriever.SupabasePgvectorKeywordRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.supabase.document_store.SupabasePgvectorDocumentStore
init_parameters:
connection_string:
type: env_var
env_vars:
- SUPABASE_DB_URL
strict: false
embedding_dimension: 768
top_k: 10
Using the Component in a Pipeline
# haystack-pipeline
components:
document_store:
type: haystack_integrations.document_stores.supabase.document_store.SupabasePgvectorDocumentStore
init_parameters:
connection_string:
type: env_var
env_vars:
- SUPABASE_DB_URL
strict: false
embedding_dimension: 768
retriever:
type: haystack_integrations.components.retrievers.supabase.keyword_retriever.SupabasePgvectorKeywordRetriever
init_parameters:
document_store: document_store
top_k: 10
connections: []
inputs:
query:
- retriever.query
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | The keyword query string to search for. |
filters | Optional[Dict[str, Any]] | Filters to apply at query time. |
top_k | Optional[int] | Maximum number of documents to retrieve. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of matching documents ranked by relevance. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | SupabasePgvectorDocumentStore | The Supabase pgvector document store to retrieve documents from. | |
filters | Optional[Dict[str, Any]] | None | Default filters applied to all searches. |
top_k | int | 10 | Maximum number of documents to return. |
filter_policy | FilterPolicy | FilterPolicy.REPLACE | How to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The keyword query string. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters to apply. |
top_k | Optional[int] | None | Maximum number of documents to retrieve. Overrides the init-time value. |
Related Information
Was this page helpful?