AzureDocumentDBFullTextRetriever
Retrieve documents from an AzureDocumentDBDocumentStore using BM25 full-text search. This component is useful for keyword-based retrieval pipelines that rely on lexical matching rather than dense vector similarity. Azure DocumentDB is Azure Cosmos DB for MongoDB (vCore), so this component is also relevant if you're searching for "Cosmos DB" or "vCore".
Full-text search in Azure DocumentDB is currently a gated preview and must be enabled on the cluster before using this retriever. You must also configure a full_text_search_index on the document store.
Key Features
- Performs BM25 full-text search using Azure DocumentDB's gated-preview text search capability.
- Accepts a single query string or a list of query strings.
- Supports optional fuzzy matching with configurable
maxEditsfor typo tolerance. - Configurable filter policy to merge or replace filters at query time.
- Supports both synchronous and asynchronous execution.
Configuration
Add Workspace-Level Integration
- Click your profile icon and choose Settings.
- Go to Workspace>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in the current workspace.
Add Organization-Level Integration
- Click your profile icon and choose Settings.
- Go to Organization>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in all workspaces in the current organization.
- Make sure full-text search is enabled on your Azure DocumentDB cluster (gated preview).
- Set the
AZURE_DOCUMENTDB_CLUSTER_NAMEenvironment variable with your cluster name, or provide amongo_connection_stringfor local development. - Configure an
AzureDocumentDBDocumentStorewith afull_text_search_indexpointing to your text search index. - Drag the
AzureDocumentDBFullTextRetrievercomponent onto the canvas from the Component Library. - Connect the retriever output to downstream components such as
PromptBuilder.
Connections
AzureDocumentDBFullTextRetriever receives a query string (or list of strings) as input. It outputs a list of Document objects ranked by BM25 relevance that you can connect to PromptBuilder or other downstream components.
Source Code
To check this component's source code, open full_text_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
AzureDocumentDBFullTextRetriever:
type: haystack_integrations.components.retrievers.azure_documentdb.full_text_retriever.AzureDocumentDBFullTextRetriever
init_parameters:
document_store: AzureDocumentDBDocumentStore
top_k: 5
Using the Component in a Pipeline
# haystack-pipeline
components:
document_store:
type: haystack_integrations.document_stores.azure_documentdb.document_store.AzureDocumentDBDocumentStore
init_parameters:
database_name: haystack
collection_name: documents
full_text_search_index: haystack_fts_index
retriever:
type: haystack_integrations.components.retrievers.azure_documentdb.full_text_retriever.AzureDocumentDBFullTextRetriever
init_parameters:
document_store: document_store
top_k: 5
inputs:
query:
- retriever.query
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | List[str] | The keyword query string or list of strings to search for. |
fuzzy | Optional[Dict[str, int]] | Azure DocumentDB fuzzy search options, for example {"maxEdits": 1} for typo tolerance. |
filters | Optional[Dict[str, Any]] | Filters to apply when retrieving documents. |
top_k | Optional[int] | The maximum number of documents to retrieve. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents ranked by BM25 relevance from the document store. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | AzureDocumentDBDocumentStore | The Azure DocumentDB document store to retrieve documents from. Requires full_text_search_index to be set. | |
filters | Optional[Dict[str, Any]] | None | Default filters to apply when retrieving documents. |
top_k | int | 10 | The maximum number of documents to retrieve. |
filter_policy | FilterPolicy | FilterPolicy.REPLACE | How to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | List[str] | The keyword query string or list of strings. | |
fuzzy | Optional[Dict[str, int]] | None | Fuzzy search options such as {"maxEdits": 1}. |
filters | Optional[Dict[str, Any]] | None | Runtime filters to apply. |
top_k | Optional[int] | None | Maximum number of documents to retrieve. Overrides the init-time value. |
Related Information
Was this page helpful?