SupabaseBucketDownloader
Download files from a Supabase Storage bucket and return them as ByteStream objects ready for further processing in indexing pipelines.
Key Features
- Downloads files in-memory from a Supabase Storage bucket.
- Returns
ByteStreamobjects withfile_pathandbucket_nameset in metadata. - Supports filtering downloads by file extension (for example,
[".pdf", ".txt"]). - Automatically detects MIME types from file extensions.
- Initializes the Supabase client lazily on first use via
warm_up().
Configuration
- Drag the
SupabaseBucketDownloadercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set
supabase_urlto your Supabase project URL, for examplehttps://<project-ref>.supabase.co. - Set
supabase_keyas a secret calledSUPABASE_SERVICE_KEY. Use the service role key to access private buckets. For instructions, see Add Secrets. - Set
bucket_nameto the name of the Supabase Storage bucket to download files from.
- Set
- Optionally, set
file_extensionson the Advanced tab to filter downloads.
Connections
SupabaseBucketDownloader receives a sources list of file paths within the bucket at runtime. It outputs a list of ByteStream objects that you can connect to a document converter such as PyPDFToDocument or TextFileToDocument.
Source Code
To check this component's source code, open supabase_bucket_downloader.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
SupabaseBucketDownloader:
type: haystack_integrations.components.downloaders.supabase.supabase_bucket_downloader.SupabaseBucketDownloader
init_parameters:
supabase_url: https://<project-ref>.supabase.co
supabase_key:
type: env_var
env_vars:
- SUPABASE_SERVICE_KEY
strict: false
bucket_name: my-documents
file_extensions:
- .pdf
- .txt
Using the Component in a Pipeline
# haystack-pipeline
components:
downloader:
type: haystack_integrations.components.downloaders.supabase.supabase_bucket_downloader.SupabaseBucketDownloader
init_parameters:
supabase_url: https://<project-ref>.supabase.co
supabase_key:
type: env_var
env_vars:
- SUPABASE_SERVICE_KEY
strict: false
bucket_name: my-documents
file_extensions:
- .pdf
converter:
type: haystack.components.converters.PyPDFToDocument
init_parameters: {}
connections:
- sender: downloader.streams
receiver: converter.sources
inputs:
sources:
- downloader.sources
outputs:
documents: converter.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
sources | List[str] | List of file paths within the bucket to download, for example ["folder/file.pdf", "notes.txt"]. |
Outputs
| Parameter | Type | Description |
|---|---|---|
streams | List[ByteStream] | A list of ByteStream objects, one per successfully downloaded file. Each stream has meta["file_path"] and meta["bucket_name"] set. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
supabase_url | str | The URL of your Supabase project, for example https://<project-ref>.supabase.co. | |
supabase_key | Secret | Secret.from_env_var("SUPABASE_SERVICE_KEY") | The Supabase API key used to authenticate requests. Use the service role key for private buckets. |
bucket_name | str | The name of the Supabase Storage bucket to download files from. | |
file_extensions | Optional[List[str]] | None | Optional list of file extensions to filter downloads, for example [".pdf", ".txt"]. Extensions are matched case-insensitively. When None, all files are downloaded. |
Related Information
Was this page helpful?