Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

Data Privacy and LLM Providers

To stay in control, it's crucial to know where your content is stored, and what leaves your boundary when a pipeline calls a model provider. Learn about retention, provider calls, and what you can control in pipeline design.


Retention of Queries and Traces​

Each pipeline run can create a search history entry and an associated trace. You can delete search history at any time. Trace data remains available for 90 days. To learn more, see What's the Retention Policy for Search History and Pipeline Logs?.

Calls to External Model Providers​

Most generative and embedding components call your chosen provider (OpenAI, Anthropic, Amazon Bedrock, Azure OpenAI, Google Gemini, and others). From a privacy perspective:

  • You add provider credentials as secrets or integrations at the organization or workspace level. For more details, see Secrets and Integrations.
  • Prompt text, retrieved documents, and chat history sent to a model come from your pipeline YAML, not from a hidden platform prompt.
  • Provider-side logging, training, and data use follow that provider's terms and your agreement with them.

For supported providers and connection setup, see Supported Connections and Integrations.

Reducing Data Sent to Models​

You control what reaches a large language model (LLM) through pipeline design:

  • Limit retrieval (top_k, filters, metadata) so only relevant chunks are passed to the generator.
  • Use workspace or organization secrets so production keys differ from development keys.
  • Deploy indexing and query pipelines in VPC configurations so document content stays in storage you operate.

If you need minimal platform retention, combine VPC storage, proactive deletion of search history, and external observability with your own retention policy.