
Curated guide to zero-data-retention configurations for LLM APIs. Covers provider-specific ZDR endpoints, threat models, compliance mappings, and self-hosting patterns for engineers in regulated industries.
Last updated: April 2026
A practical guide to keeping your data private when using LLM APIs. Covers zero-retention endpoints, self-hosting, compliance requirements, and data protection patterns for engineers in regulated industries.
"Zero-retention" is not a single feature — it is a bundle of technical controls + contract terms ensuring customer content (prompts, outputs, files) is not stored at rest by the vendor. Different approaches offer different trade-offs:
flowchart TD
Start(["Need Private AI?"]) --> Q1{"Can you\nself-host?"}
Q1 -->|"Yes, have GPUs"| SH["Self-Host Open Weights\n(Llama 4 · DeepSeek · Mistral · Qwen)"]
Q1 -->|"Yes, CPU only"| OL["Ollama + Quantized Models\n(7B–14B on consumer hardware)"]
Q1 -->|No| Q2{"Need frontier\nmodel quality?"}
Q2 -->|Yes| Q3{"Regulatory\nrequirements?"}
Q2 -->|No| Q4{"Budget\nconstrained?"}
Q3 -->|"HIPAA / FedRAMP"| Cloud["Azure OpenAI · AWS Bedrock\n+ Private Endpoints + BAA"]
Q3 -->|"Multi-provider"| GW["OpenRouter · Cloudflare AI Gateway\nwith ZDR routing"]
Q3 -->|"Single provider OK"| Direct["Direct ZDR Contract\n(OpenAI · Anthropic · Google)"]
Q4 -->|Yes| Budget["Fireworks · Together AI\n(open-weights, low cost, ZDR included)"]
Q4 -->|"Not really"| Fast["Groq · Fireworks · Together\nZDR toggle in dashboard"]
style Start fill:#4a90d9,stroke:#2c5f8a,color:#fff
style SH fill:#2ecc71,stroke:#1a9c54,color:#fff
style OL fill:#2ecc71,stroke:#1a9c54,color:#fff
style Cloud fill:#e67e22,stroke:#b3611a,color:#fff
style GW fill:#9b59b6,stroke:#7a3d92,color:#fff
style Direct fill:#3498db,stroke:#2471a3,color:#fff
style Budget fill:#1abc9c,stroke:#148f77,color:#fff
style Fast fill:#1abc9c,stroke:#148f77,color:#fff
| Approach | Privacy Strength | Model Quality | Operational Cost | Setup Complexity |
|---|---|---|---|---|
| Self-hosted (air-gapped) | Strongest | Open-weight only | Hardware + ops | High |
| Self-hosted (VPC) | Very strong | Open-weight only | Cloud GPU cost | Medium |
| Cloud ZDR + Private Link | Strong (contractual) | Frontier models | API pricing | Low-Medium |
| SaaS ZDR API | Good (contractual) | Frontier models | API pricing | Low |
| Gateway with ZDR routing | Good (delegated) | Multi-provider | API + gateway fee | Low |
Before choosing an approach, understand what you're protecting against:
| Threat | Description | Mitigated By |
|---|---|---|
| Training data leakage | Your prompts/outputs used to train the provider's models | ZDR contract, API-tier (not free-tier), self-hosting |
| Abuse monitoring retention | Provider stores prompts for safety review (often 30 days) | ZDR/MAM opt-out, self-hosting |
| Employee access | Provider staff can view your data during incident response | ZDR + BYOK encryption, self-hosting |
| Subpoena / legal discovery | Government or legal requests to the provider for your data | Self-hosting, data residency controls, no-retention contract |
| Breach at provider | Provider's systems compromised, your data exfiltrated | No-retention (nothing to steal), self-hosting, encryption at rest |
| Your own logging | Your infra (proxies, APM, error trackers) logs sensitive prompts | DLP proxy, log redaction, audit your pipeline |
| Prompt injection exfiltration | Malicious input causes LLM to leak data via tool calls | Output scanning, least-privilege tools, sandboxing |
flowchart LR
User["User Input"] --> App["Your App"]
subgraph YourInfra["Your Infrastructure"]
App --> Logs1["App Logs ⚠️"]
App --> DLP["DLP / PII Proxy"]
DLP --> GW["API Gateway"]
GW --> Logs2["Gateway Logs ⚠️"]
end
subgraph Provider["LLM Provider"]
GW --> Inference["Model Inference\n(in-memory)"]
Inference --> Abuse["Abuse Monitor\n(0–30 day retention)"]
Inference --> Training["Model Training\n(opt-out or ZDR)"]
end
Inference --> Response["Response"]
Response --> App
style Logs1 fill:#e74c3c,stroke:#c0392b,color:#fff
style Logs2 fill:#e74c3c,stroke:#c0392b,color:#fff
style Abuse fill:#f39c12,stroke:#d68910,color:#fff
style Training fill:#e74c3c,stroke:#c0392b,color:#fff
style DLP fill:#2ecc71,stroke:#1a9c54,color:#fff
style Inference fill:#3498db,stroke:#2471a3,color:#fff
Red = risk points where data can be retained. Green = protection layer. ZDR eliminates the provider-side risks; DLP/proxy eliminates your-side risks.
store parameter is always treated as false, even if set to true in requestsstore parameter functional — for orgs that need data retention but reduced monitoringZDR-Eligible Endpoints:
/v1/chat/completions, /v1/responses, /v1/images/*, /v1/embeddings, /v1/audio/*, /v1/moderations, /v1/completions, /v1/realtime
NOT ZDR-Eligible:
Assistants API (/v1/assistants, /v1/threads, /v1/vector_stores), Conversations API, Files, Fine-tuning, Batches, Evals, Background mode (/v1/responses with background: true), Hosted containers (Code Interpreter)
Additional Controls:
eu.api.openai.com), AU (au.api.openai.com) — requires ZDR amendment, 10% cost uplift