SaaS RAG security: isolate tenant retrieval and customer data
Design tenant-safe SaaS retrieval-augmented generation with authorization-aware search, isolated vector data, trusted metadata, source citations and leakage tests.
In this guide
How do you secure a multi-tenant RAG system?
A secure SaaS retrieval-augmented generation (RAG) system authorizes a user's access before retrieving customer content, carries that verified scope into search, and checks the returned chunks again before sending them to a model. A vector filter or tenant label helps implement that boundary, but it is not a substitute for application authorization. Treat retrieved text as untrusted input because it can contain private data, stale permissions or instructions aimed at the model.
Derive retrieval scope from authenticated membership
Resolve the caller, active organization, role and permitted records in trusted server code. Build the retrieval filter from that authorization result; never let a browser supply a tenant ID, document ACL, user group or filter expression that broadens access. If the system cannot establish the caller's tenant and permissions, stop before search and generation.
Enforce access at retrieval time and on every returned chunk
Apply tenant and document permissions in the vector query or a separate authorized retrieval layer, then verify each result's tenant, document state and current ACL before constructing model context. AWS documents metadata filtering and identity-based ACL patterns for particular Bedrock integrations; exact capabilities differ by store. A post-generation check is too late because the model has already seen the content.
Choose storage boundaries that match the risk
Compare separate indexes or collections, tenant namespaces, and a shared index with mandatory server-side metadata filters. Consider administrator blast radius, noisy-neighbor effects, encryption and key options, backup and deletion behavior, operational complexity and provider guarantees. A namespace is useful defense in depth, but broad service credentials can still cross namespaces unless the store enforces that boundary.
| Caller and verified tenant | Document permission source | Index/filter boundary | Pre-generation checks | Leakage test owner |
|---|---|---|---|---|
| Tenant member asking about a private file | ||||
| Support user with approved access | ||||
| Former member after access revocation |
How should SaaS teams ingest and update RAG documents?
Attach trustworthy provenance and permission metadata
For each chunk retain a stable document ID, tenant ID, source, version, ingestion time, classification and the permission reference needed for retrieval. Validate metadata at ingestion and reject missing or conflicting tenant ownership. Do not infer an ACL from a filename or user-editable tag; map source-system permissions through a controlled, reviewed process.
Synchronize permission changes, deletion and re-indexing
Treat an ACL change, tenant transfer, document deletion or user offboarding as a security event for every derived chunk, cached answer and embedding. Use a reconciliation job that can find stale copies and retry safely; mark content unavailable while updates are incomplete. Test that deletion and revocation reach replicas, backups and downstream indexes according to the product's documented retention rules.
Keep retrieval context narrow and traceable
Retrieve only the few authorized passages needed for the task, cap context size and preserve chunk-to-source references. Avoid sending an entire tenant corpus to the model for convenience. Record which document versions informed an answer so an incident reviewer can reconstruct exposure without logging the full sensitive prompt or response by default.
How do you test RAG tenant isolation and retrieval quality?
Run adversarial cross-tenant retrieval tests
Seed tenants A and B with distinctive canary phrases, then search as each tenant across direct questions, paraphrases, filters, hybrid search, reranking and fallback paths. Assert that no result, citation, snippet, log or generated answer reveals the other tenant's canary. Include anonymous callers, suspended tenants, revoked memberships and malformed metadata.
Test injection and poisoned-source cases separately
Place instruction-like text in an otherwise ordinary authorized document and confirm the model treats it as quoted source data, not as permission to reveal secrets or call a tool. Check ingestion review, source provenance and removal workflow for poisoned or compromised documents. Filtering search results does not neutralize malicious instructions inside a result that the user is allowed to read.
Monitor access decisions without collecting excess content
Track tenant-scoped retrieval counts, denied requests, missing ACL metadata, stale-index lag, unusual broad searches and citation failures. Use opaque IDs and limited retention; keep prompts and retrieved text out of routine logs unless a reviewed incident workflow requires them. Alert on a filter unexpectedly disappearing or a retrieval service switching to an unscoped fallback.
Multi-tenant RAG security FAQs
Does adding tenant_id metadata to vector chunks prevent data leaks?
No. The application must derive the filter from verified identity and permissions, enforce it in the retrieval path, validate returned chunks, and test every fallback and cache path. Metadata is useful only when the query cannot be widened by an untrusted caller.
Is a separate vector database required for every customer?
Not always. Separate stores offer a clearer isolation boundary but add cost and operational work. A shared store can be appropriate when its authorization filters, service identities, deletion behavior and isolation tests are strong enough for the data and threat model.
Can the model decide whether a user may see a retrieved document?
No. Access belongs to the application's trusted authorization layer. The model can summarize content already authorized for the caller; its generated reasoning or refusal is not an access-control decision.
Does RAG prevent prompt injection?
No. Retrieved content can contain indirect prompt injection. Keep permissions outside the prompt, treat source text as untrusted, limit model capabilities and test malicious documents as a separate threat.
Related practical guides
Related issue guides
Sources and publication record
Draft prepared 27 September 2026; engineering, security and editorial review pending · Sources checked .
- LLM08:2025 Vector and Embedding WeaknessesOWASP Gen AI Security Project
- LLM01:2025 Prompt InjectionOWASP Gen AI Security Project
- OWASP Top 10 for LLM Applications 2025OWASP Gen AI Security Project
- Connect to SharePoint data sources with access control listsAmazon Web Services Bedrock Knowledge Bases
- Set up a knowledge base with a security configurationAmazon Web Services Bedrock Knowledge Bases
- Metadata filtering for knowledge basesAmazon Web Services Bedrock Knowledge Bases
- AWS SaaS Lens: Preventing cross-tenant accessAmazon Web Services
- AWS SaaS Lens: Testing multi-tenant SaaS reliabilityAmazon Web Services
- OWASP Cheat Sheet: AuthorizationOWASP Foundation
- OWASP Cheat Sheet: LoggingOWASP Foundation
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)National Institute of Standards and Technology