
Tenant Isolation on AWS: Follow the Data Beyond Login π
Learn how to isolate tenant data across PostgreSQL, background jobs, caches, and file downloads in a multi-tenant SaaS on AWS.
Series: Building SaaS on AWS: From Architecture to Operations (3 articles)
- 2Tenant Isolation on AWS: Follow the Data Beyond Login π This article
Tenant Isolation on AWS: Follow the Data Beyond Login
A user signs in successfully, opens a document URL, and receives another customer's file. Authentication worked. The application still failed at its most important boundary.
In a shared document portal, tenant isolation must survive every hop: HTTP requests, database queries, background jobs, object downloads, and support tools. This article walks through a proposed isolation design for fictional restaurant groups sharing one application.
Treat tenant context as verified server-side data
Suppose a user belongs to two restaurant groups. The request selects one group, but the server must verify that selection against the user's memberships and the action being attempted.
The resulting context might contain a user ID, tenant ID, permitted branch IDs, and an authorization decision. Construct it after validating the session. Pass it explicitly to business operations rather than reading an arbitrary tenant header inside each database helper.
AWS distinguishes tenant isolation from simply authenticating and authorizing users: access must also be constrained to the correct tenant's resources. AWS guidance: tenant isolation and authorizationΒ
For our portal, a useful rule is that a document lookup takes both a document ID and a verified tenant context. Returning βnot foundβ for inaccessible documents can avoid confirming another tenant's document exists, but the internal audit trail should still record the denied operation without exposing document contents.
Put a boundary in the database
In a pooled PostgreSQL design, application filters are necessary but easy to omit. Row-level security provides another enforcement layer. AWS recommends setting tenant context at runtime and applying RLS to tables containing tenant data. AWS guidance: PostgreSQL RLSΒ
The policy below illustrates the idea for an existing table with a UUID
tenant_id. It is not a complete database setup:1
2
3
4
5
6
7
8
9
10
ALTER TABLE documents ENABLE ROW LEVEL SECURITY;
ALTER TABLE documents FORCE ROW LEVEL SECURITY;
CREATE POLICY documents_tenant_scope ON documents
USING (
tenant_id = current_setting('app.tenant_id', true)::uuid
)
WITH CHECK (
tenant_id = current_setting('app.tenant_id', true)::uuid
);For each operation, start a transaction, set the verified tenant ID with a parameterized call to
set_config('app.tenant_id', $1, true), and execute the queries on that same transaction connection. The true argument makes the setting transaction-local. Avoid a session-wide tenant setting that could survive connection reuse.Use an application role without superuser or
BYPASSRLS privileges, keep migration credentials separate, and test using the actual application role. Table owners normally bypass RLS unless forced; privileged roles can still bypass it. AWS Database Blog: RLS behavior and caveatsΒ RLS does not verify who should be allowed to set the context. The trusted application must do that. Branch permissions also need their own enforcement inside the tenant boundary.
Check the boundaries outside SQL
For this portal, I would use tenant-prefixed S3 object keys, but treat the prefix as an organization convention until an access policy or trusted download service enforces it. A caller must not be able to choose an arbitrary object key for signing.
Cache keys need the same discipline. A cache entry called
dashboard:branch-01 can collide across tenants. Include tenant identity and any additional authorization dimensions that affect the response.Background jobs should carry a tenant ID and resource ID from a trusted producer. Before processing, the worker checks that the resource belongs to that tenant. Restrict queue producers with IAM; a tenant field in a message is not proof of authority by itself.
Support tooling deserves a separate access path with explicit scope and audit records. A hidden βadmin bypassβ in normal user code makes an otherwise careful design difficult to reason about.
Build a negative test matrix
Create two tenants and deliberately reuse branch names and document display names. Then exercise these cases:
| Attempt | Expected result |
|---|---|
| Tenant A requests tenant B's document ID | No document data returned |
| A request omits tenant context | Operation fails closed |
| A pooled connection serves A, then B | No context carries over |
| A job combines tenant A with B's resource | Worker rejects the mismatch |
| A download request supplies another object's key | No URL is issued |
| A user loses branch access | Subsequent requests honor the change |
Run these tests against detail pages, lists, exports, downloads, and worker paths. Testing only the visible dashboard leaves too many alternate routes unexplored.
Tenant isolation becomes maintainable when every data path has a named enforcement point and an observable failure mode. The next article applies that same discipline to retries and background processing.
Series: Building SaaS on AWS: From Architecture to Operations (3 articles)
- 2Tenant Isolation on AWS: Follow the Data Beyond Login π This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article