In a multi-tenant SaaS product, one query that forgets its tenant filter can show one customer another customer’s data. Code review won’t reliably catch it, and that one omission can become a breach notice. The fix is to stop relying on every developer remembering WHERE tenant_id = ? and to enforce tenant scope in one place that refuses to run a query without it.
Omnisnia is the multi-tenant CRM built by Nandeshou, written in Go with Gin, GORM and PostgreSQL (with pgvector for embeddings). Every customer organization is a tenant, and nearly every table holds tenant-owned data. This post explains how its data layer makes tenant isolation the default rather than something each query has to remember.
Why a hand-written filter is the wrong control
A tenant predicate written by hand has to be repeated in every query against every entity, by every developer, forever. That covers list endpoints, lookups by ID, bulk updates, deletes, exports and background jobs. The failure is silent. A missing filter does not throw an error; it returns more rows than it should, and a test that only checks “my tenant’s record came back” still passes.
In a multi-tenant data layer, scoping should be:
- Automatic for every tenant-owned model, including models added next year.
- Fail-closed, so a query that arrives without a tenant is refused rather than run unscoped.
- Explicit to bypass, so the rare operation that must cross tenants says so in a way a reviewer can find.
Callbacks on every operation
GORM lets you register callbacks that run before its built-in query, row, create, update and delete steps. Omnisnia registers one for each. Each callback first asks whether the model has a tenant column. If it doesn’t, as with users, who must be found at login before any tenant is known, it lets the query through. If the model is tenant-owned, the callback needs an active tenant from the request context, or it records an error that stops the statement.
A simplified version of the pattern:
var ErrNoTenant = errors.New("no active tenant")
func Register(db *gorm.DB) error {
cb := db.Callback()
if err := cb.Query().Before("gorm:query").Register("tenancy:query", scopeRead); err != nil {
return err
}
if err := cb.Delete().Before("gorm:delete").Register("tenancy:delete", scopeRead); err != nil {
return err
}
if err := cb.Update().Before("gorm:update").Register("tenancy:update", scopeUpdate); err != nil {
return err
}
return cb.Create().Before("gorm:create").Register("tenancy:create", stampCreate)
}
func scopeRead(db *gorm.DB) {
if db.Statement.Schema == nil || db.Statement.Schema.LookUpField("TenantID") == nil {
return // not tenant-owned data
}
tenantID, ok := activeTenant(db) // records ErrNoTenant when missing
if !ok {
return
}
db.Statement.AddClause(clause.Where{Exprs: []clause.Expression{
clause.Eq{Column: clause.Column{Table: clause.CurrentTable, Name: "tenant_id"}, Value: tenantID},
}})
}
Each operation adds something beyond the read filter:
- Create overwrites the record’s tenant with the active one. A caller who sends a different tenant ID in the request body still writes into their own tenant.
- Update adds the tenant filter and pins the tenant column, so an update cannot move a record into another tenant.
- Upsert gets its own guard. When a create carries an existing primary key, GORM’s
Save()can fall back to an insert-or-update on conflict, and that would overwrite a row regardless of who owns it. The create callback refuses that case when the key belongs to another tenant.
This design depends on the error, not the filter. A missing tenant does not mean “no filter”. It means the statement fails with an error the caller has to handle.
An explicit escape hatch for system work
Some work legitimately has no tenant: finding a user at sign-in, the first-admin bootstrap and schema migrations. Background jobs often need to list tenants before working inside each one.
For these, server-internal code marks the context as a system operation. Request handlers never create one, so a real request that arrives without a tenant is always refused. The system scope is the only way past the fail-closed check, and it is a separate, named call rather than a missing value:
func AsSystem(ctx context.Context) context.Context {
return context.WithValue(ctx, systemKey{}, true)
}
// A background job lists tenants as the system, then works inside each one.
func processAll(ctx context.Context, db *gorm.DB) error {
var tenants []Tenant
if err := db.WithContext(AsSystem(ctx)).Find(&tenants).Error; err != nil {
return err
}
for _, t := range tenants {
scoped := WithTenant(ctx, t.ID)
if err := processTenant(scoped, db); err != nil {
return err
}
}
return nil
}
The Omnisnia event relay follows that shape. It lists tenants under the system scope, then handles each tenant’s queue under that tenant’s scope, so everything downstream sees the right tenant.
Three outcomes for a tenant-owned query: scoped, explicitly system, or refused.
Because the bypass is a single named function, every use can be found with one search and reviewed. That is the point of making it explicit: crossing the tenant boundary is a deliberate, visible act in the code, never an accident of omission.
Two kinds of path deserve their own review in any design like this. Models without a tenant column pass through the callbacks by design, so each one should be a conscious decision. Raw SQL skips model-aware callbacks entirely, so a raw query that touches tenant data has to write its tenant predicate itself and refuse to run without a scope.
As general good practice, you can also log each system-scope call with a short reason. Then the audit trail records why the boundary was crossed, as well as where in the code.
Tests that prove the boundary
Run these tests against real PostgreSQL rather than a mock, because the point is that the generated WHERE clause actually reaches the database. The suite seeds rows in two tenants and asserts that:
- a create that asks for another tenant lands in the active tenant;
- a list, and a fetch by primary key, never return the other tenant’s row;
- a cross-tenant update or delete affects zero rows and leaves the target untouched;
- a cross-tenant upsert is refused;
- a query, create or delete with no tenant at all fails with the no-tenant error and returns nothing.
The last of these matters most, because forgetting to scope is the failure the whole design exists to prevent:
func TestUnscopedAccessIsRefused(t *testing.T) {
db := openTestDB(t)
var rows []Client
err := db.WithContext(context.Background()).Find(&rows).Error
if !errors.Is(err, ErrNoTenant) {
t.Fatalf("unscoped list: err = %v, want ErrNoTenant", err)
}
if len(rows) != 0 {
t.Fatalf("unscoped list returned %d rows, want 0", len(rows))
}
}
Row-level security as a second layer
PostgreSQL row-level security (RLS) can enforce the same rule inside the database. According to the PostgreSQL documentation (version 18, current as of October 2026), a table with RLS enabled and no policies uses a default-deny policy, so no rows are visible or can be modified. That makes RLS fail-closed by design. The usual multi-tenant pattern sets the tenant per transaction and checks it in a policy:
ALTER TABLE clients ENABLE ROW LEVEL SECURITY;
ALTER TABLE clients FORCE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON clients
USING (tenant_id = current_setting('app.tenant_id', true)::bigint)
WITH CHECK (tenant_id = current_setting('app.tenant_id', true)::bigint);
-- At the start of each transaction, from the application
-- (the third argument makes the setting transaction-local):
SELECT set_config('app.tenant_id', '<tenant-id>', true);
If the setting was never made, current_setting(..., true) returns null and no row matches. On a pooled connection where it was set in an earlier transaction, it reads as an empty string and the cast fails. Either way, a forgotten tenant returns nothing. Use set_config() rather than SET LOCAL when the value comes from a bind parameter, because SET does not accept one.
RLS has its own exceptions. Superusers and roles with BYPASSRLS always bypass it. Table owners do too, unless the table uses FORCE ROW LEVEL SECURITY. So the application should connect as a role that owns nothing. In exchange, RLS covers what ORM callbacks cannot: raw SQL, ad hoc scripts, and any future code path that talks to the database directly.
We treat the two as layers, not alternatives. ORM callbacks give fast, testable feedback in application code and clear errors at the point of failure. RLS, where you add it, ensures that a missed path still returns nothing.
How we can help
We review multi-tenant SaaS platforms for tenant isolation across the ORM, raw SQL, background jobs and the database. Where it helps, we add fail-closed enforcement and tests that prove it. Talk to us if you want an outside check on your tenant boundary.