Skip to main content

Directory module

Identity, organisations, users, roles, and permissions. Every other module reads identity.* for the principal performing a request — this module is the foundation everything else depends on.

  • Source: apps/directory/
  • Schema: identity
  • Project Tracker prefix: DIR-*
  • Hosting: root zone in the multi-zone topology — constellation.planetb2b.com is served by Directory, which rewrites /projects/* and /catalog/* to the sub-zones.

1. Purpose​

Directory is the platform's identity backbone. It answers who is asking, and what are they allowed to do for every request in every module — no search, project update, or wiki read happens without Directory's say-so.

Key capabilities

  • Tenants, member organisations (with hierarchy and verification workflows), and users
  • Per-tenant roles and permissions — hybrid RBAC/ABAC, evaluated on every request
  • Organisation certifications with expiry tracking, and capability tags used for supplier matching
  • Tenant policy glue: locale, time zone, currency

Who uses it

  • Tenant administrators managing organisations, users, and roles
  • Compliance staff keeping supplier certifications current
  • Every other module, implicitly, on every authenticated request

For product context, see What is Constellation? and the Tenant types reference. The rest of this page is the technical reference.

Directory owns who in Constellation: tenants, organisations, users, memberships, roles, permissions, and the small slice of policy-glue (locale, time zone, currency) that the platform reads in every request. Every authenticated request flows through Directory's middleware — createConstellationAuthMiddleware — even when the destination is a sub-zone like Catalog or Project Tracker. Tenant locale resolution, the active-membership lookup that withTenantAuth performs, and the user/role/permission tables every other module joins to are all owned here.

2. Component diagram (C4 L3)​

Call direction is strict: API Route → Tool → Service → Repository. Routes never call services or repositories directly. Workflows and event handlers must call tools or services, never repositories. See route-wrapping and module-isolation.

3. Schema​

identity — owned by Directory. Other modules may read it directly (raw SQL, no Prisma cross-schema joins, per module-isolation) but may not issue direct DML against it.

Where another module needs a transactional identity mutation, Directory exposes it as a SECURITY DEFINER function in the identity schema, created and versioned by a Directory migration. The caller invokes it as ordinary SQL, so the write joins the caller's own transaction — which is what makes an atomic cross-schema operation possible at all (invitation acceptance, for instance, must commit the membership and the invitation status together, and an HTTP call cannot enlist in the caller's transaction).

The boundary is about ownership, not privilege stripping: the shared runtime role still holds direct DML on identity.* (Directory itself needs it), so what keeps a consuming module's identity mutations routed through Directory-owned, Directory-versioned code is convention plus CI, not revoked table grants. The CI gate is npm run check:identity-writes, and it scans apps/project-tracker/{src,prisma,scripts} and apps/wiki/{src,scripts,evals}. The wiki's local seed scripts are the one sanctioned exception: an audited identity bootstrap bounded by ADR-042. Project Tracker's seeds and operator scripts are a grandfathered, pinned inventory (ADR-057). Catalog issues no identity writes at present; a module that starts to should be added to the gate's scanned roots in the same change. Each function re-establishes the tenant guard that RLS cannot apply to an RLS-exempt owner, and each derives its EXECUTE grant from the full set of privileges it exercises, so calling it can never achieve more than the caller could do directly. That guarantee is time-bounded: the conjunction is evaluated when the migration RUNS, so the grant is never ISSUED to a role that could not already do it, but it does not follow the grantee afterwards — narrowing a role's identity.* access must re-apply these migrations (which revoke every non-owner and re-derive) or revoke the function grant, in the same change. That derivation is preceded by a full EXECUTE reset — every non-owner grantee is revoked, not just PUBLIC — because ALTER DEFAULT PRIVILEGES grants the runtime role EXECUTE on each function the moment it is created, and a role-specific grant left in place would let a role that fails the privilege conjunction reach the function anyway. Current functions: grant_invitation_access, provision_user, merge_user_metadata, update_organisation_profile, bootstrap_tenant. Rationale and constraints: ADR-025.

The write boundary has a read counterpart: pre-tenant decision points that tenant-scoped RLS would blind. identity.resolve_member_tenant_by_slug (migration 045, PLT-960) is the newest of the confined identity_bootstrap-owned resolver family — given a slug and a user it returns the caller's own ACTIVE-membership ACTIVE tenant plus whether any roles exist, which is what lets Project Tracker's organisation-creation pre-check resume a half-provisioned tenant instead of burning its slug. A foreign or free slug returns nothing, so the constellation_app EXECUTE grant discloses nothing about tenants the caller is not in. Both halves of the boundary — write-function ownership and the resolvers' owner/grant/policy shapes — are pinned by scripts/verify-identity-boundary.ts.

Related seeding semantics (PLT-960): ensureDefaultRoles serializes per tenant with a transaction-scoped advisory lock, because the resume path makes concurrent seeders reachable and the roles uniqueness constraint deliberately does not deduplicate canonical GLOBAL rows (scope_id IS NULL) — unserialized, both writers pass the find-then-create and leave duplicate canonical roles that only manual repair removes. The seeding transaction budgets maxWait for a winner's full ceiling, and the seed endpoint carries a route-specific maxDuration for the serialized wait.

4. Entities​

AggregatePurpose
tenantsTop-level isolation boundary. Owns default_locale, enabled_locales, time_zone, currency.
organisationsMember companies inside a tenant. Hierarchical via parent_id.
usersPrincipals. tenant_id is the user's "home tenant". clearance holds the platform ladder (§4.1).
user_tenant_membershipsActive-tenant resolution. withTenantAuth validates membership before any tenant-scoped call.
roles, permissionsRBAC. role_permissions and user_roles are the join tables.
tenant_appsPer-tenant app enablement — the tenant half of the app-entitlement seam (§12).
api_keysIssued programmatic credentials, resolved by jti on every request. ephemeral separates machine-minted OAuth bridge tokens from user-issued keys (§13).
organisation_certificationsCompliance state with expiry → drives the credential.expired event.
organisation_competence_domainsCapability tags for organisation matching (used by Catalog supplier offer matching).

4.1 Stored clearance vocabulary​

identity.users.clearance is the stored clearance on a user's Directory profile. The ladder it draws from is owned by @constellation-platform/auth-core and by public.clearance_rank(), which RLS policies evaluate inside the database.

Two things read a clearance, and they are not the same value. Request-time authorization uses the validated clearance claim on the caller's JWT, supplied by the auth provider (hasSufficientClearance(jwt.clearance, …) in auth-core, and the RLS session GUC). Nothing synchronises this column into issued tokens, so updating the row does not change what the token-driven path allows.

A second class of consumer reads the stored row directly. Today there is exactly one that makes an access decision from it (Directory's own admin user page merely displays the value): Project Tracker's coordinator-consult masking (PLT-212, coordinator-consults.service.ts), which loads the viewer's identity.users.clearance and redacts consult excerpts classified above it. For that surface, updating the row can change what the user sees — with a caveat worth knowing: that read is tenant-scoped by RLS against the user's home tenant, so for a viewer acting through a membership in a different tenant it can return no row, and the viewer is then treated as UNCLASSIFIED. That direction over-masks rather than over-discloses, but it means the stored value is not always the value that surface uses.

So: same vocabulary, two independently-sourced values. When you change one, say which. A new consumer should prefer the JWT claim unless it genuinely needs the stored profile value, and should say why in its own code.

The stored vocabulary is five values, ascending:

UNCLASSIFIED < INTERNAL < RESTRICTED < CONFIDENTIAL < SECRET

INTERNAL was inserted at rank 1 (PLT-497 for the platform ladder, PLT-498 for this column). The relative order of the original four is deliberately unchanged: that order is what every stored classification means relative to every stored clearance, so reordering it would silently reinterpret live data in both directions with no accompanying data migration.

Enforcement is a single CHECK constraint, users_clearance_check. Widening it means dropping and restating the complete list — PostgreSQL cannot add a value to a CHECK, and a second constraint would leave the older, narrower one active and the new value rejected.

Directory stores and reads the value; it does not yet expose a way to assign one. Every user is created UNCLASSIFIED, and the admin user page renders the clearance read-only.

5. Domain events​

Published from src/server/events/directory.events.ts via the transactional outbox under the directory.* namespace. The full payload schemas live in @constellation/contracts and are catalogued in the Domain events index.

directory.tenant.created, directory.tenant.updated, directory.organisation.created, directory.organisation.verified, directory.organisation.verification.requested, directory.organisation.verification.rejected, directory.organisation.verification.completed, directory.user.created, directory.user.suspended, directory.user_credential.updated, directory.role.assigned, directory.role.unassigned, directory.role.permission_changed, directory.role.deleted, directory.credential.expired, directory.credential.expiry.approaching, directory.credential.expiry.imminent, directory.qualification.updated.

6. Public API​

The Directory API reference is not yet wired into this site — adding Zod→OpenAPI to apps/directory/ is tracked under INF-27. Illustrative routes:

  • POST /api/organisations — create a new organisation in the active tenant. Worked-example of the full request lifecycle is in Request lifecycle.
  • GET /api/users/:userId — read a user with role assignments.
  • GET /api/auth/memberships — bootstrap call after login; resolves the user's set of { tenantId, organisationId } pairs for the org switcher, alongside a status of "ok" or "degraded". status is load-bearing, not decorative (INF-392): the endpoint answers 200 in both cases, and "degraded" means an identity read failed, so the accompanying empty list establishes nothing and must never be rendered as "this user belongs to no organisations". Only "ok" may be trusted — treat an absent status (a deployment predating INF-392) as unknown rather than healthy, and any unrecognised value as non-authoritative. The lookup resolves under the caller's effective identity id, so an invited user whose auth subject differs from their identity row still gets their real memberships.
  • GET /api/organisations/switchable — organisations the signed-in user can switch into. Distinct from GET /api/organisations, which is the tenant-wide paginated admin list behind the /organisations management page. See Which organisations a user can switch into below.
  • GET /api/search?q=<query> — tenant-scoped search over organisations + users (DIR-85, PT-625 Stage 1), returning grouped hits for the shared GlobalSearch bar. Synthetic agent accounts (agent+<class>@constellation.local) are excluded from user hits; each group is capped by limit (1–25, default 10).
  • GET /api/v1/users/lookup?q=<query>&limit=<n> — users eligible to own a tenant-scoped resource in the acting tenant (PLT-518), for entity pickers such as the wiki page-owner field. Returns id, name, displayName, avatarUrl and a Directory-derived isAgent flag, and never an email. See Owner-candidate lookup — it is deliberately not the same population as /api/search.
  • POST /api/auth/api-keys / GET /api/auth/api-keys / DELETE /api/auth/api-keys/:id — issue, list, and revoke the caller's own API keys (DIR-1). Issue and revoke are session-only (assertSessionCaller): an API key may not mint or revoke another — that is the no-chaining property. Listing is open to any authenticated token kind, since reading one's own key metadata is not an escalation. The list returns the caller's user-issued keys — active, expired and revoked — but excludes machine-minted OAuth bridge tokens (§13). This is the endpoint the Project Tracker settings page reads.

Administrative read authorization​

Administrative reads require the existing directory.read permission at the tool boundary, including users and batch users, membership details, credentials, organisations and their certifications/competence domains, roles, permissions, tenant details, and Directory search. Versioned and unversioned adapters use the same tools. Tenant listing additionally retains its platform-operator check. Audit reads retain their separate audit permissions and scope checks.

Tenant detail allows an administrative reader to inspect their own acting tenant. A different requested tenant ID additionally requires platform-operator authority before the resource lookup, even when the database bypasses RLS. The operator's existing RLS context is retained. Role detail always filters by the acting tenant in SQL before loading the role's permissions.

Migration 015 seeds directory.read and grants it to legacy administrators; canonical Admin roles select the full permission catalog. No new read permission is needed. Auth bootstrap, the caller's own memberships/keys/sessions, switchable organisations, and the minimal owner picker keep their existing narrower rules.

User list/detail reads omit clearance, nationalCaveats and authProviderId unless the caller also holds the existing user.update permission. Credential responses, including creation, updates and verification, apply the same additional permission to documentUrl. The shared user SQL projection and row mapper default to safe fields; internal authentication and mutation lookups explicitly retain the full record. /api/users/batch always returns only id, name, email and avatarUrl, even for administrators.

Credential PATCH/DELETE/verify and membership role assignment/removal perform ownership pre-reads through separate helpers authorized by the mutation action. A principal with only the relevant mutation grant does not need directory.read.

Owner-candidate lookup​

GET /api/v1/users/lookup answers one question: who may own a tenant-scoped resource in the tenant I am acting in? It exists because no other read surface can answer it.

Eligibility — the active-membership-or-home-tenant rule. A user is a candidate when either they hold an ACTIVE membership in the acting tenant, or their home tenant (identity.users.tenant_id) is that tenant. These are the two paths wiki.validate_page_owner() applies at write time (wiki migration 022, PLT-308), mirrored here on purpose: a picker narrower than the save-time guard hides legitimate owners, and a picker wider offers candidates the write then rejects.

The candidate set is a subset of what that guard accepts, not an exact match, by two deliberate narrowings: soft-deleted users are excluded (they are not candidates for new ownership), and tenant liveness is required on both paths, where the guard applies it only to the membership one.

The second path is not redundant. It covers system and synthetic principals that own content but hold no membership row. The first path is the multi-org one, and it is the reason this endpoint had to exist: identity.users is FORCE-RLS keyed on the home tenant, so a colleague who belongs to your tenant but whose home tenant is another one is invisible to every tenant-scoped read — including GET /api/search and GET /api/users/batch.

It is not /api/search, and the difference is deliberate.

GET /api/searchGET /api/v1/users/lookup
Populationhome tenant onlyACTIVE membership or home tenant
Synthetic agent accountsexcludedincluded — they own pages
Email in the responseyesno — only the derived isAgent flag
Empty queryrejected (q is required)browses the first page

Two resources rather than one mode flag, because a single response shape carrying two different visibility rules is how a caller ends up trusting the wrong one.

isAgent is derived inside Directory from the synthetic-agent email convention (agent+<class>@constellation.local, INF-113/INF-125) and computed in SQL, so the address never leaves the database.

Authorization for the owner picker is tenant membership via authedRoute. It remains available to ordinary members for collaboration. Its wider population is accepted because each candidate has an ACTIVE membership in the acting tenant or is an ACTIVE home-tenant principal, and the projection contains only identity, display names, avatar and the derived agent flag. It returns no email or personnel security attributes. Administrative /api/search requires directory.read (DIR-88); that restriction does not apply to this intentionally minimal picker.

Query semantics. q is optional; absent or blank browses. When present it matches name and displayName (never email), prefix matches sorting ahead of substring matches, then name ASC with an id tiebreak so a capped page is deterministic. limit is 1–25, default 10.

Implementation. The eligibility rule lives in the Directory-owned SECURITY DEFINER function identity.lookup_owner_candidates (Directory migration 042), owned by the confined identity_bootstrap role. It refuses to read any tenant other than the caller's established app.tenant_id session context, so holding EXECUTE on it is not a cross-tenant enumeration capability. The shared wire contract is OwnerCandidateSchema in @constellation/contracts.

Which organisations a user can switch into​

This is a property of the user's profile, not of the app they are looking from: if a user can reach an organisation, they can reach it from Directory, catalog, wiki and project-tracker alike. All four zones therefore answer it from one shared implementation — listSwitchableOrganisations in @constellation-platform/auth-next — even though each serves it from its own route, because every zone must also work as a standalone preview deployment where Directory is not reachable at the root.

Two properties are easy to get wrong and are worth stating explicitly:

  • Membership status is not enough — and the gate is in the resolver, not the list. Both list_active_memberships and resolve_active_membership require the membership's tenant to be ACTIVE, not just the membership row (migration 032, PLT-477), so an organisation in a SUSPENDED or DECOMMISSIONED tenant is not offered here and cannot be entered by any other id source either. The helper itself no longer filters tenant status: with the resolvers gated, its own clause could not fire, and a per-request query that reads as a security control while doing nothing is worse than its absence.

    Superseding PLT-476's acting-tenant exemption. Until PLT-477 this list carried an exemption for the caller's own acting tenant. That was not a design choice — resolve_active_membership still resolved a membership in a disabled tenant, so hiding the current organisation would have emptied the acting-tenant slot and made the first entry a FOREIGN tenant the switcher would reseed x-active-org from, i.e. a silent tenant switch. Gating the resolvers removed the condition the exemption existed for, and it was deleted. It must not be re-added: doing so would re-admit an acting tenant the auth boundary now refuses. Project-tracker used its own copy of this logic until PLT-483 and applied no tenant gate whatsoever.

  • The acting organisation affects ORDERING, not membership. The list spans every tenant the caller belongs to; the acting organisation decides which entry sorts first (and seeds x-active-org when no cookie is set).

The rationale and the deferred alternative — a single Directory-owned endpoint consumed cross-zone — are recorded in ADR-023.

Creating an organisation, and why the switcher offers it in one zone only​

Two different things are both called "creating an organisation", and only one of them produces something you can switch into.

Directory POST /api/organisations (/organisations/new)Project Tracker POST /api/organisations
Createsone organisation inside your current tenanta new tenant + its organisation
Grants you a membershipnoyes, ACTIVE
Appears in the switch list abovenoyes, immediately
What it is forrecording an organisation in your tenant's directory — a partner, supplier or sub-tier entity, with a type, an optional parent, and a verification workflowstarting a new workspace you are the Admin of

Because the switch list is membership-scoped, an organisation created through Directory's form never appears in it. That is not a defect in the form — it is what a tenant-local register is — but it does mean the form cannot back a switcher action. So Directory, Catalog and Wiki show no "Create Organisation" entry in the switcher; Project Tracker does (PLT-627). Directory's form remains available and unchanged, from the Organisations management page.

It also cannot be repaired by having that form grant the creator a membership. identity.user_tenant_memberships is UNIQUE (user_id, tenant_id), so a caller who already belongs to an organisation in that tenant can hold no second membership there, and repointing the existing row would move them out of their current organisation. One membership per tenant is the reason "switch organisation" is in practice "switch tenant", and the reason a create action in the switcher can only mean "create a new tenant".

7. Layers + call direction​

LayerPathMay import from
API Routessrc/app/api/Tools only
Toolssrc/server/tools/Services, Policies, Events
Servicessrc/server/services/Repositories, Policies, @constellation-platform/db
Repositoriessrc/server/repositories/@constellation-platform/db (Prisma client + raw SQL)
Policiessrc/server/policies/@constellation-platform/auth-core
Workflowssrc/server/workflows/Tools or Services (never repositories)
Eventssrc/server/events/@constellation-platform/events publish() + outbox

Enforced at PR time by scripts/check-route-wrapping.ts — invoked via npm run check:routes, which scans apps/directory/src/app/api and apps/catalog/src/app/api. Every withAuth must be paired with withTenantAuth. Plus ESLint import boundaries.

8. Code entry points​

9. Known exceptions / pitfalls​

  • Directory is the root zone in production. It serves constellation.planetb2b.com and rewrites /projects/* → constellation-platform.vercel.app and /catalog/* → constellation-catalog.vercel.app. Multi-zone wiring lives in apps/directory/next.config.ts.
  • Tenant locale policy lives here. tenants.default_locale ∈ tenants.enabled_locales is enforced by a DB CHECK; the form validates client-side too. users.preferred_locale = NULL means "inherit tenant default". The locale resolver wires into the platform auth middleware via the localeResolver option — only Directory wires it today; Catalog and Project Tracker stay on the existing path until they opt in. See Multilanguage (i18n).
  • No Prisma models for cross-module identity reads. Other apps that need to read identity.users / identity.organisations use raw SQL via repositories like Project Tracker's IdentityUserRepo — never Prisma cross-schema joins.
  • (admin) route group is platform-operator-only. Tenant admins use /settings, not /tenants/[id]. The two pages enforce different RBAC scopes.

Pre-tenant-context identity resolution (RLS bootstrap)​

/api/auth/me and the org switcher must read identity.users / identity.user_tenant_memberships before app.tenant_id is set — the membership row is precisely what validates the tenant. Under a non-bypass runtime role those tables are filtered to zero rows by FORCE RLS, so the lookups go through SECURITY DEFINER functions owned by the NOLOGIN identity_bootstrap role: resolve_active_membership / list_active_memberships (migration 016), resolve_identity_user (migration 024, DIR-72), and resolve_principal_liveness (migration 043, PLT-833) — the principal-check liveness read, which has to see a user row owned by a tenant other than the acting one, because an API key minted while acting in a secondary tenant carries that tenant as its claim. The bootstrap RLS policies are meant to match only inside those functions, and they are deliberately not uniform. users_bootstrap_select and memberships_bootstrap_select (migration 016) use USING pg_has_role(current_user, 'identity_bootstrap', 'USAGE'): they open for any session whose current role inherits identity_bootstrap's privileges — which is exactly why the invariant below keeps constellation_app out of that role. tenants_bootstrap_select (migration 024) uses USING current_user = 'identity_bootstrap': it opens only when the session's current role is identity_bootstrap — inside a SECURITY DEFINER function it owns, or after an explicit SET ROLE identity_bootstrap, which requires a membership carrying the SET option (the owner-plumbing migrations 029b/039a/041/043 grant that to the migrator only transiently: acquire, transfer ownership, release). Inherited membership alone does not unlock tenant rows. Migration 043 pins all three quals as they are stored, because that distinction is security-sensitive.

Both membership resolvers require the membership's tenant to be ACTIVE, not just the membership row (migration 032, PLT-477). See Tenant status semantics below.

resolve_identity_user is id-first with a fail-closed email fallback: it matches by primary key first and only falls back to email when exactly one non-deleted row matches. identity.users.email is unique only per tenant (and the synthetic agent emails exist in every tenant), so a shared address resolves to no row rather than hydrating an arbitrary tenant.

Since migration 051 (PLT-1010) it also returns clearance, so /api/auth/me reports the tier the caller's identity.users row actually carries. Before that, no resolver returned the column and every interactive session resolved as UNCLASSIFIED — the tier a classification gate then compared against, whatever the row said. The value is validated against the ClearanceLevels ladder on the way out (parseViewerClearance in @constellation-platform/auth-core) and fails closed to UNCLASSIFIED. An off-ladder viewer clearance is not itself a fail-open hazard — clearance_rank() answers -1 for it and the RLS comparison then denies every storable classification — but it is a tier nobody asserted, and normalising it once here keeps ClearanceLevel true at runtime instead of leaving each consumer to interpret an unknown string. (The direction that genuinely fails open is the RESOURCE classification; see Glossary.)

  • Invariant: the app runtime role (constellation_app) must NOT be a member of identity_bootstrap. That membership makes the bootstrap "see-all" policies fire for every direct app query, leaking users/memberships cross-tenant (DIR-72; the residual INF-60 grant was revoked by migration 028 — DIR-79).
  • Runtime role (INF-60, cut over). All of the above only enforces isolation once the app connects as a non-bypass role: under a rolbypassrls = true role PostgreSQL silently disables every policy, even with FORCE ROW LEVEL SECURITY, and these functions become a no-op for isolation. Production now connects as constellation_app (NOBYPASSRLS), and that is enforced rather than assumed — all four apps call assertNonBypassRole() from @constellation-platform/db in src/instrumentation.ts on startup, so a DATABASE_URL repointed at a BYPASSRLS role fails the boot instead of silently serving cross-tenant data. DIRECT_URL deliberately stays on a BYPASSRLS role, because migrations need it. See Deployment.

Acting-tenant authority hydration (PLT-717)​

withTenantAuth resolves the acting tenant through identity.resolve_membership_authority (migration 041) — the same membership predicate, dual-id lookup and tenant-liveness gate as resolve_active_membership, extended with the acting tenant's authority: role_names (DISTINCT names of the user's GLOBAL-scope identity.user_roles grants, with the roles join bound on both ur.tenant_id and r.tenant_id) and has_canonical_admin (a granted role whose complete persisted identity is the canonical Admin: name Admin, scope GLOBAL, unscoped).

And clearance (migration 051, PLT-1010). It is not a per-tenant fact — the column belongs to the user, so the value is the same whichever tenant the caller is acting in — but it rides this lookup because this is the one database call the tenant-auth path already makes per request, so carrying it costs a column rather than a round trip. withValidatedTenant writes it onto ctx.user.clearance in the same construction that rewrites tenant_id, org_id and roles, which is what makes a classification gate compare against the principal's real tier. The join is keyed on the resolved membership's user_id, so it can only ever be the clearance of the principal whose membership was just validated.

Precedence: the row fills an absent claim; a present, ladder-valid claim is left alone. In production the two rules are indistinguishable, because nothing mints a clearance claim — neither Supabase app_metadata, nor buildMockPlatformJwt, nor Directory's API-key issuance — so every real session and every real API key receives the row. The claim channel stays available to a deliberately-minted credential, and applyOptionalClaims validates it against the same ladder. If Directory ever starts minting one, it must be derived from this column.

Since migration 044 (DIR-118) the schema binds that pair too: identity.user_roles (role_id, tenant_id) is a composite foreign key to identity.roles (id, tenant_id) (ON DELETE CASCADE, NOT DEFERRABLE, alongside the original role_id → roles(id) key), so a tenant-A assignment naming a tenant-B role is refused at the INSERT with SQLSTATE 23503 — in bypass-capable contexts too (seeds, repair scripts, the privileged test harnesses), where RLS is inert. The both-sides join in every reader stays as defence in depth for databases that have not applied 044 yet. Above the schema, RoleService.assignRoleToUser resolves roles with an explicit acting-tenant predicate and returns the same ValidationError (HTTP 400) for a foreign or nonexistent roleId before any side effect (DIR-76). The response does not reveal cross-tenant existence, and the FK remains the final backstop.

ctx.user.roles is then rewritten alongside tenant_id/org_id, per token kind:

tokenKindActing = attested claim tenantActing ≠ attested claim
session (or undefined)DB-hydratedDB-hydrated
api_keyclaim roles unchangedDB-hydrated
internal, legacy_api_keyclaim roles unchangedclaim roles unchanged

DB-hydrated = the grant names, plus the TENANT_ADMIN sentinel when the canonical Admin is held (only the DB can prove canonicity — a custom tenant role merely named Admin never buys blanket access), plus the claim's platform-operator names (PLATFORM_ADMIN / platform_admin, the one deliberately cross-tenant authority), plus the user baseline.

Consequences worth knowing:

  • A session's flat app_metadata.roles claim is no longer a tenant authorization input. Before PLT-717 nothing populated that claim, which locked every normally-provisioned user out of wiki writes (INF-333) — and had it been backfilled, a tenant-less admin claim would have granted admin in every tenant the user could switch into (INF-334). Grant roles through Directory (identity.user_roles), scoped to the tenant that granted them. Hydration re-reads the grants per request, so the role names update on the next request — but permission sets evaluated through auth-core's in-process cache (Directory's canAccess) can lag up to PERMISSION_CACHE_TTL_MS (default 5 minutes) per route instance, because no event dispatcher runs in Directory's serving processes to deliver the invalidation. This matters most for revocation (DIR-115).
  • API keys keep exactly the authority Directory issued into them in their attested issuing tenant (DIR-87); acting in a different tenant via x-act-as-org hydrates from that tenant's DB grants instead, so tenant-specific key authority cannot ride across a switch.
  • Fail-closed: a resolver error (including the function missing because migration 041 has not been applied) denies with a 403 and a server-side error log — apply the migration before deploying code that calls it.

Tenant status semantics​

identity.tenants.status is one of ACTIVE, SUSPENDED, DECOMMISSIONED. The ladder is a lifecycle distinction, not an access tier — any non-ACTIVE status is a full deny at the auth boundary, and the specific non-ACTIVE value carries no access meaning.

StatusAccess at the auth boundaryReversibleData
ACTIVEfull—live
SUSPENDEDnone (403)yes — flipping back to ACTIVE restores it exactlypreserved untouched
DECOMMISSIONEDnone (403)terminal by conventionpreserved but abandoned

Enforced in three places, all with the same predicate (PLT-477):

  • identity.resolve_active_membership and identity.list_active_memberships (migration 032) — the resolvers withTenantAuth and every other membership consumer run. This gates the tenant a request is scoped into, whichever id source named it: the x-active-org cookie, the x-act-as-org header on the API-key path, or the JWT tenant_id claim.
  • createPrincipalCheck in @constellation-platform/auth-next — returns TENANT_DISABLED. This gates the attested claims.tenant_id, which is not always the tenant finally scoped into, so it is a complement to the above rather than a substitute. (Historical note: apps/wiki wired no principal check until INF-372; all three zones now do, on both bearer paths.)
  • identity.mcp_refresh_principal_live (migration 029) — refuses to re-mint an MCP token for a disabled tenant.

Read-only-but-visible for SUSPENDED is deliberately not implemented. A read-only tier is not something a membership resolver can express — it would need a capability threaded onto PlatformJWT, honoured by every mutating route in four modules, and enforced at the DB layer to be worth anything. Shipping the deny first does not foreclose it.

Suspending a tenant now revokes access at the auth boundary on the caller's next request — every tenant-scoped route reached through withTenantAuth, plus PT's own acting-org resolution. (Access reached by addressing a RESOURCE directly by id is gated on the paths PLT-477 audited; the sweep of the rest is PLT-638.) Before PLT-477 it did not: only revoking every user_tenant_memberships row did, which is a different and destructive operation. A user whose only membership is in a disabled tenant lands in the pre-existing zero-membership state — authentication still works (/api/auth/* is a public path) and no x-active-org cookie is seeded from the dead tenant; only tenant-scoped routes 403. Recovery is an operator action with no data loss.

10. MCP OAuth authorize endpoint (DIR-75)​

Directory acts as the identity provider for the MCP OAuth authorization code grant flow. This enables Claude Connectors, ChatGPT OAuth mode, Codex, and other MCP clients to authenticate users and receive identity claims via a standards-compliant OAuth handshake.

Endpoints​

MethodPathPurpose
GET/api/auth/mcp/authorizeBrowser-facing authorize redirect handler. Validates the redirect_uri client-aware (INF-286) before any session check: the first-party pt CLI client (constellation-pt-cli) may use RFC 8252 loopback redirects (http://127.0.0.1:<port>/callback or http://[::1]:<port>/callback — loopback literals only, any port); every other client stays on the MCP_REDIRECT_ORIGIN_ALLOWLIST exact-origin check. Writes mcp_authz_pending signed cookie; redirects to /login?redirect=/api/auth/mcp/authorize if unauthenticated.
POST/api/auth/mcp/authorizeConsent form submission. Validates CSRF double-submit cookie; on Allow issues the authorization code; on Deny records a denial audit entry.
GET/mcp/authorizeConsent page. Server component that reads the mcp_authz_pending and mcp_csrf cookies (both set by the GET Route Handler) and renders the consent UI (client_id, redirect_uri host, scopes, Allow/Deny) with the CSRF nonce embedded in the form.
POST/api/internal/auth/mcp-tokenServer-to-server only. The MCP AS (INF-164) sends this to exchange the one-time code for identity claims. Protected by X-Mcp-Client-Secret (constant-time check); no browser session required.
POST/api/auth/mcp/cli-tokenPublic client (INF-286). Token endpoint for the pt CLI / local MCP (client_id: constellation-pt-cli). No client secret: authorization_code is PKCE-authenticated (S256, verified inside the SECURITY DEFINER consume), refresh_token/revoke by possession of the rotating refresh token. Responses set Cache-Control: no-store + Pragma: no-cache (RFC 6749 §5.1).

Active CLI sessions (DIR-104)​

Self-service management of the refresh-token families above: one entry per npx pt login (no machine identifier is stored, so two logins from one machine are two sessions), letting a user end a session on a machine they no longer control. Both endpoints accept browser session tokens only (tokenKind === 'session') — a CLI-minted bridge token must not be able to enumerate or end the sessions of the credential chain it belongs to, the same "no chaining" property DIR-1 applies to API-key management.

MethodPathPurpose
GET/api/auth/cli-sessionsThe caller's own live sessions, aggregated from identity.mcp_cli_refresh_tokens by family_id: createdAt (consent), lastRotatedAt (newest rotation), expiresAt (absolute family expiry — rotation copies it rather than extending it) and rotationCount. Never hash material — migration 030 column-restricts the runtime role's SELECT so token_hash / rotated_from are unreachable.
DELETE/api/auth/cli-sessions/:familyEnds one session. Calls identity.revoke_mcp_cli_refresh_family and, in the same transaction, revokes that family's live bridge access tokens. Returns { familyId, refreshTokensRevoked, accessTokensRevoked }. A family that is unknown, already ended, owned by a tenant peer, or the caller's own in a different tenant all return 404 — never 403, so the endpoint is not an existence oracle.

identity.api_keys.cli_family_id (migration 035) is what makes the second write possible: mintPtAccessToken stamps it on the backing row for the two CLI paths (initial exchange and refresh), leaving it NULL for user-issued keys and the connector paths, which belong to no family. Without that link, revoking a bridge token never ended the session — the CLI minted a replacement at its next refresh — and revoking the family alone left the current access token usable for the rest of its TTL.

Propagation is the platform-standard ~30 seconds (the 30 s principal-liveness cache in createPrincipalCheck), not instant. A CLI access token is also now capped to its family's absolute expiry, and a refresh in the family's final ~44 s is refused as invalid_grant, so a bridge token can never outlive the session that produced it and become unrevokable from this surface. That 44 s is derived, not chosen: the CLI treats a token with 30 s or less left as unusable, up to ~10 s can elapse between the refresh being authorised and the token being signed, a further ~2 s is held back for the audit write and the commit that follow, and 2 s more covers the two separate points at which the remaining lifetime is floored to whole seconds. A session in that final window is still listed and still endable — the refusal is about minting a replacement, not about hiding the session.

pt CLI public-client flow (INF-286)​

The first-party CLI client uses the same consent screen with an RFC 8252 loopback redirect_uri (http://127.0.0.1:<port>/callback, any port — allowed only for constellation-pt-cli) plus PKCE. Consent-time roles are resolved server-side from the authoritative provider (AuthProvider.getUserInfo — Supabase app_metadata.roles in prod), never the session JWT. The exchange mints a PT bridge token (label pt-cli-oauth) and a rotating refresh token family (identity.mcp_cli_refresh_tokens, hashes only, ~30-day absolute lifetime). Refresh re-runs the DIR-87 fail-closed chain behind a family row lock; reuse outside a one-shot ~60 s grace window revokes the family. Role or membership changes through Directory's mutation tools invalidate the user's grants + refresh families (mcp_authority_invalidated, audit-critical), forcing re-consent under current roles.

Consented grants and the S2S re-mint (DIR-80 / DIR-87)​

The token exchange also returns a short-lived PT bridge token (ptAccessToken, DIR-80) and — since DIR-87 — persists the consented identity (tenant_id, user_id, client_id, org_id, roles, scope) in identity.mcp_grants (FORCE-RLS, unique per (tenant_id, user_id, client_id)), inside the same transaction as the exchange.

The refresh mode (grant_type: "refresh" on the same endpoint) re-mints the bridge token for a connector-held grant with no browser session. It re-verifies the user's ACTIVE membership and signs the fresh token exclusively with the stored consent-time roles — the refresh request carries no roles. Missing grant record, offboarded user, or a cross-tenant lookup all fail closed with invalid_grant. The mode is enabled by default; MCP_PT_REMINT_ENABLED=false (or 0) is the explicit kill-switch (403 refresh_disabled). Each re-mint emits a routine auditAction (mcp_pt_token_reminted) with the request correlation id.

The authorization-code exchange accepts an optional expected_scope request field — the scope the connector displayed and validated at /authorize. Directory compares it against the consumed code's scope at the top of the exchange transaction (before any consent audit, grant persist, or bridge-token mint) and fails closed with 400 { "error": "invalid_scope" } when it differs or is absent. Because the check throws inside the same transaction, a mismatch rolls the whole exchange back — the one-time code is not consumed, no mcp_oauth_consent audit is written, and no identity.api_keys bridge token is minted. This binds the issued grant to the scope the user actually saw, atomically, so a tampered intermediate authorize URL cannot leave behind an orphaned (unusable) bridge token. The connector maps invalid_scope to access_denied (unchanged client UX). The re-mint mode has no consent screen to echo, so it never carries expected_scope. Reject-on-absent means Directory must deploy in lockstep with (or after) the connector that forwards the field.

On a successful token exchange, auditCritical() is called with:

  • action = 'mcp_oauth_consent'
  • resourceType = 'mcp_client'
  • resourceId = client_id
  • module = 'directory'
  • changes = { scope, redirect_uri, client_id }

On a denied consent, auditAction() is called with action = 'mcp_oauth_consent_denied'.

Environment variables​

VariablePurpose
MCP_CLIENT_SECRETShared secret validated with timingSafeEqual on POST /api/internal/auth/mcp-token.
MCP_REDIRECT_ORIGIN_ALLOWLISTComma-separated exact scheme+host allowlist (e.g. https://mcp.planetb2b.com,https://constellation-pt-mcp.vercel.app). Editing it requires a Directory redeploy to recompile.
MCP_COOKIE_SECRETHMAC-SHA256 key with a dual use: it signs the mcp_authz_pending cookie and keys the code_hash HMAC for identity.mcp_authorization_codes. Rotating it invalidates outstanding authorization codes as well as in-flight cookies. mcp_csrf is an unsigned random double-submit nonce and is NOT signed with this secret.

Trust model​

The real trust anchor is the redirect_uri allowlist — the one-time authorization code is only ever 302'd to an allowlisted origin. A spoofed client_id cannot exfiltrate the code because it is always delivered to the already-verified origin. This is stated explicitly on the consent screen UI.

11. Default role seeding (PT-564)​

Directory owns the authoritative, audited capability for seeding the six platform default roles on every tenant. This replaces the earlier Project Tracker–side bootstrapTenant call, which was non-compliant with the constitution (identity writes must be Directory-owned, per-tenant, and auditCritical-emitting).

The six default roles​

RolePermission selector
AdminAll permissions in the catalog.
Project ManagerAll permissions whose resource is in the project-management set (fields, gates, initiatives, projects, stages, risks, tasks, etc.), plus coordinator.use.
Viewerprojects.read, tasks.read, initiatives.read, comments.create, coordinator.use.
Guest Customerprojects.read.own, tasks.read, initiatives.read, comments.create, files.read.
Project Customerprojects.read, tasks.read, initiatives.read, comments.create, files.read.
Project ApproverProject Customer set + tasks.approve.

Every role above additionally carries the three app-access permissions described in §12, and agents.problem_reports.file (Admin and Project Manager also agents.problem_reports.triage), described under Agent problem-report permissions. That is load-bearing: ensureDefaultRoles replaces a role's permission set, so a spec omitting them would revoke them on the next run. The same rule put risks in the Project Manager set (DIR-131): Project Tracker migration 008_risks_permissions.sql granted risks.<action> to every role that held stages.<action> when it ran, so most live Project Managers hold all four while tenants created afterwards hold none (production at DIR-131: 9 of 14 tenants held them). A spec without risks would have revoked them on the first reconcile; with it, the reconcile preserves those grants and adds the missing ones where a Project Manager lacked them.

The three customer-collaboration roles (Guest Customer, Project Customer, Project Approver) gained initiatives.read in PT-564 — the previous PT-seeded defaults omitted it, which was the root cause of the bug. (Viewer already carried it.)

coordinator.use (PLT-225) lets a member consult the coordinator and read an initiative's shared decision log without being an initiative delegate. Project Manager and Viewer carry it as an exact pair; Admin gets it through its all-permissions selector. The three customer-collaboration roles deliberately do not carry it, because the log holds other members' questions and answers. Migration 054_coordinator_use_permission.sql seeds it and grants it to no role, for the reason §12 gives for app.agents.read: a grant must co-commit auditCritical(), directory.role.permission_changed and MCP-authority invalidation, which migration SQL cannot reach. An existing tenant's roles therefore hold it only after ensureDefaultRoles next runs there (the backfill below). The coordinator honours it only for an ACTIVE, non-guest member, through a GLOBAL-scope assignment, on an UNCLASSIFIED initiative — see coordinator.initiative_access_admits in packages/platform/coordinator/migrations/021_coordinator_use_access.sql. Grant it to a custom role only alongside org-wide projects.read. A holder reads every UNCLASSIFIED initiative's coordinator sessions (including their page context), consult answers, consult-staged action previews and token usage, and those were produced under their authors' project scope. A holder needs no initiatives.read to open the coordinator: Project Tracker's /coordinator pages load GET /api/coordinator/initiatives, which lists only the initiatives that predicate admits.

These six exact names at GLOBAL scope with no scopeId are canonical system roles. Directory exposes them as isSystem: true and rejects generic creation, rename, permission-replacement, and delete operations. Custom roles, including same-named roles at a non-global scope, remain mutable. Canonical permission changes flow only through ensureDefaultRoles.

Tool: ensureDefaultRoles​

import { ensureDefaultRoles } from '@/server/tools/index';

const result = await ensureDefaultRoles(ctx, {
adminUserId: 'uuid-of-user', // optional — assign Admin role
actor: { id: 'system:backfill', type: 'SYSTEM' }, // optional — audit actor
});
// → { created: string[], updated: string[], adminAssigned: boolean }

Idempotent. Runs inside a single withTenantContext. Emits auditAction CREATE per new role, and auditCritical UPDATE_PERMISSIONS + role.permission_changed only for roles whose permission set actually changes (an unchanged re-run emits neither — no audit/event noise), plus optionally auditCritical ASSIGN_ROLE + role.assigned for the admin assignment.

Endpoint​

POST /api/v1/tenants/:id/seed-default-roles

  • Requires create:role + assign:permission + assign:role — the full set the operation performs, so a holder of only create:role can't self-assign Admin via adminUserId.
  • The caller may only seed their own tenant (params.id === user.tenant_id).
  • Optional body: { adminUserId?: string (uuid) }.
  • Response: { data: { created: string[], updated: string[], adminAssigned: boolean } }.

Called by Project Tracker on org creation (cross-module client call to Directory).

Backfill script​

Existing tenants that pre-date this change can be backfilled with:

# Dry run — writes its report to ./default-role-plan.json
DATABASE_URL=postgresql://... npm run db:backfill-default-roles -- --tenant <uuid>
DATABASE_URL=postgresql://... npm run db:backfill-default-roles -- --all-active

# Repair — only ever executes the plan that was reviewed
DATABASE_URL=postgresql://... npm run db:backfill-default-roles -- \
--tenant <uuid> --write --plan default-role-plan.json

Both commands are read-only by default and print a human plan plus a JSON report. --write is rejected without --plan <dry-run report path>: the repair replays the reviewed report and aborts if the set of ACTIVE tenants, any tenant's canonical drift, or the permission catalog moved since it was produced. Write mode calls ensureDefaultRoles per ACTIVE tenant as a SYSTEM actor, re-inspects after each write, exits non-zero on residual drift or tenant failure, and is idempotent. The manual Default role backfill GitHub workflow preserves the plan and result as artifacts and requires a protected Staging or Production environment before loading a write credential. The full staging/production procedure is in .ai/runbooks/default-role-production-repair.md.

Source locations​

Agent problem-report permissions (PLT-694)​

Migration 053_agent_problem_report_permissions.sql seeds two permissions for the Agents zone's capability-gap surface:

PermissionDefault rolesMeaning
agents.problem_reports.fileAll six (Admin through its all selector; the other five as an exact pair)File an agent problem report for the acting tenant.
agents.problem_reports.triageAdmin (through all) and Project Manager (as an exact pair) — no other roleRead and transition the tenant-wide problem-report queue. The Agents zone is to honour it only from a GLOBAL assignment of a GLOBAL role (PLT-694 S5).

Both are granted as exact resource/action pairs, never as the whole agents.problem_reports resource, so a later action under that resource is not inherited by every role that only needed to file. The triage scope rule belongs to the Agents zone's own identity lookup (PLT-694 S5), not to the permission row: Directory lets a GLOBAL role be assigned at PROJECT, PROGRAMME or MODULE scope, and the generic resolver flattens those into the same permission string. A project-scoped Project Manager therefore holds the string but is refused triage once that lookup ships.

The migration grants nothing to existing roles. A migration grant is unaudited authority, and during the deploy window the outgoing build's ensureDefaultRoles reconciles canonical roles back to its old specs anyway. Tenants provisioned by the new build get both permissions from DEFAULT_ROLE_SPECS. Existing roles get them only from the audited post-deploy repair, run once the new Directory deployment is serving and before any ensureDefaultRoles reconcile of an existing tenant (db:backfill-default-roles, db:provision-service-principals, re-provisioning a partner tenant). The reconcile replaces a canonical role's permission set and invalidates its assignees' MCP authority whenever that set changes, so reconciling first forces the pt login the repair avoids; after the repair it finds nothing to change:

DATABASE_URL=postgresql://… npm run db:backfill-agent-problem-report-permissions -w @constellation/directory \
-- --tenant <uuid> [--tenant <uuid> …]
  • Explicit tenants only (constitution §1): name every tenant in the estate with its own --tenant <uuid>. The repair never discovers tenants, refuses to run with none, and fails an id that is not a tenant. Before reading a tenant it commits an agent_problem_report_permission_repair.started audit entry; after every repair has committed it re-reads each started tenant under a .verified entry, and a tenant whose transaction rolled back gets .failed. A .verified entry records SUCCESS only for a tenant that passes the run (it exists, its repair committed, no role is still incomplete) and FAILURE otherwise, so the audit log agrees with the exit code.
  • Every role of each named tenant, custom roles included, gains file — except a role named for a registered SERVICE integration profile (SERVICE_PROFILE_ROLE_NAMES, INF-372), at any scope. Key issuance requires that role's grants to equal its profile, so granting it file would lock issuance for the multi-LLM reviewer and the spec-wiki sync. The canonical Admin and Project Manager (GLOBAL scope, no scopeId) also gain triage. A same-named role at a narrower scope does not.
  • Each changed role gets one auditCritical(UPDATE_PERMISSIONS) entry naming exactly the permissions added, plus one directory.role.permission_changed event, in the same transaction as the grant. If either side effect fails, the grant rolls back.
  • No MCP-authority invalidation. Unlike the DIR-98 app-access repair, this one does not call invalidateMcpAuthorityForRole. Backfilling file touches every role in every tenant, so invalidating would force a platform-wide pt login for a purely additive grant. It is not needed either: a consented MCP grant stores role names, not permissions, and permissions resolve from role assignments at request time, so the new permission reaches existing sessions without a new login.
  • A clean re-run records only its started/verified reads — no permission change and no event. Exit 0 means every named tenant verified with no role incomplete; 1 means a tenant failed, a role is still incomplete, or a tenant could not be verified — its re-read threw, or it recorded a FAILURE verification the rest of the result does not already show, as a tenant repaired and then deleted before the closing pass does; 2 means it did not run (no --tenant, a malformed or repeated one, an unrecognised argument, a missing DATABASE_URL, an invalid timeout, or migration 053 absent).
  • It is registered in the release runbook's cutover-repairs table, and each tenant transaction honours DIRECTORY_REPAIR_TX_TIMEOUT_MS (DIR-130).

Source: default-roles.ts (AGENT_PROBLEM_REPORT_PERMISSIONS, AGENT_PROBLEM_REPORT_TRIAGE_ROLE_NAMES), backfill-agent-problem-report-permissions.ts, and its real-Postgres suite agent-problem-report-permissions-repair.integration.test.ts.

12. App entitlement (DIR-98)​

Directory owns the data behind app-level entitlement — "may this (tenant, user) use this application at all?" — which is independent of, and enforced in front of, the row-level tenant RLS + classification every module already applies. Both dimensions are mandatory; neither substitutes for the other.

Two halves, ANDed by the shared resolver:

HalfSource of truthMeaning
Tenantidentity.tenant_apps — (tenant_id, app) unique, enabled booleanThe tenant has provisioned the app.
UserA per-app permission held through the user's rolesThe user may enter the app.

identity.tenant_apps is RLS-forced with a single tenant_apps_tenant_isolation FOR ALL policy on identity.row_in_current_tenant(tenant_id). app is CHECK-constrained to the known applications (five since migration 048 added agents — PLT-739), so adding another is deliberately a migration; the TypeScript mirror is apps/directory/src/lib/apps.ts.

Every tenant gets its rows automatically. An AFTER INSERT trigger on identity.tenants (trg_tenants_seed_apps) seeds every app in identity.default_app_keys() enabled, so no tenant-creation path — Directory's own, project-tracker's bootstrapTenant, the provisioning script, seed scripts — has to know about entitlement. The trigger is SECURITY INVOKER and is RLS-safe by construction: tenants_insert already requires app.tenant_id to equal the row's own id, which is exactly what tenant_apps' WITH CHECK needs. The FK is ON DELETE CASCADE, so deleting a tenant removes its entitlement rows.

Permission namespace​

AppPermissionGranted to
Wikiapp.wiki.readEvery role (031 blanket backfill)
Catalogapp.catalog.readEvery role (031 blanket backfill)
Projectsapp.projects.readEvery role (031 blanket backfill)
Directorydirectory.readAdmin-shaped roles only
Agentsapp.agents.readAdmin roles only, via audited reconciliation (048 seeds)

The app.* namespace is deliberately disjoint from every module's data permissions. In particular the Projects entitlement is app.projects.read, never projects.read — the latter already exists as Project Tracker's "View all projects" data permission, which Guest Customer must not hold. app.* also works as a domain wildcard meaning "all apps". Directory keeps the pre-existing directory.read (its original app-access permission, granted to admin roles only) rather than gaining a duplicate under the new namespace.

Not every app permission is blanket-granted. Migration 031's every-role backfill covered exactly the three apps that were open to everyone before entitlement existed — that breadth was behaviour preservation, not the norm. app.agents.read (seeded by migration 048, PLT-739) is granted only to admin-shaped operator roles: the agents module is a control-plane surface with no pre-existing access to preserve. And the grant is not made in the migration at all — a role-permission grant is an audit-critical authority change that must co-commit auditCritical(), the directory.role.permission_changed event, and MCP-authority invalidation, none of which migration SQL can reach — so it flows through the audited app-layer paths: the canonical Admin acquires it at the next role reconciliation (its { kind: 'all' } selector; run the default-role backfill workflow post-deploy), and any further operator-role grant is an interactive role-management change or a per-role audited backfill in the backfill-app-access-grants.ts shape. app.agents.read is deliberately excluded from identity.app_access_resources(), SEEDED_APP_ACCESS_PERMISSIONS, and APP_ACCESS_CORE in default-roles.ts — each of those means "blanket-granted to every role", and the grant sweeps read them as such.

Fail-closed contract​

The consuming resolver treats an unresolved entitlement as denied, and omits the app silently rather than returning an error that would reveal whether it is provisioned. That is why the introducing migration is default-enable: it records the access every principal already had — every app enabled for every existing tenant, and the three founding app.* permissions granted to every existing role — so enabling enforcement changes nobody's access. (For agents, tenant-level enablement is likewise default-enable, but the fail-closed resolver will deny non-admin users by design — there is no pre-existing access to preserve.) A role created outside ensureDefaultRoles (e.g. hand-created in role management) starts with no permissions, so a user whose only role is such a role will not resolve entitled; grant the app permissions explicitly in that case.

Source locations​

13. API-key lifecycle: user-issued vs ephemeral bridge tokens (DIR-1 / DIR-103)​

identity.api_keys holds two populations that look alike and must be treated differently. Both are api_key-kind JWTs signed with API_KEY_JWT_SECRET, and both require a row here — PT's createPrincipalCheck resolves the token's jti against this table and rejects a token with no row as KEY_NOT_FOUND.

User-issued keyEphemeral bridge token
Created byPOST /api/auth/api-keys (a human, in Settings)mintPtAccessToken — the OAuth flows
ephemeralfalsetrue
Labeluser-chosen, free textpt-cli-oauth (CLI) / mcp-pt-bridge (connector)
TTL1–365 days, capped by role — see belowMCP_PT_TOKEN_TTL, default 1 h, hard cap 24 h
Listed by GET /api/auth/api-keysyes — active, expired and revokedno
Deleted automaticallynevereligible 7 days past expiry, swept by a later mint

TTL caps by role (DIR-105)​

A user-issued key's lifetime depends on whether the issuing principal is administrative:

PrincipalDefault TTLMaximum TTLEnv override
Administrative14 days30 daysAPI_KEY_TTL_MAX_ADMIN_DAYS
Everyone else90 days365 daysAPI_KEY_TTL_MAX_USER_DAYS

Administrative keys are short-lived precisely because they carry the issuing user's full authority. PLT-78 is marked complete, but its promised token-role cap is absent from the current permission paths; PLT-1125 owns both downstream token-bound enforcement and issuance-time capability selection. Requesting more than the applicable maximum is a ValidationError, not a silent truncation.

Both maxima are deployment-tunable and each is clamped to the 365-day platform ceiling (API_KEY_MAX_TTL_DAYS), so 30 is the administrative default maximum, not an absolute limit. A default is clamped down whenever an operator sets the matching maximum below it, so lowering a cap shortens keys rather than rejecting every request that omits expiresInDays.

"Administrative" is decided by one predicate — hasAdministrativeRole in @constellation-platform/auth-core. It matches Directory's canonical full-access role Admin plus the five historic variants admin, tenant_admin, TENANT_ADMIN, platform_admin, PLATFORM_ADMIN. Matching is case-sensitive: Admin and admin are distinct role identities here, and ADMIN is neither.

DIR-105. Until 2026-08-01 this gate used a private three-name list that omitted canonical Admin — the role granted the entire permission catalog — so a full-access principal received the 365-day ceiling. The predicate now lives in auth-core and is derived from ADMIN_EQUIVALENT_ROLES so the two cannot drift apart again.

That auth-core set is deliberately narrower: it grants blanket access in the simplified permission path, and a JWT roles claim carries names only, so a tenant-created role merely named Admin must not buy access. The TTL predicate may carry the name precisely because it restricts.

A custom role holding full permissions under some other name still receives the ordinary user caps — name matching is a proxy, tracked for replacement in DIR-106.

Why the split (DIR-103). A bridge token is minted per npx pt login and per access-token rotation, so an active CLI session produces roughly one row per hour. Before DIR-103 those rows were listed as if they were credentials the user managed, and never removed — the settings page filled with pt-cli-oauth entries badged "Expired" and the table grew without bound. They are OAuth session artifacts, so they are now excluded from the list and pruned.

Revoking a bridge token from a key list would also have been misleading: it does not end the CLI session. The refresh family in identity.mcp_cli_refresh_tokens is untouched, and the revocation's 401 drives PTClient into CredentialManager.onAuthError, which forces a refresh regardless of the stored token's exp and retries — so the row is replaced on the user's very next command. Ending a session means revoking the family (identity.revoke_mcp_cli_refresh_family, what pt logout does).

The sweep is audited. Deleting rows from an authentication table is a security-sensitive mutation, so the prune primitive returns the identities it removed and mintPtAccessToken writes an auditCritical entry (action: DELETE, resourceType: api_key, actorType: SYSTEM) naming them — id, jti, user, label and expiry — inside the same savepoint as the delete, so the two commit or roll back together. A sweep that deletes nothing writes no entry, so an idle mint adds no audit noise.

Pruning is opportunistic, not scheduled. Each mint sweeps up to 50 of its own tenant's ephemeral rows that are more than 7 days past expires_at — the same self-cleaning shape identity.rotate_mcp_cli_refresh_token uses for refresh-token tombstones, so no cron is involved and a tenant that stops minting stops accumulating. A row therefore becomes eligible at 7 days and is removed by whichever mint comes next; a tenant that never mints again keeps its remaining rows indefinitely, which is harmless because nothing lists or authenticates them.

What protects a live token is the expires_at bound, not the retention window: an unexpired row is never a candidate at any retention. The seven days are for forensics — they keep last_used_at readable for a while after a token dies, which the audit trail does not carry (it records the mint, not subsequent use). The sweep never matches a user-issued key, and it runs inside a SAVEPOINT, degrading to a logged no-op on failure — cleanup can never be the reason a pt login or a token refresh fails.

The floor and the cap are enforced inside the function, not supplied by the caller: holding EXECUTE cannot waive the forensic window by asking for zero retention, nor widen the bounded sweep into a tenant-scale delete.

The delete is least-privilege. It lives in the identity.prune_ephemeral_api_keys SECURITY DEFINER primitive, and the runtime role's table-level DELETE on identity.api_keys is revoked — the same design migration 030 uses for mcp_cli_refresh_tokens. A table grant would have been tenant-scoped by api_keys_delete but not scoped to ephemeral, to expiry, or to any row bound, so it would have let any bug in Directory hard-delete a live user-issued key. The four invariants (ephemeral-only, tenant-bound, past expiry + retention, bounded row count) are enforced in the database rather than only in app SQL, and the two thresholds are clamped inside the function so holding EXECUTE cannot choose the policy.

UPDATE is column-scoped to last_used_at, revoked_at, revoked_by, revoke_reason — the only columns the runtime writes after insert. That is what makes the DELETE revoke meaningful: with a table-wide UPDATE grant the protection was bypassable in two steps, by setting a user-issued key to ephemeral = true with a backdated expires_at and then invoking the primitive, at which point the row satisfies every invariant.

Deploy-window repair — the only path that classifies existing rows. The migration defines the classifier but deliberately does not invoke it: a blanket sweep would reclassify rows across every tenant with neither an explicit tenant_id nor an auditCritical() emission, and that helper is not reachable from SQL (it writes the audit row and publishes to the events outbox). So classification happens here, and running it is not optional — until it does, pre-existing bridge rows keep ephemeral = false and stay listed and unpruned.

After the deploy is also the correct moment for it. Migrations are applied to staging and production before the release PR opens, so for a period the new column exists while the old mintPtAccessToken — which does not set ephemeral — is still serving; a sweep at migrate time could not see the rows that code is about to mint. Once after the deploy is live, run:

DIRECT_URL=postgresql://… npm run db:backfill-ephemeral-api-keys -w @constellation/directory

It sweeps per tenant, inside each tenant's own transaction, emitting an auditCritical record for every tenant it actually changes — constitution §1 requires a cross-tenant operation to carry an explicit tenant_id and an audit entry, and ephemeral is a classification with security meaning. It calls the migration's own classifier (identity.backfill_ephemeral_api_keys), so the repair and the migration cannot drift; that function is withheld from the runtime role. It then verifies no candidates remain in any tenant and exits non-zero if any do, or if any tenant errored. It requires the privileged DIRECT_URL role and preflights for it, refusing a role that can neither bypass RLS nor act as superuser. Two independent reasons, neither of them a silent no-op: the runtime role does not hold EXECUTE on the classifier at all (the same withholding described above), so it would raise permission denied for function; and the final verification is deliberately tenant-agnostic, so under RLS it would see nothing and could report clean while candidates remained. What confines a privileged run to one tenant at a time is the explicit tenant_id predicate on each write, not RLS. The step is registered in the release runbook's post-deploy repair table, which is what a release runner actually follows.

Per-jti and per-id lookups are deliberately not filtered by ephemeral: hiding a row from a list must never change whether its token authenticates or can be revoked.

Source locations​

14. Service principals & on-behalf-of key issuance (INF-372)​

identity.users carries a principal model distinguishing people from machines:

  • principal_kind — HUMAN (default) | SERVICE | AGENT. Settable only by migration, seed, or operator script: the user create/update schemas explicitly reject the field (an attempt is a validation error, never a silent strip), so no request path can change what a row is.
  • integration_key — the immutable per-integration selector, NOT NULL exactly when principal_kind = 'SERVICE' (CHECK-bound), unique per tenant (partial unique index). Grant profiles key on it — never on email, which stays user.update-mutable and is presentation-only. A reserved-namespace guard in UserService additionally refuses a non-SERVICE row taking a ci+…@constellation.local address, and a SERVICE principal's email is immutable through the user API: renaming one onto a human-controlled address would let any email-fallback identity resolver (they select by address, not by principal_kind) hand that human the integration's identity.
  • The synthetic agent+…@constellation.local rows (INF-113/125) are AGENT: seedAgentUsersForTenant stamps new ones explicitly, and the audited post-deploy repair (db:backfill-agent-principal-kinds) classifies the pre-existing ones across every tenant. Migration 047 deliberately reclassifies nothing — a principal's kind is security-meaningful, so the change carries a tenant and an audit entry, neither of which a migration can produce (the migration-033 precedent).
  • The agent namespace is reserved too (DIR-153). Project Tracker tells an agent from a human by the agent+…@constellation.local address alone — the human-only pending-action gates, the agent-only WIP and dependency gates, claimedBy, audit actorType — so UserService refuses creating a row in that namespace (every API-created row is HUMAN) or renaming any row into it, matched case-insensitively, as the ci+… guard is. An agent row's email is immutable through the user API in every direction: off the namespace, onto another agent address, or to another casing. The freeze keys on the address as well as the kind, because a long-lived database the audited repair has not reached can still hold HUMAN-kinded agent rows, and so can rows written by raw SQL that sets no kind. Both refusals are a 400 VALIDATION_ERROR; the row's other fields and its status stay editable. The partner-provisioning script refuses either reserved namespace as its --admin-email. The database CHECK behind both guards is DIR-123.

On-behalf-of issuance. POST /api/auth/api-keys/service — a separate, strict endpoint — mints a key for a SERVICE principal: the token carries the target's sub and its role names resolved from identity.user_roles (never the caller's, and never the dead users.roles column), at a fixed 90-day TTL (expiresInDays is rejected here, as is any unrecognised field — the schema is .strict() and targetUserId is required, so a mistyped selector is a 400 rather than a key quietly issued for the caller). The caller must be a browser session (the no-chaining rule, asserted at the route and inside the tool) and hold the api_key.issue_for_service permission through the user_roles → role_permissions → permissions chain — deliberately not checkPermission, whose PLATFORM_ADMIN short-circuit never consults the chain.

Target gate: exact grant-profile equality. The target must be SERVICE, ACTIVE, not soft-deleted, same-tenant, hold an ACTIVE org-bound membership, and its persisted grants must exactly equal the profile registered for its integration_key in the code registry (SERVICE_INTEGRATION_PROFILES): exactly one GLOBAL, tenant-bound user_roles assignment; a GLOBAL, tenant-bound role with the profile's exact name; a (resource, action) multiset equal to the profile's; every linked permissions.condition NULL; and hasAdministrativeRole rejected on top. Authority-affecting drift therefore locks issuance instead of shipping silently; identity-preserving recreation (same shape under new UUIDs) passes by design. Equality — not "is this role administrative" — because DIR-106 showed classification is not decidable here (canonical Admin holds 75/77 catalog permissions and no *.* row).

Profile role names stay unique tenant-wide, in three legs. Token consumers resolve permissions by role name with no scope filter and union every match, so a second same-named role at any scope would widen every outstanding service key. (1) The role API refuses creating or renaming onto a registered profile name while any role in the tenant carries it (409, any scope). Deliberately not an outright reservation: wiki:page:write is also the wiki interim's public authorization vocabulary — an ordinary tenant provisions its least-privilege wiki-writer role by creating exactly that name, and must stay able to create the first one. (2) Provisioning refuses to adopt a pre-existing same-named role that anyone other than the service principal holds — a role squatted ahead of provisioning is never legitimized with service authority; the operator deletes it (or renames it away from the reserved name, which stays allowed) and re-runs. (3) Issuance refuses while a same-name collision exists in the tenant. (4) A provisioned profile role's name cannot be released — renamed away or deleted — while the integration has live service keys: outstanding tokens carry the role name and consumers resolve it live, so freeing the name would hand their authority to whatever role claims it next (revoke the integration's keys first; with zero live keys the release stays allowed, which is the squat-remediation path). The same live-key gate refuses every authority-affecting mutation on a SERVICE principal while keys are live — assigning any additional role, creating a membership, reactivating or re-binding one (a membership grants act-as-org authority the issuance-time count never saw); suspensions and revocations narrow and stay allowed, and user-level reactivation is gated identically. The gate reads only the acting tenant's own row, RLS-scoped and tenant-predicated, and refuses uniformly when it cannot resolve the principal — foreign-home and absent are indistinguishable in the refusal. Nothing classifies another tenant's user on the shared runtime role, so there is no cross-tenant classification oracle; the one thing a foreign tenant could otherwise have widened is a pre-existing cross-tenant membership row, which §1 forbids regardless of the principal's kind. The gate holds a principal-scoped advisory lock that issuance and provisioning share, so a mutation cannot race a mint. Permission edits on the profile role itself remain the ruled 2026-08-01 trade-off (operator revoke-and-re-issue procedure; auto-revoke tracked as DIR-122). Claims, releases, issuance, and assignments of profile-named roles all serialize on one tenant+name advisory lock (tenant id case-canonicalized in the key). Together: a squatted or colliding role can never feed a service key, while tenant wiki vocabulary keeps working.

Admin list/revoke. GET /api/auth/api-keys?serviceUserId=<uuid> lists a SERVICE principal's keys (metadata only, tenant-scoped); DELETE /api/auth/api-keys/[id]?serviceOwned=true revokes a SERVICE-owned key. Both require the same permission; a HUMAN-owned key is refused — human keys stay strictly owner-revocable. Issuance and privileged revocation write auditCritical entries with changes.onBehalfOf = { subjectUserId, integration } — stable identifiers, no email in the payload.

Provisioning is an operator step, not a migration. Migration 047_service_principals.sql ships only the tenant-independent parts (columns, CHECKs, the partial unique index, and the global ('api_key','issue_for_service') permission row) — and reclassifies no existing user row. The tenant-bound principal/role/grants/membership per integration are provisioned by npm run db:provision-service-principals -w @constellation/directory -- --tenant <uuid> --org <uuid> — idempotent, one audited transaction per profile, post-validated through the real issuance-path check, and it invokes ensureDefaultRoles so canonical Admin (selector all) picks up the new permission. Operational runbook: CI credentials § Service principals.

Source locations​

15. Purging a disposable tenant (DIR-136)​

Directory owns the only way to remove a tenant's identity.* rows. identity.purge_disposable_tenant(uuid) (migration 050_disposable_tenant_purge.sql) deletes every one of them, in foreign-key order, inside the caller's transaction, and returns per-table deleted row counts as jsonb so the caller can record exactly what vanished in its own audit entry.

It exists so a module that owns a disposable test tenant can tear it down without writing the shared identity schema directly — which constitution §2 reserves to Directory, and which ADR-025 says is done through a SECURITY DEFINER function on this boundary.

It only ever touches a tenant created as disposable. identity.tenants carries disposable boolean NOT NULL DEFAULT false, and a BEFORE UPDATE trigger refuses any change to it — so the value can be established by the INSERT that creates the tenant and never afterwards. Every tenant that already exists, in every environment, is therefore un-purgeable, and a mistyped id naming a real tenant is refused by the database rather than by the calling script. A tenant not created disposable cannot be converted; create a new one.

Nobody can call it at runtime, and that is deliberate. Migration 050 grants EXECUTE to the function's owner (identity_writer) and to the applying role, then revokes it from every other non-superuser role — including the ALTER DEFAULT PRIVILEGES grant db-init.sql hands the shared runtime login at creation time. An earlier revision derived the deployment's runtime caller and granted it, so the boundary would be reachable on a deployment whose login is not called constellation_app. Review retired that by demonstration: the same role holds INSERT on identity.tenants, so with app.tenant_id set first it could create a disposable = true tenant and immediately purge it — a committed identity purge with no audit entry anywhere. PLT-1102 issues the runtime grant together with the auditCritical() wrapper that discharges the constitution's unwaived obligation, which is the only place that audit can live: a migration in this repository cannot write an audit row.

Three refusals the function itself raises, and each is deliberate — plus one failure it does not raise, kept in the list below because a caller meets it the same way and would otherwise mistake it for a bug:

  • the tenant is not disposable — identity.tenants.disposable must be true, and it can only have been set at INSERT. This is the refusal a mistyped production tenant id meets. It fires before anything is DELETED — the function reads the tenant row to decide it, and the catalog completeness check runs earlier still, so "before any read" would be wrong;
  • the acting tenant must match — app.tenant_id must be set and equal the argument, in its canonical lower-case text form, because the row-level-security policies the function reads through compare tenant_id::text case-sensitively while uuid comparison is not;
  • cross-tenant dependants — if a user or an organisation of the tenant owns children scoped to a different tenant, the purge refuses rather than destroying another tenant's data through a foreign key. Eight tables can hold such a child: the five cascading children of identity.users (cross-tenant membership is supported, so this is ordinary), the two of identity.organisations, whose tenant_id is never checked against the organisation's, and identity.user_tenant_memberships. That eighth is worth naming separately, because it is the odd one and was missing from this page until review caught it: its foreign key is NO ACTION rather than cascading, so without the check it would not destroy another tenant's row — it would abort the purge with a raw 23503 instead of this function's named refusal. The check runs once, immediately before the two parent deletes, with FOR UPDATE locks on the tenant, its organisations and its users taken before any delete — the locks are what stop a child being inserted mid-flight, and the late position is what catches one MOVED to another tenant after the purge began, which no lock available to this boundary can prevent. It reads across the tenant boundary through six SELECT-only policies added for it — without them it would count zero on six of the eight and pass on precisely the tenants it exists to refuse;
  • NOT a refusal, but reached the same way — rows left in other schemas. wiki.*, projects.*, coordinator.* and agents.* reference identity.tenants with NO ACTION, so a caller that has not removed its own rows first gets a raw foreign-key violation from PostgreSQL, not a message from this function. It is listed here because the caller's experience is the same — the purge aborts and nothing is deleted — but it is a constraint doing its job rather than a guard this boundary implements, and Directory deliberately does not read another module's data to pre-empt it. Review flagged an earlier version of this list for blurring that line.

Three foreign keys are the exception to that last point, and they are not fixed here. projects.feedback_counters.tenant_id cascades from identity.tenants, and projects.issue_queues.created_by_id and projects.issue_resolution_outcomes.created_by_id null out from identity.users — so those rows change without Project Tracker's cleanup running and without appearing in the returned counts. Changing a projects foreign key is Project Tracker's to make (PT-1024); what Directory asserts, at apply time and in CI, is that no fourth such foreign key appears.

Completeness, and its one edge. The purge is complete against every writer the platform has: children of the tenant's users and organisations cannot be inserted while it runs (their parents are locked), cannot be moved out unseen (the check is re-asked once, immediately before the two adjacent parent deletes), and nothing with a foreign key to identity.tenants can outlive the tenant. Two tables carry no such foreign key — mcp_authorization_codes and mcp_cli_refresh_tokens — so a row inserted after the final scan, carrying this tenant's id while pointing at another tenant's user, would survive. The caller cannot close this window on its own. An earlier version of this page said to run the purge in a SERIALIZABLE transaction; that does not help, because PostgreSQL's serializability guarantee holds only among transactions that are themselves SERIALIZABLE, and the writers of those two tables run at the default READ COMMITTED. What removes the anomaly is DIR-137, which makes a cascading child's tenant_id follow its parent's.

Audit is the caller's. The function writes no audit.* row: auditCritical() is TypeScript that writes the immutable entry and publishes audit.entry.created to the outbox, and no migration in this repository has ever written an audit row. Because the function joins the caller's transaction, a caller invokes it and calls auditCritical() with the returned counts, and both commit or roll back as one.

Ownership. It is owned by identity_writer, like every other function on this boundary, which gains SELECT and DELETE on the fifteen tables it purges. Membership in that role must therefore be read as unrestricted cross-tenant identity write and delete authority — see ADR-025 § Amendment (DIR-136), which also records why a dedicated confined owner is not available: migration 039a's owner allowlist is frozen and replaying it is a supported recovery path.

Source locations​

See also​