Deployment
Constellation is deployed on Vercel as a multi-zone topology:
| Zone | Host | Vercel project |
|---|---|---|
| Directory (root) | constellation.planetb2b.com | constellation-directory |
| Project Tracker | /projects/* (rewritten) | constellation-platform |
| Catalog | /catalog/* (rewritten) | constellation-catalog |
| Docs | docs.planetb2b.com | constellation-docs |
Directory is the root zone and rewrites requests for /projects/* and /catalog/* to the other Vercel projects.
NEXT_PUBLIC_BASE_PATH
Sub-zone apps read this env var at build time:
| Environment | Value | Behaviour |
|---|---|---|
| Production | /projects or /catalog | Multi-zone: served under prefix, rewrites from Directory |
| Preview | unset | Standalone: served at / for preview deploys |
| Local dev | unset | Standalone: turbo dev on app-local port |
fetch() and basePath
Next.js auto-prepends basePath to <Link> and router.push, but not to fetch(). Use the apiUrl() helper from each app's src/lib/api-url.ts:
import { apiUrl } from '@/lib/api-url';
const res = await fetch(apiUrl('/api/people')); // '/projects/api/people' in prod
Migration gate on release PRs
Every PR targeting main (release branches and hotfixes) runs the Verify Release Migrations CI job. It fails the merge if either staging (DEV_DATABASE_URL) or production (PROD_DATABASE_URL) is missing a migration that's on the branch. The job is read-only — it never applies migrations. Apply outstanding migrations to both databases via the Supabase dashboard SQL editor or the Supabase MCP (apply_migration / execute_sql) before opening the release PR.
CONCURRENTLY cannot go through a transactionScan the migrations you are about to apply first — recursively, since an entry can be a directory:
grep -rl CONCURRENTLY apps/<module>/prisma/migrations/<entry>
PostgreSQL rejects CREATE INDEX CONCURRENTLY / DROP INDEX CONCURRENTLY outright inside a transaction block, so a tool that wraps what it is given does not merely apply them more slowly — the migration fails. Those files are written deliberately without BEGIN;/COMMIT; to avoid holding an ACCESS EXCLUSIVE lock on a busy table for the length of a full index build, and they say so in their header.
Apply them with a session that does not wrap, and always with ON_ERROR_STOP=1 — psql's default is to report an error and carry on, which lets a migration fail its DDL and still run its final schema_migrations insert, recording an incomplete migration as applied:
psql "$PROD_DATABASE_URL" -v ON_ERROR_STOP=1 -f <file>
A migration entry can hold more than one file — 033_tasks_backlog_grooming is migration.sql plus post_migration.sql, and the grep points only at the one containing CONCURRENTLY. Apply every .sql in the entry in runner order (migration.sql first) and record the tracking row only once they have all succeeded; stamping the entry after applying one file marks the whole migration as done while the rest never ran.
The raw psql path also bypasses the module runner's tracking-row write, and not every migration self-inserts. Check the file for a schema_migrations insert and add one yourself if it has none, or the release gate stays red against a database that is already up to date — chained to the apply with && so it cannot stamp a migration that failed, into the owning module's schema (see scripts/check-migrations.ts), and keyed on the migration entry name (a directory-based migration is tracked by its directory, not by the inner .sql file).
If you use the dashboard or MCP instead, run the concurrent statement on its own and then confirm the index is valid: a failed concurrent build leaves an INVALID index behind, which a later re-run guarded by IF NOT EXISTS will skip straight past.
SELECT indexrelid::regclass, indisvalid FROM pg_index WHERE NOT indisvalid;
Full procedure: .claude/skills/release-and-migrations/SKILL.md § Migrations Before the PR.
See scripts/check-migrations.ts (invoked with --strict from the gate) for the implementation.
Platform packages carry their own migrations
Migrations do not live only under apps/<module>/prisma/migrations. A shared package can own a schema too, and @constellation-platform/jobs owns jobs — its files are in packages/platform/jobs/migrations, applied by the package's own db:migrate, and tracked in jobs.schema_migrations like any module's. npx turbo run db:migrate picks it up with everything else; turbo.json orders @constellation/project-tracker#db:migrate after it, because both would otherwise create jobs.queue and race.
Baseline adoption on databases that already have jobs.queue
Project Tracker's grandfathered 013_jobs_queue.sql created jobs.queue long before the package had a runner, so on staging, on production and on any machine that has run PT's db:setup, the table exists while jobs.schema_migrations does not. Re-applying 001_jobs.sql there is not what you want, and neither is skipping it silently.
The runner therefore baseline-adopts 001: it records the file as applied without executing it, but only after comparing the deployed table against a reference the same server builds from 001's own declaration — columns, constraints, indexes, RLS flags and every policy's command, kind, role scope and predicate. The reference is a closed set, so an object neither 001 nor 002 declares is refused rather than tolerated: an extra permissive policy is a cross-tenant read path that applying 001 would not remove, and an extra UNIQUE index rejects enqueues the schema permits.
Three outcomes, and each is loud:
- matches —
001is recorded as applied and never executed, then002runs normally; - differs repairably (a missing index, RLS not forced, a policy
001recreates) — adoption is declined and001is applied, which is written entirely inIF NOT EXISTS/DROP POLICY IF EXISTSform; - differs unrepairably — the runner aborts and names the difference. Nothing is recorded. Reconcile the database by hand before re-running; a recorded migration is a claim every later migration acts on.
Every cross-tenant drain audits its claim and its settle
Five drains claim jobs.queue rows across tenants: the wiki's /api/cron/process-jobs and Project Tracker's notification, agent-run, run-score and cycle-retrospective drains. Since PLT-1342 all five run one shared transaction machine, runInWorkerBypassTransaction in @constellation-platform/jobs, on the confined constellation_jobs_worker login. What an operator can rely on:
- Every claim and every mutating settle writes an
auditCritical()entry in its own transaction —wiki.job.claimed/wiki.job.settled, orprojects.job.claimed/projects.job.settled— attributed to the job's tenant, with actorSYSTEMand resource typejobs.queue. The entry records lifecycle fields only (job type, status, attempts), never the payload or the last error. - An audit or outbox failure rolls the queue change back. A refused claim leaves the job
pendingfor the next tick; a refused settle leaves itrunningfor the stale-claim lease. Soaudit.audit_entriesandevents.outboxprivileges are a hard prerequisite of the drain, which is why its preflight refuses to start without them. - A settle is refused, and writes nothing, when the job has moved on since it was claimed or observed — reclaimed at a higher
attempts, or no longerrunning. The drain logs it and does not count the job as drained. - A suspended tenant's jobs are still claimed and audited. Attribution deliberately bypasses Project Tracker's tenant-liveness gate (ADR-044 § D7); gating it would stall the job type for every tenant. Org-less notifications are audited under the all-zeros tenant they are stored under, which no tenant session reads.
Runtime-role grants travel with the migration
002_worker_claim_policy_and_runtime_grants.sql grants constellation_app schema USAGE plus SELECT, INSERT, UPDATE on jobs.queue — not DELETE — and read-but-not-write on jobs.schema_migrations. This is deliberate and it is why enqueuing works on a deployed database at all: RLS grants nothing on its own, and locally the privileges came from scripts/db-init.sql, which is a dev-only fixture. Without them the symptom is 42501 out of PostgresJobQueue.add() and out of every worker claim.
The grants are guarded on the role existing, and followed by hard postconditions — including a negative one, because REVOKE reports success whether or not a grant existed and so says nothing about write access reaching the role through PUBLIC or through a role it is a member of. If the migration role lacks authority to grant, the file raises instead of being recorded: the shared runner suppresses any statement that fails with 42501 and logs it as a skip, which would otherwise leave the ledger claiming work that never happened.
audit migration 013: apply it in one transaction
Applies to Dedicated Cloud, On-Prem, and any database that has not yet applied packages/platform/db/migrations/013_audit_chain_heads_rls.sql — and to any database that records it but was never verified. SaaS staging and production applied it in one transaction on 2026-09-13, starting from no policies on the table, and its check passed on both; they need nothing. (PLT-1225)
013 turns on tenant row-level security for audit.chain_heads. It creates each of its three policies only when that policy name is missing. Then it enables and forces RLS. Only after that does it check that the table carries exactly those three policies with its definitions. The file opens no transaction of its own. So when it is applied one statement at a time (psql -f without -1), the ENABLE and FORCE have already committed when that check runs. If the table already had a same-named policy defined differently, or an extra permissive policy, the check fails but RLS stays on with that policy in force. For example, a leftover USING (true) select policy lets every role read every tenant's chain heads. A same-named policy that is stricter than 013's fails the same check and takes the same repair. Without ON_ERROR_STOP psql keeps going after the failure and records 013 as applied, so the migration ledger shows nothing wrong.
Before applying, check the table. On a database that has not applied 013 this returns f | f | {}. If policies is not empty, go to the repair below instead of applying:
SELECT c.relrowsecurity, c.relforcerowsecurity,
array(SELECT polname FROM pg_policy WHERE polrelid = c.oid ORDER BY polname) AS policies
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'audit' AND c.relname = 'chain_heads';
Apply 013 on its own, in one transaction, after 001–012:
psql "$DIRECT_URL" -X -v ON_ERROR_STOP=1 -1 \
-c "SET LOCAL lock_timeout = '5s'" \
-f packages/platform/db/migrations/013_audit_chain_heads_rls.sql
-1 puts the -c and the whole file in one transaction, so if the check fails the ENABLE is rolled back too. -X skips ~/.psqlrc, where settings such as AUTOCOMMIT or ON_ERROR_ROLLBACK could change what commits. Use -1 for this file only, not the whole directory: 004_audit_partitioning.sql opens and commits its own transaction, and -1 does not work as intended on a file that does.
Run from the repository root, npm run db:migrate also applies 013 in one transaction. It reaches packages/platform/db/migrations through Project Tracker's migration runner, which applies 013 with --single-transaction. It sets no lock timeout, though. On a live database its ACCESS EXCLUSIVE request can wait indefinitely behind an open transaction, with every audit write queued behind it, so use the command above there. A checkout from before PLT-1225 applies 013 one statement at a time; use the command above on those too.
Creating a policy and enabling RLS each take an ACCESS EXCLUSIVE lock on audit.chain_heads, held until the transaction ends. That blocks reads as well as writes, so audit writes that advance the hash chain wait for it. The lock request also queues behind any transaction already using the table. Apply in a quiet window. lock_timeout limits each lock wait on its own, not the apply as a whole. If any one lock cannot be acquired within 5 seconds, the apply gives up (exit 3, table unchanged). It does not bound how long audit writes are blocked. They wait while the apply waits for its locks, and those waits can add up past 5 seconds. Once it holds the table lock, they wait until its transaction commits or rolls back. It is set with SET LOCAL inside the transaction rather than through PGOPTIONS, which a connection string carrying its own options= silently overrides.
If 013 fails with tenant RLS was enabled on audit.chain_heads but recording the tracking row in audit.schema_migrations was refused, do not follow that message's advice to insert the row by hand. Under -1 the message is wrong: the rollback has already undone the ENABLE, so a hand-inserted row would record 013 with RLS off. Grant the applying role both privileges the ledger insert needs, INSERT and SELECT (filename) (its ON CONFLICT (filename) reads that column): GRANT INSERT, SELECT (filename) ON audit.schema_migrations TO <applying role>. Then apply again.
Verify, both after applying and on any database that already records 013 but was never checked. The first query must return t | t with the three names below, and this one must return exactly these rows:
SELECT policyname, permissive, roles, cmd, qual, with_check
FROM pg_policies
WHERE schemaname = 'audit' AND tablename = 'chain_heads'
ORDER BY policyname;
| policyname | permissive | roles | cmd | qual | with_check |
|---|---|---|---|---|---|
chain_heads_tenant_isolation | PERMISSIVE | {public} | SELECT | ((tenant_id)::text = current_setting('app.tenant_id'::text, true)) | empty |
chain_heads_tenant_isolation_insert | PERMISSIVE | {public} | INSERT | empty | ((tenant_id)::text = current_setting('app.tenant_id'::text, true)) |
chain_heads_tenant_isolation_update | PERMISSIVE | {public} | UPDATE | ((tenant_id)::text = current_setting('app.tenant_id'::text, true)) | ((tenant_id)::text = current_setting('app.tenant_id'::text, true)) |
Repair, when the preflight shows a policy, 013's policy check refuses it, or the verification fails. First inspect the policies that are there. Then replace them all in one transaction. The command drops every policy on the table and re-runs 013, which recreates its three and checks them again. It takes the same ACCESS EXCLUSIVE lock for the length of that one transaction, so run it in a quiet window too:
psql "$DIRECT_URL" -X -v ON_ERROR_STOP=1 -1 \
-c "SET LOCAL lock_timeout = '5s'" \
-c "DO \$\$ DECLARE p record; BEGIN FOR p IN SELECT polname FROM pg_policy WHERE polrelid = 'audit.chain_heads'::regclass LOOP EXECUTE format('DROP POLICY %I ON audit.chain_heads', p.polname); END LOOP; END \$\$" \
-f packages/platform/db/migrations/013_audit_chain_heads_rls.sql
Never insert the audit.schema_migrations row by hand to get past the check.
Why 013 itself is not fixed. It has shipped and been applied. The ledger matches migrations by filename alone, so an edited 013 would never run on a database that already recorded the original. A new migration cannot close the gap either: it runs after 013, when RLS has already been switched on.
Upgrade prerequisite: the wiki INTERNAL classification tier
Applies to Dedicated Cloud and On-Prem. On SaaS the platform controls the rollout and there is nothing to do.
The INTERNAL classification tier reaches the wiki across two releases, and
they must be installed in order:
| Release | What it does |
|---|---|
| N — the release carrying PLT-502 | Every wiki build can read INTERNAL. Nothing can create one: the wiki.classification domain and every write schema still reject it. |
| N+1 — the release carrying PLT-977 | Widens the domain and the write schemas, so INTERNAL pages become creatable. |
Release N is the minimum version from which release N+1 may be installed.
Upgrading straight from N-1 to N+1 skips the waypoint: the outgoing build then
runs a four-value event validator while the incoming one can already publish an
INTERNAL payload, and the outgoing dispatcher rejects it. Deliveries are
retried rather than lost, but a repeatedly-failing row holds a slot at the head
of an oldest-first outbox scan until the old build retires, so a long-tailed
upgrade pays throughput for it.
Nothing enforces this mechanically on these tiers — the operator chooses when and from what version to upgrade, and this repository has no sequential- upgrade gate. That is why it is written here as a prerequisite rather than assumed. The reasoning, and the residuals it accepts, are in ADR-032 § The rolling-deploy story.
The exposure begins when somebody classifies a page INTERNAL, which is an
operator action against this documented prerequisite — not something the upgrade
performs on its own. A deployment that cannot accept that should not take the
tier: raise the case rather than skipping the waypoint, since the strategy lives
in the artifact and cannot be selected per deployment.
Environment variables per Vercel project
| Vercel project | Var | Value | Environments |
|---|---|---|---|
constellation-platform | NEXT_PUBLIC_BASE_PATH | /projects | Required on production. Standalone preview/dev unset (serves at /) |
constellation-catalog | NEXT_PUBLIC_BASE_PATH | /catalog | Required on production. Standalone preview/dev unset |
constellation-directory | PROJECTS_ZONE_URL | https://constellation-platform.vercel.app | Required on production. Optional elsewhere (falls back to localhost defaults) |
constellation-directory | CATALOG_ZONE_URL | https://constellation-catalog.vercel.app | Required on production. Optional elsewhere (falls back to localhost defaults) |
The resolveBasePath() helper in scripts/resolve-base-path.ts fails fast if NEXT_PUBLIC_BASE_PATH is missing on Vercel production — without it the sub-zone app deploys at / and breaks multi-zone routing.
Database credentials are scoped per project and per environment
A Vercel environment variable is scoped to an explicit set of environments, and in this repo each
app zone's DATABASE_URL / DIRECT_URL is set on that zone's own Vercel project. A deployment
reads whichever value is scoped to the environment it runs in — which may or may not be the same
value production reads, because one variable can be scoped across several environments.
Two consequences, both established while triaging INF-590:
- The blast radius of a bad credential is a per-project, per-environment question. The zones
that hold a
DATABASE_URLagainst the shared database are the ones with a Prisma datasource:apps/{directory,project-tracker,catalog,wiki,agents}/prisma/schema.prisma. Derive the set from those rather than from a hand-copied list — the zone table at the top of this page is not it, and is known to be incomplete. Two deployed projects are outside the set:constellation-docsis a static Docusaurus build reaching no database, andconstellation-pt-mcpkeeps its OAuth state in a dedicated Neon store underOAUTH_DATABASE_URL.DIRECT_URLis a separate, more privileged credential — the broadBYPASSRLSrole used by migrations, local scripts and theCRON_SECRET-gated cron-bypass routes at runtime — so it is not invalidated by rotating theconstellation_appruntime password, and rotating it reaches cron routes as well as migrations. - A database error from a preview deployment is not automatically a production incident — but
only the SERVER and EDGE tags can tell you that.
sentry.server.config.tsandsentry.edge.config.tsreadVERCEL_ENVat runtime, so theirenvironmenttag is the real deployment tier. ⚠️ Project Tracker'ssentry.client.config.tsreads the same variable inside the client bundle, where Next.js inlines onlyNEXT_PUBLIC_*— so in the browser it isundefinedand falls back toNODE_ENV, which isproductionon a preview build too. A browser error from a Project Tracker preview is therefore taggedenvironment: production(INF-605). The wiki's and the agents zone's browser tags are not affected: each app'snext.config.tsfixesNEXT_PUBLIC_ERROR_REPORTING_ENVIRONMENTfromVERCEL_ENVat build time (PLT-1323, PLT-1370), pending the production checks in PLT-1369 and PLT-1442. Each reports only onceNEXT_PUBLIC_SENTRY_DSNis set on its own Vercel project, and every event it sends is scrubbed to the adapter's allowlist: no exception message, breadcrumb or request URL leaves. The primary exception carries instead a description the scrub derives from the error's structure — its Prisma code and SQLSTATE, or its application error code and status, with a fixed label (PLT-1497). And even on the server the tag names the tier, not the credential: a variable scoped across several environments carries one value. Read the tag, note which runtime produced the event, then read that variable's scopes in the Vercel project before deciding the blast radius.
Note that a saved value reaches only deployments created after it, so a changed credential is not in use until the affected deployments are rebuilt. Writing the full rotation procedure — which scopes, in which order, and how each is verified — needs Vercel operational access and is not attempted here.
RLS-bypass startup guard (assertNonBypassRole)
Each app's src/instrumentation.ts calls assertNonBypassRole() (from @constellation-platform/db) on server startup. PostgreSQL silently disables every Row-Level Security policy — even with FORCE ROW LEVEL SECURITY set on every table — when the connected role has rolsuper = true or rolbypassrls = true. The guard is the fail-fast tripwire for that: if DATABASE_URL is ever repointed at a SUPERUSER/BYPASSRLS role, the app refuses to boot instead of silently serving cross-tenant data. Production connects as the non-bypass constellation_app runtime role, so the guard passes; it only trips on a regression.
| Var | Value | Effect |
|---|---|---|
CONSTELLATION_ALLOW_BYPASS_RLS | 1 | Fail-closed opt-out — attested CI only. Skips the guard ONLY on an attested ephemeral GitHub Actions runner: it takes effect only when CI=true and GITHUB_ACTIONS=true and DATABASE_URL is a loopback host (localhost/127.0.0.1/::1) and no Vercel signal (VERCEL/VERCEL_ENV) is present. Every deployment tier fails that attestation — SaaS (Vercel), Dedicated Cloud, and On-Prem (Docker/K8s) set neither CI var and use a remote database — so a stray value here is ignored on any deployment and the guard always enforces. (Vercel is additionally an absolute veto, even under full attestation.) Set it only where a privileged/bypass role is used deliberately (the CI end-to-end tests serving the apps as a superuser). Any other value, or unset, always enforces. |
Local dev needs nothing: .env.local.example already points the apps at constellation_app. Migration scripts legitimately connect as a privileged role and never call the guard.
Speed Insights — pinned paths, and where the samples are expected to land
Directory, Catalog and Project Tracker mount <SpeedInsights> directly in their root layouts
with both props pinned. Wiki mounts the platform adapter instead (apps/agents carries no
telemetry at all):
import { WebVitals } from '@constellation-platform/telemetry';
<WebVitals />;
WebVitals takes no props — it owns both transport paths. The three direct call sites still pass
them by hand:
<SpeedInsights
scriptSrc="/_vercel/speed-insights/script.js"
endpoint="/_vercel/speed-insights/vitals"
/>
Migrating those three to the adapter, and retiring the hand-passed props, is PLT-1106.
Neither prop is optional: dropping either stops collection, and the two fail differently.
The v2 SDK defaults
to a randomised /<hash>/… path for both the script and the collector. Under the multi-zone
rewrite the browser resolves those against the root zone's host, where the script path 404s and
the collector path falls through to the Next.js catch-all and answers HTML with status 200 —
so sendBeacon reports success, DevTools shows a green request, and Vercel records nothing.
A missing scriptSrc is at least loud — the 404 surfaces as a console error, which is how
INF-54 found it. A missing endpoint is the dangerous one: it fails silently and green, with
no console error and no failing gate, and the only symptom is a dashboard nobody is watching.
Both halves reached production once each and were fixed by hotfix (INF-54, #543 and #545).
Directory's rewrites cover /projects/*, /catalog/*, /wiki/* and /agents/* — but not
/_vercel/*. Both the script request and the vitals beacon therefore reach Directory's
deployment, so Directory's dashboard is where sub-zone field data is expected to land, and
the first place to look for it. That follows from the routing, but which project Vercel finally
credits has not yet been confirmed against a deploy; treat it as the expected outcome rather than
an established one until someone checks.
If it holds, it is dataset location and mixing, not data loss: the beacon carries the page path,
so /wiki/* and /projects/* routes stay separable within that project — filter by path. Giving
each zone its own project would need a zone-owned absolute endpoint (cross-origin, so CORS) or
the SDK's dsn prop, and is a cross-app decision rather than a per-zone one.
Vercel project wiring
The four Vercel projects share a build root (the monorepo) but each has its own Root Directory set to apps/<name> so Vercel runs each build from inside that workspace. installCommand and buildCommand are left blank on three of them (Vercel auto-detects); the docs project overrides them to use Turborepo, see apps/docs/vercel.json. All projects use the framework default output directory (.next for the three Next apps, build for Docusaurus) — apps/docs/vercel.json declares outputDirectory: "build" explicitly, but it matches the Docusaurus default.