Skip to main content

Deployment

Constellation is deployed on Vercel as a multi-zone topology:

ZoneHostVercel project
Directory (root)constellation.planetb2b.comconstellation-directory
Project Tracker/projects/* (rewritten)constellation-platform
Catalog/catalog/* (rewritten)constellation-catalog
Docsdocs.planetb2b.comconstellation-docs

Directory is the root zone and rewrites requests for /projects/* and /catalog/* to the other Vercel projects.

NEXT_PUBLIC_BASE_PATH​

Sub-zone apps read this env var at build time:

EnvironmentValueBehaviour
Production/projects or /catalogMulti-zone: served under prefix, rewrites from Directory
PreviewunsetStandalone: served at / for preview deploys
Local devunsetStandalone: turbo dev on app-local port

fetch() and basePath​

Next.js auto-prepends basePath to <Link> and router.push, but not to fetch(). Use the apiUrl() helper from each app's src/lib/api-url.ts:

import { apiUrl } from '@/lib/api-url';
const res = await fetch(apiUrl('/api/people')); // '/projects/api/people' in prod

Migration gate on release PRs​

Every PR targeting main (release branches and hotfixes) runs the Verify Release Migrations CI job. It fails the merge if either staging (DEV_DATABASE_URL) or production (PROD_DATABASE_URL) is missing a migration that's on the branch. The job is read-only — it never applies migrations. Apply outstanding migrations to both databases via the Supabase dashboard SQL editor or the Supabase MCP (apply_migration / execute_sql) before opening the release PR.

A migration containing CONCURRENTLY cannot go through a transaction

Scan the migrations you are about to apply first — recursively, since an entry can be a directory:

grep -rl CONCURRENTLY apps/<module>/prisma/migrations/<entry>

PostgreSQL rejects CREATE INDEX CONCURRENTLY / DROP INDEX CONCURRENTLY outright inside a transaction block, so a tool that wraps what it is given does not merely apply them more slowly — the migration fails. Those files are written deliberately without BEGIN;/COMMIT; to avoid holding an ACCESS EXCLUSIVE lock on a busy table for the length of a full index build, and they say so in their header.

Apply them with a session that does not wrap, and always with ON_ERROR_STOP=1 — psql's default is to report an error and carry on, which lets a migration fail its DDL and still run its final schema_migrations insert, recording an incomplete migration as applied:

psql "$PROD_DATABASE_URL" -v ON_ERROR_STOP=1 -f <file>

A migration entry can hold more than one file — 033_tasks_backlog_grooming is migration.sql plus post_migration.sql, and the grep points only at the one containing CONCURRENTLY. Apply every .sql in the entry in runner order (migration.sql first) and record the tracking row only once they have all succeeded; stamping the entry after applying one file marks the whole migration as done while the rest never ran.

The raw psql path also bypasses the module runner's tracking-row write, and not every migration self-inserts. Check the file for a schema_migrations insert and add one yourself if it has none, or the release gate stays red against a database that is already up to date — chained to the apply with && so it cannot stamp a migration that failed, into the owning module's schema (see scripts/check-migrations.ts), and keyed on the migration entry name (a directory-based migration is tracked by its directory, not by the inner .sql file).

If you use the dashboard or MCP instead, run the concurrent statement on its own and then confirm the index is valid: a failed concurrent build leaves an INVALID index behind, which a later re-run guarded by IF NOT EXISTS will skip straight past.

SELECT indexrelid::regclass, indisvalid FROM pg_index WHERE NOT indisvalid;

Full procedure: .claude/skills/release-and-migrations/SKILL.md § Migrations Before the PR.

See scripts/check-migrations.ts (invoked with --strict from the gate) for the implementation.

Platform packages carry their own migrations​

Migrations do not live only under apps/<module>/prisma/migrations. A shared package can own a schema too, and @constellation-platform/jobs owns jobs — its files are in packages/platform/jobs/migrations, applied by the package's own db:migrate, and tracked in jobs.schema_migrations like any module's. npx turbo run db:migrate picks it up with everything else; turbo.json orders @constellation/project-tracker#db:migrate after it, because both would otherwise create jobs.queue and race.

Baseline adoption on databases that already have jobs.queue​

Project Tracker's grandfathered 013_jobs_queue.sql created jobs.queue long before the package had a runner, so on staging, on production and on any machine that has run PT's db:setup, the table exists while jobs.schema_migrations does not. Re-applying 001_jobs.sql there is not what you want, and neither is skipping it silently.

The runner therefore baseline-adopts 001: it records the file as applied without executing it, but only after comparing the deployed table against a reference the same server builds from 001's own declaration — columns, constraints, indexes, RLS flags and every policy's command, kind, role scope and predicate. The reference is a closed set, so an object neither 001 nor 002 declares is refused rather than tolerated: an extra permissive policy is a cross-tenant read path that applying 001 would not remove, and an extra UNIQUE index rejects enqueues the schema permits.

Three outcomes, and each is loud:

  • matches — 001 is recorded as applied and never executed, then 002 runs normally;
  • differs repairably (a missing index, RLS not forced, a policy 001 recreates) — adoption is declined and 001 is applied, which is written entirely in IF NOT EXISTS / DROP POLICY IF EXISTS form;
  • differs unrepairably — the runner aborts and names the difference. Nothing is recorded. Reconcile the database by hand before re-running; a recorded migration is a claim every later migration acts on.

Every cross-tenant drain audits its claim and its settle​

Five drains claim jobs.queue rows across tenants: the wiki's /api/cron/process-jobs and Project Tracker's notification, agent-run, run-score and cycle-retrospective drains. Since PLT-1342 all five run one shared transaction machine, runInWorkerBypassTransaction in @constellation-platform/jobs, on the confined constellation_jobs_worker login. What an operator can rely on:

  • Every claim and every mutating settle writes an auditCritical() entry in its own transaction — wiki.job.claimed / wiki.job.settled, or projects.job.claimed / projects.job.settled — attributed to the job's tenant, with actor SYSTEM and resource type jobs.queue. The entry records lifecycle fields only (job type, status, attempts), never the payload or the last error.
  • An audit or outbox failure rolls the queue change back. A refused claim leaves the job pending for the next tick; a refused settle leaves it running for the stale-claim lease. So audit.audit_entries and events.outbox privileges are a hard prerequisite of the drain, which is why its preflight refuses to start without them.
  • A settle is refused, and writes nothing, when the job has moved on since it was claimed or observed — reclaimed at a higher attempts, or no longer running. The drain logs it and does not count the job as drained.
  • A suspended tenant's jobs are still claimed and audited. Attribution deliberately bypasses Project Tracker's tenant-liveness gate (ADR-044 § D7); gating it would stall the job type for every tenant. Org-less notifications are audited under the all-zeros tenant they are stored under, which no tenant session reads.

Runtime-role grants travel with the migration​

002_worker_claim_policy_and_runtime_grants.sql grants constellation_app schema USAGE plus SELECT, INSERT, UPDATE on jobs.queue — not DELETE — and read-but-not-write on jobs.schema_migrations. This is deliberate and it is why enqueuing works on a deployed database at all: RLS grants nothing on its own, and locally the privileges came from scripts/db-init.sql, which is a dev-only fixture. Without them the symptom is 42501 out of PostgresJobQueue.add() and out of every worker claim.

The grants are guarded on the role existing, and followed by hard postconditions — including a negative one, because REVOKE reports success whether or not a grant existed and so says nothing about write access reaching the role through PUBLIC or through a role it is a member of. If the migration role lacks authority to grant, the file raises instead of being recorded: the shared runner suppresses any statement that fails with 42501 and logs it as a skip, which would otherwise leave the ledger claiming work that never happened.

audit migration 013: apply it in one transaction​

Applies to Dedicated Cloud, On-Prem, and any database that has not yet applied packages/platform/db/migrations/013_audit_chain_heads_rls.sql — and to any database that records it but was never verified. SaaS staging and production applied it in one transaction on 2026-09-13, starting from no policies on the table, and its check passed on both; they need nothing. (PLT-1225)

013 turns on tenant row-level security for audit.chain_heads. It creates each of its three policies only when that policy name is missing. Then it enables and forces RLS. Only after that does it check that the table carries exactly those three policies with its definitions. The file opens no transaction of its own. So when it is applied one statement at a time (psql -f without -1), the ENABLE and FORCE have already committed when that check runs. If the table already had a same-named policy defined differently, or an extra permissive policy, the check fails but RLS stays on with that policy in force. For example, a leftover USING (true) select policy lets every role read every tenant's chain heads. A same-named policy that is stricter than 013's fails the same check and takes the same repair. Without ON_ERROR_STOP psql keeps going after the failure and records 013 as applied, so the migration ledger shows nothing wrong.

Before applying, check the table. On a database that has not applied 013 this returns f | f | {}. If policies is not empty, go to the repair below instead of applying:

SELECT c.relrowsecurity, c.relforcerowsecurity,
array(SELECT polname FROM pg_policy WHERE polrelid = c.oid ORDER BY polname) AS policies
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'audit' AND c.relname = 'chain_heads';

Apply 013 on its own, in one transaction, after 001–012:

psql "$DIRECT_URL" -X -v ON_ERROR_STOP=1 -1 \
-c "SET LOCAL lock_timeout = '5s'" \
-f packages/platform/db/migrations/013_audit_chain_heads_rls.sql

-1 puts the -c and the whole file in one transaction, so if the check fails the ENABLE is rolled back too. -X skips ~/.psqlrc, where settings such as AUTOCOMMIT or ON_ERROR_ROLLBACK could change what commits. Use -1 for this file only, not the whole directory: 004_audit_partitioning.sql opens and commits its own transaction, and -1 does not work as intended on a file that does.

Run from the repository root, npm run db:migrate also applies 013 in one transaction. It reaches packages/platform/db/migrations through Project Tracker's migration runner, which applies 013 with --single-transaction. It sets no lock timeout, though. On a live database its ACCESS EXCLUSIVE request can wait indefinitely behind an open transaction, with every audit write queued behind it, so use the command above there. A checkout from before PLT-1225 applies 013 one statement at a time; use the command above on those too.

Creating a policy and enabling RLS each take an ACCESS EXCLUSIVE lock on audit.chain_heads, held until the transaction ends. That blocks reads as well as writes, so audit writes that advance the hash chain wait for it. The lock request also queues behind any transaction already using the table. Apply in a quiet window. lock_timeout limits each lock wait on its own, not the apply as a whole. If any one lock cannot be acquired within 5 seconds, the apply gives up (exit 3, table unchanged). It does not bound how long audit writes are blocked. They wait while the apply waits for its locks, and those waits can add up past 5 seconds. Once it holds the table lock, they wait until its transaction commits or rolls back. It is set with SET LOCAL inside the transaction rather than through PGOPTIONS, which a connection string carrying its own options= silently overrides.

If 013 fails with tenant RLS was enabled on audit.chain_heads but recording the tracking row in audit.schema_migrations was refused, do not follow that message's advice to insert the row by hand. Under -1 the message is wrong: the rollback has already undone the ENABLE, so a hand-inserted row would record 013 with RLS off. Grant the applying role both privileges the ledger insert needs, INSERT and SELECT (filename) (its ON CONFLICT (filename) reads that column): GRANT INSERT, SELECT (filename) ON audit.schema_migrations TO <applying role>. Then apply again.

Verify, both after applying and on any database that already records 013 but was never checked. The first query must return t | t with the three names below, and this one must return exactly these rows:

SELECT policyname, permissive, roles, cmd, qual, with_check
FROM pg_policies
WHERE schemaname = 'audit' AND tablename = 'chain_heads'
ORDER BY policyname;
policynamepermissiverolescmdqualwith_check
chain_heads_tenant_isolationPERMISSIVE{public}SELECT((tenant_id)::text = current_setting('app.tenant_id'::text, true))empty
chain_heads_tenant_isolation_insertPERMISSIVE{public}INSERTempty((tenant_id)::text = current_setting('app.tenant_id'::text, true))
chain_heads_tenant_isolation_updatePERMISSIVE{public}UPDATE((tenant_id)::text = current_setting('app.tenant_id'::text, true))((tenant_id)::text = current_setting('app.tenant_id'::text, true))

Repair, when the preflight shows a policy, 013's policy check refuses it, or the verification fails. First inspect the policies that are there. Then replace them all in one transaction. The command drops every policy on the table and re-runs 013, which recreates its three and checks them again. It takes the same ACCESS EXCLUSIVE lock for the length of that one transaction, so run it in a quiet window too:

psql "$DIRECT_URL" -X -v ON_ERROR_STOP=1 -1 \
-c "SET LOCAL lock_timeout = '5s'" \
-c "DO \$\$ DECLARE p record; BEGIN FOR p IN SELECT polname FROM pg_policy WHERE polrelid = 'audit.chain_heads'::regclass LOOP EXECUTE format('DROP POLICY %I ON audit.chain_heads', p.polname); END LOOP; END \$\$" \
-f packages/platform/db/migrations/013_audit_chain_heads_rls.sql

Never insert the audit.schema_migrations row by hand to get past the check.

Why 013 itself is not fixed. It has shipped and been applied. The ledger matches migrations by filename alone, so an edited 013 would never run on a database that already recorded the original. A new migration cannot close the gap either: it runs after 013, when RLS has already been switched on.

Upgrade prerequisite: the wiki INTERNAL classification tier​

Applies to Dedicated Cloud and On-Prem. On SaaS the platform controls the rollout and there is nothing to do.

The INTERNAL classification tier reaches the wiki across two releases, and they must be installed in order:

ReleaseWhat it does
N — the release carrying PLT-502Every wiki build can read INTERNAL. Nothing can create one: the wiki.classification domain and every write schema still reject it.
N+1 — the release carrying PLT-977Widens the domain and the write schemas, so INTERNAL pages become creatable.

Release N is the minimum version from which release N+1 may be installed. Upgrading straight from N-1 to N+1 skips the waypoint: the outgoing build then runs a four-value event validator while the incoming one can already publish an INTERNAL payload, and the outgoing dispatcher rejects it. Deliveries are retried rather than lost, but a repeatedly-failing row holds a slot at the head of an oldest-first outbox scan until the old build retires, so a long-tailed upgrade pays throughput for it.

Nothing enforces this mechanically on these tiers — the operator chooses when and from what version to upgrade, and this repository has no sequential- upgrade gate. That is why it is written here as a prerequisite rather than assumed. The reasoning, and the residuals it accepts, are in ADR-032 § The rolling-deploy story.

The exposure begins when somebody classifies a page INTERNAL, which is an operator action against this documented prerequisite — not something the upgrade performs on its own. A deployment that cannot accept that should not take the tier: raise the case rather than skipping the waypoint, since the strategy lives in the artifact and cannot be selected per deployment.

Environment variables per Vercel project​

Vercel projectVarValueEnvironments
constellation-platformNEXT_PUBLIC_BASE_PATH/projectsRequired on production. Standalone preview/dev unset (serves at /)
constellation-catalogNEXT_PUBLIC_BASE_PATH/catalogRequired on production. Standalone preview/dev unset
constellation-directoryPROJECTS_ZONE_URLhttps://constellation-platform.vercel.appRequired on production. Optional elsewhere (falls back to localhost defaults)
constellation-directoryCATALOG_ZONE_URLhttps://constellation-catalog.vercel.appRequired on production. Optional elsewhere (falls back to localhost defaults)

The resolveBasePath() helper in scripts/resolve-base-path.ts fails fast if NEXT_PUBLIC_BASE_PATH is missing on Vercel production — without it the sub-zone app deploys at / and breaks multi-zone routing.

Database credentials are scoped per project and per environment​

A Vercel environment variable is scoped to an explicit set of environments, and in this repo each app zone's DATABASE_URL / DIRECT_URL is set on that zone's own Vercel project. A deployment reads whichever value is scoped to the environment it runs in — which may or may not be the same value production reads, because one variable can be scoped across several environments.

Two consequences, both established while triaging INF-590:

  • The blast radius of a bad credential is a per-project, per-environment question. The zones that hold a DATABASE_URL against the shared database are the ones with a Prisma datasource: apps/{directory,project-tracker,catalog,wiki,agents}/prisma/schema.prisma. Derive the set from those rather than from a hand-copied list — the zone table at the top of this page is not it, and is known to be incomplete. Two deployed projects are outside the set: constellation-docs is a static Docusaurus build reaching no database, and constellation-pt-mcp keeps its OAuth state in a dedicated Neon store under OAUTH_DATABASE_URL. DIRECT_URL is a separate, more privileged credential — the broad BYPASSRLS role used by migrations, local scripts and the CRON_SECRET -gated cron-bypass routes at runtime — so it is not invalidated by rotating the constellation_app runtime password, and rotating it reaches cron routes as well as migrations.
  • A database error from a preview deployment is not automatically a production incident — but only the SERVER and EDGE tags can tell you that. sentry.server.config.ts and sentry.edge.config.ts read VERCEL_ENV at runtime, so their environment tag is the real deployment tier. ⚠️ Project Tracker's sentry.client.config.ts reads the same variable inside the client bundle, where Next.js inlines only NEXT_PUBLIC_* — so in the browser it is undefined and falls back to NODE_ENV, which is production on a preview build too. A browser error from a Project Tracker preview is therefore tagged environment: production (INF-605). The wiki's and the agents zone's browser tags are not affected: each app's next.config.ts fixes NEXT_PUBLIC_ERROR_REPORTING_ENVIRONMENT from VERCEL_ENV at build time (PLT-1323, PLT-1370), pending the production checks in PLT-1369 and PLT-1442. Each reports only once NEXT_PUBLIC_SENTRY_DSN is set on its own Vercel project, and every event it sends is scrubbed to the adapter's allowlist: no exception message, breadcrumb or request URL leaves. The primary exception carries instead a description the scrub derives from the error's structure — its Prisma code and SQLSTATE, or its application error code and status, with a fixed label (PLT-1497). And even on the server the tag names the tier, not the credential: a variable scoped across several environments carries one value. Read the tag, note which runtime produced the event, then read that variable's scopes in the Vercel project before deciding the blast radius.

Note that a saved value reaches only deployments created after it, so a changed credential is not in use until the affected deployments are rebuilt. Writing the full rotation procedure — which scopes, in which order, and how each is verified — needs Vercel operational access and is not attempted here.

RLS-bypass startup guard (assertNonBypassRole)​

Each app's src/instrumentation.ts calls assertNonBypassRole() (from @constellation-platform/db) on server startup. PostgreSQL silently disables every Row-Level Security policy — even with FORCE ROW LEVEL SECURITY set on every table — when the connected role has rolsuper = true or rolbypassrls = true. The guard is the fail-fast tripwire for that: if DATABASE_URL is ever repointed at a SUPERUSER/BYPASSRLS role, the app refuses to boot instead of silently serving cross-tenant data. Production connects as the non-bypass constellation_app runtime role, so the guard passes; it only trips on a regression.

VarValueEffect
CONSTELLATION_ALLOW_BYPASS_RLS1Fail-closed opt-out — attested CI only. Skips the guard ONLY on an attested ephemeral GitHub Actions runner: it takes effect only when CI=true and GITHUB_ACTIONS=true and DATABASE_URL is a loopback host (localhost/127.0.0.1/::1) and no Vercel signal (VERCEL/VERCEL_ENV) is present. Every deployment tier fails that attestation — SaaS (Vercel), Dedicated Cloud, and On-Prem (Docker/K8s) set neither CI var and use a remote database — so a stray value here is ignored on any deployment and the guard always enforces. (Vercel is additionally an absolute veto, even under full attestation.) Set it only where a privileged/bypass role is used deliberately (the CI end-to-end tests serving the apps as a superuser). Any other value, or unset, always enforces.

Local dev needs nothing: .env.local.example already points the apps at constellation_app. Migration scripts legitimately connect as a privileged role and never call the guard.

Speed Insights — pinned paths, and where the samples are expected to land​

Directory, Catalog and Project Tracker mount <SpeedInsights> directly in their root layouts with both props pinned. Wiki mounts the platform adapter instead (apps/agents carries no telemetry at all):

import { WebVitals } from '@constellation-platform/telemetry';

<WebVitals />;

WebVitals takes no props — it owns both transport paths. The three direct call sites still pass them by hand:

<SpeedInsights
scriptSrc="/_vercel/speed-insights/script.js"
endpoint="/_vercel/speed-insights/vitals"
/>

Migrating those three to the adapter, and retiring the hand-passed props, is PLT-1106.

Neither prop is optional: dropping either stops collection, and the two fail differently. The v2 SDK defaults to a randomised /<hash>/… path for both the script and the collector. Under the multi-zone rewrite the browser resolves those against the root zone's host, where the script path 404s and the collector path falls through to the Next.js catch-all and answers HTML with status 200 — so sendBeacon reports success, DevTools shows a green request, and Vercel records nothing.

A missing scriptSrc is at least loud — the 404 surfaces as a console error, which is how INF-54 found it. A missing endpoint is the dangerous one: it fails silently and green, with no console error and no failing gate, and the only symptom is a dashboard nobody is watching. Both halves reached production once each and were fixed by hotfix (INF-54, #543 and #545).

Sub-zone samples are expected under the ROOT zone's Vercel project

Directory's rewrites cover /projects/*, /catalog/*, /wiki/* and /agents/* — but not /_vercel/*. Both the script request and the vitals beacon therefore reach Directory's deployment, so Directory's dashboard is where sub-zone field data is expected to land, and the first place to look for it. That follows from the routing, but which project Vercel finally credits has not yet been confirmed against a deploy; treat it as the expected outcome rather than an established one until someone checks.

If it holds, it is dataset location and mixing, not data loss: the beacon carries the page path, so /wiki/* and /projects/* routes stay separable within that project — filter by path. Giving each zone its own project would need a zone-owned absolute endpoint (cross-origin, so CORS) or the SDK's dsn prop, and is a cross-app decision rather than a per-zone one.

Vercel project wiring​

The four Vercel projects share a build root (the monorepo) but each has its own Root Directory set to apps/<name> so Vercel runs each build from inside that workspace. installCommand and buildCommand are left blank on three of them (Vercel auto-detects); the docs project overrides them to use Turborepo, see apps/docs/vercel.json. All projects use the framework default output directory (.next for the three Next apps, build for Docusaurus) — apps/docs/vercel.json declares outputDirectory: "build" explicitly, but it matches the Docusaurus default.