Multi-LLM PR review
A cross-family second opinion on every PR. Runs in CI, posts findings as inline diff comments anchored to the flagged lines, and never approves or requests changes — humans own merge.
Why a second opinion
Same-family review (Claude reviewing Claude) shares blind spots. A different model family catches different things — rule hallucinations, missed tenancy wrapping, dropped error paths, security regressions one family is structurally biased toward missing. The cost is one extra LLM call per PR push, bounded by a per-PR token budget; the benefit is a second pair of eyes that consistently disagrees with the first.
Independence depends on which reviewers actually run for the PR class in front of you, and that is not uniform. The authoritative matrix is AGENTS.md § Which review gates cover which PR — read it there rather than reasoning from this page, which is about the roster rather than about coverage. Three consequences bear on the roster choice:
- On an ordinary source PR to
develop, the cross-model loop pairs the families by direction (Claude-led reviewed by Codex, Codex-led by Claude), so by merge both families will have seen the diff. A roster from either is then a second opinion from a family the PR would consult anyway — which is what makes family separation, rather than headline capability, the interesting axis. Timing caveat: this job runs on every push while that loop is agent-initiated at final HEAD, so the claim is about the eventual pre-merge review set, not about what has already happened when this job fires. - On a PR where the loop is skipped — the
codex-reviewskill scopes the mandate to source changes and allows[skip codex-review]on docs-only, config-only and trivial diffs, while a docs or config diff is not one of this workflow's exemptions — so for an eligible PR this job may be the only independent review, and an OpenAI roster is not redundant at all. This workflow has deliberate skips of its own (drafts, forks, Dependabot,release/*andhotfix/*heads, a base other thandevelopor anepic/*branch, a zero-file diff, and[skip ai-review]; § How to skip a PR and § Failure modes are authoritative). A PR that hits both skip sets gets no independent review from either, however green its checks look. - ⚠️ On
release/*andhotfix/*heads, this workflow is skipped too. For a hotfix,Verify Codex Evidenceis exempt as well, so neither required independent-review gate covers that diff — the workflow's own skip message puts it more bluntly still ("NOTHING independently reviews this diff — review it by hand"), though strictly what is absent is the gates: GitHub's native Copilot auto-review is an org setting outside this workflow and, when enabled, still reviews every push. Treat a hotfix as ungated and review it by hand. The roster is irrelevant there; do not read this page as assurance that those PRs were covered.
So the roster should be defensible under both of the first two cases: where the loop runs, family separation is the axis that matters; where it does not, the roster's own review quality is carrying the PR alone.
INF-63 shipped the workflow scaffold. INF-126 shipped the reviewer script that the workflow invokes. INF-155 added SSE streaming, inline diff comments, and neutral check-run signalling. The reviewer has been enabled in CI since 2026-06-02 (initially on openai/gpt-5.5) and runs on every PR push. The roster has changed since — see Current roster for the live value and How it was enabled for historical context.
What it does
On every PR push (opened, synchronize, reopened, ready_for_review):
.github/workflows/multi-llm-review.ymlchecks gating (feature flag, secret, escape-hatch markers, draft / fork / dependabot exclusions).- If gating passes, it invokes
scripts/multi-llm-review.tswith the PR number, base SHA, head SHA, the configured model list, and the per-PR token budget. - The script:
- Computes
git diff <base>...<head>. - Reads
.ai/constitution.mdand any.ai/specs/SPEC-*.mdpaths it can extract from the PR body (regex + existence-check). - For each configured model: builds a prompt (constitution + specs + diff + reviewer rules), streams the response from the Vercel AI Gateway (SSE), parses the model's findings from a fenced JSON block, and posts findings as inline diff comments via the GitHub Reviews API.
- Tracks token usage across models; aborts cleanly with a "budget exhausted" notice on the PR review if the budget would be exceeded.
- When findings exist, creates a neutral check-run in the PR checks panel so findings are visible at a glance without blocking merge.
- Computes
Every finding is required by the prompt to cite the spec heading or AGENTS.md / constitution clause it flags. Uncited findings are dropped on the client side — reviewer noise is the failure mode we care most about.
How findings are posted
Findings are posted as inline diff comments anchored directly to the changed line where the issue was found. This means reviewers see the annotation right on the relevant code, not in a separate comment thread.
Anchorable findings — those whose path:line location falls within an actual changed hunk — appear as inline comments on the RIGHT (new) side of the diff.
Non-anchorable findings — those with no location, a hallucinated path, or a line number outside any changed hunk — fall back into the top-level review body summary. This ensures that off-target model guesses do not cause the entire review POST to fail (the GitHub API rejects inline comments on lines not in the diff).
A neutral check-run named Multi-LLM review findings is created in the PR checks panel whenever findings exist. Neutral never blocks merge — it provides an informational signal ("N findings posted — advisory") without turning the check red. Clean reviews (zero findings) and skipped/error runs do not create a check-run; the existing green status is preserved.
How the idle-timeout model works (streaming)
Gateway calls use SSE streaming with an idle/inactivity timer rather than a hard wall-clock timeout. The timer resets on every received chunk — a slow-but-progressing reasoning trace is never killed. Only a truly stalled upstream (no bytes arriving for GATEWAY_IDLE_TIMEOUT_MS = 60_000 ms) aborts the request.
This is strictly better than the previous wall-clock cap for reasoning models like openai/gpt-5 that may take well over 60 seconds to produce a full response on large diffs, but do so steadily without stalling. A stall (network partition, gateway unresponsive mid-generation) is still caught quickly via the idle timer.
The workflow's timeout-minutes: 15 provides an outer wall-clock guard for the job as a whole (raised from 8 by INF-249 — a sequential agentic panel can legally spend 5 × 180 s exploration calls plus two extraction attempts per model). Inside it, a soft run deadline (MULTI_LLM_REVIEW_RUN_DEADLINE_MS, workflow default 10 min) finalizes the run gracefully instead of letting the job timeout kill it mid-model: models that have not started are skipped with an explicit reason (a partial panel still synthesizes and posts), and a mid-review expiry stops tool exploration and jumps straight to the forced findings extraction (stoppedReason: 'deadline'). Per-call gateway timeouts under a deadline are clamped to the remaining run time plus a 30 s finalization grace, so a call admitted just before expiry cannot drag the run past the job timeout.
How it was enabled (one-time repo-admin step)
The reviewer was enabled on 2026-06-02. The one-time setup was:
- Generated a Vercel AI Gateway token with access to
openai/gpt-5.5and added it as theVERCEL_AI_GATEWAY_TOKENrepo secret. - Set
MULTI_LLM_REVIEW_ENABLED=trueas a repo variable. - Set
MULTI_LLM_REVIEW_MODELS=openai/gpt-5.5as a repo variable.
The reviewer now runs on every PR push automatically. To tune the knobs:
MULTI_LLM_REVIEW_MODELS— comma-separated model ids. The repo variable is the live roster; see Current roster.MULTI_LLM_REVIEW_TOKEN_BUDGET— max total tokens per PR push. Default:100000.
Current roster
MULTI_LLM_REVIEW_MODELS is the live roster, and the repo variable is the only source of truth. This page records what it was set to and why; it is not consulted at runtime, so treat a value here as dated rather than current, and read the variable itself (gh variable list --repo B2B-Online/constellation) before acting on it. The workflow's inline default (openai/gpt-5) applies only when the variable is unset — it is a fallback, not the configured roster.
| Set | Roster | Why |
|---|---|---|
| 2026-06-02 | openai/gpt-5.5 | Initial enablement (INF-63 / INF-126). |
| 2026-07-15 | openai/gpt-5.5,openai/gpt-5.6-sol | Quality candidate rostered alongside 5.5, not replacing it (INF-249). With PANEL=false and two models each posts its own separate review — 10 + 10 on #1419 — so metrics over this window blend two models unless partitioned by the **Model:** line. Needed the Responses-API adapter: Chat Completions function tools require effective reasoning none while 5.6 defaults to medium. |
| 2026-07-20 | openai/gpt-5.6-sol | Dual roster collapsed to the single quality candidate. |
| 2026-09-21 | spacexai/grok-4.7 | Family separation from the Claude implementer and the Codex reviewer (INF-646). Reverted the same day — see below. |
| 2026-09-21 | openai/gpt-5.6-sol | Rollback. Grok could not finish a review of a real code PR inside the agentic path's 180 s per-call gateway timeout (INF-648). |
| 2026-09-23 | openai/gpt-6-sol | Successor to the 5.6 family: same reasoning + tools shape, half the price ($2/$10 per M vs $4/$20), 1.05 M context. Needed a MULTI_LLM_REVIEW_RESPONSES_MODELS prefix added — see below. |
Neither roster change needed a code change, but they needed opposite variable changes, and the difference is the thing to check before rostering anything.
⚠️ MULTI_LLM_REVIEW_RESPONSES_MODELS is a list of literal prefixes, matched with startsWith — so a new OpenAI family does not inherit the previous family's routing. openai/gpt-6-sol does not start with openai/gpt-5.6, so rostering it against the old prefix list would have sent a reasoning-plus-function-tools model down the Chat Completions path that INF-249 built the Responses adapter to avoid. The variable is now openai/gpt-5.6,openai/gpt-6, which covers both families and keeps a rollback to openai/gpt-5.6-sol a single-variable edit.
⚠️ That mistake would not have announced itself. Probing openai/gpt-6-sol against Chat Completions with the real forced-submit_findings body returns HTTP 200 with a well-formed findings block and a non-zero reasoning_tokens count — it does not 4xx, and it does not visibly misbehave on a small diff. The cost of getting this wrong is a quieter, differently-configured review, not a red check. Set the routing prefix before the roster, so no window exists in which the model is rostered unrouted.
By contrast spacexai/grok-4.7 matched no prefix deliberately: it runs Chat Completions natively. The general rule is to read the model's supported_parameters and reasoning_options at https://ai-gateway.vercel.sh/v1/models — the catalog is public and needs no auth — and decide the routing explicitly, rather than assuming the previous roster's routing still applies.
The 2026-09-21 grok-4.7 attempt, and why it was reverted
Recorded because the failure is instructive and cheap to repeat.
Sequence, because the order is the lesson. The roster variable was flipped at 18:11 UTC; the benchmark replay only started at 18:20, and the rollback came at 20:25. The change went live before it was measured — the measurement was a justification run after the fact, not a gate ahead of it. Both rosters were then replayed against all 48 mined-2026-06 cases; the harness matched 17/48 for grok against 10/48 for sol, with the harvested Copilot baseline at 2/43.
⚠️ Those are automated match scores, not recall. findingMatchesCase requires exact file, line overlap and a defect_class keyword as a whole word, so it scores vocabulary alongside detection (INF-647). Adjudicating the unmatched findings, 5 of the 7 cases scored grok-only had a sol finding in the right place the keyword gate rejected, and counting every right-place finding gives grok 26/48 vs sol 24/48 — noise on n=48. The benchmark never supported a detection advantage.
🔴 What the benchmark could not show, and what actually decided it. In the ~2h15m the roster was live, 8 of 8 runs on real code PRs failed with agentic error: Gateway timed out after 180000ms → no reviewer posted a review, while 4 of 4 runs on a docs-only PR succeeded in about two minutes. Two PRs received no independent review at all.
Three lessons for the next roster change:
- Per-call latency on a real diff is a rostering constraint, and the benchmark's workload hid it rather than its instrumentation. The 180 s per-call
GATEWAY_TIMEOUT_MSis active in the replay harness too — replay passes notimeoutMs, so every benchmark call inherits it. What replay omits is the separate 600 s run deadline. The benchmark did not trip the per-call ceiling because its cases are single-file diffs whose individual calls run nowhere near 180 s, so a benchmark built on realistically-sized diffs would have caught this. The 50 s/case for grok against 14.7 s for sol was the available signal, and it was read as a cost note rather than a feasibility limit. - A docs-only PR is the worst possible smoke test, because the constraint is diff-size dependent — the cheapest check is exactly the one that passes.
- Red is not one thing, and both target branches were red throughout — for three different reasons. Before the swap they failed at
Verify independent-reviewer verdict: the review ran, posted, and the gate flagged blocking findings. While grok was rostered they failed atPublish terminal review statuswith no review posted at all. After the rollback they returned to failing at the verdict gate — review running again. A before/after comparison on the check's colour therefore shows no regression across any of it. Compare the cause, always.
The family-separation and cost arguments survive intact; only the latency blocked it.
⚠️ A third-family roster is PAUSED, and this page is not authorization to retry it. INF-646 was paused by an explicit human decision on 2026-09-21, not merely blocked on a dependency. INF-648 — making the agentic per-call budget an idle timer, the treatment INF-155 already gave the streaming path, and configurable — is a prerequisite for any retry, not a trigger for one. Resuming requires a fresh, separately approved decision.
⚠ The xAI namespace on the gateway is spacexai/…, not xai/…. There is no xai/ prefix; an id written that way 404s.
Panel mode stays off across all of these. Running several models and posting the union lifts recall but regresses precision (INF-171), so the roster is one model at a time. Panel is a measurement tool, not the default posture.
Agentic-mode knobs (INF-169 / INF-207 / INF-170 — gated behind MULTI_LLM_REVIEW_AGENTIC)
The agentic arm (read beyond the diff, ground in repo rules, defer to SonarCloud) is off by default. When MULTI_LLM_REVIEW_AGENTIC=true:
MULTI_LLM_REVIEW_MAX_TOOL_TURNS/MULTI_LLM_REVIEW_MAX_READ_FILE_BYTES/MULTI_LLM_REVIEW_PER_REVIEW_TOKEN_CEILING/MULTI_LLM_REVIEW_KEEP_RECENT_TOOL_RESULTS— the INF-207 cost levers. Unset ⇒ the tuned script defaults (5 turns, 8KB reads, 30k ceiling, keep-recent 2).MULTI_LLM_REVIEW_GROUNDING_TOKEN_BUDGET(INF-170) — token budget for the repo rule-grounding block (AGENTS.md hard rules + linked-spec ACs + lessons + the SonarCloud summary). Unset ⇒ the script default.SONAR_TOKEN(secret, INF-170) — SonarCloud user token for the division-of-labour read (PR issues + hotspots, summarised into grounding; LLM findings overlapping a Sonar issue on file+line are deduped). Unset ⇒ the Sonar read/dedupe is skipped cleanly — the review still runs.SONAR_PROJECT_KEY(var, INF-170) — overrides the project key (env →sonar-project.properties→ the documented defaultB2B-Online_constellation).
Linked task-thread context (INF-465)
When PT_BASE_URL and PT_AUTH_TOKEN are also configured, the grounding source
resolves a linked Project Tracker task from a bracketed key in the PR title,
body, or branch. It sends the task description plus at most 20 non-deleted task
comments, newest-first, through the configured Vercel AI Gateway as part of the
review prompt. The description is trusted requirement grounding; comment bodies
are separately bounded, single-line encoded, JSON-quoted, and placed only in the
user message under a fixed system rule that treats them as untrusted data and
forbids them from suppressing findings.
The thread is budgeted from its own reserved slice of
MULTI_LLM_REVIEW_GROUNDING_TOKEN_BUDGET — 15% of the total, capped at 750
tokens and zero below a 1,500-token total, so 600 tokens at the 4,000-token
default. The slice is a pure function of the total, so a chatty task can reduce
the 20-comment slice but can never displace the hard rules, spec ACs, ticket, or
lessons the reviewer grounds on. For the same reason the thread's content is
excluded from the INF-293/INF-300 review-input fingerprint: while a thread is
being sent, further comments — new, edited, or removed — do not force a full
re-review.
What the fingerprint does carry is a single bit for whether a thread is being sent at all. So a full review is forced when that bit flips: the first comment arriving on a task that had none, or the last one going away. That flip adds or removes the trust contract in the reviewer's system rules, and a review scoped to "confirm the previous round" must not inherit rules it never ran under.
Read "being sent" literally — it means deliverable, not "the task has comments", because the bit is taken after budgeting. A thread whose newest comment alone cannot fit the reserve is omitted, so an edit that grows a comment across that size also flips the bit and costs one full review. Churn is free only while the thread stays deliverable.
Deliverability today is decided by the newest comment alone: records are
newest-first and only a leading run of whole records is kept, so one comment
larger than the reserve omits the entire thread — including shorter older
comments that were being sent before it arrived. At the 4,000-token default that
threshold is around 2,100 characters, which ordinary session-summary comments
exceed. The run logs PT task thread omitted when this happens. Whether this is
the right retention policy is tracked in INF-488.
The slice is only charged when a thread actually reaches the prompt. Every omission path below — and a thread too large for even one complete comment record plus the truncation marker — leaves the reviewer grounded byte-identically to a build without this feature, so nothing is paid for context that was never sent. The marker is counted deliberately: a thread trimmed to fit would otherwise reach the model looking complete, so a slice that leaves no room to say "this was truncated" is dropped rather than sent unlabelled.
Missing credentials, no resolvable task, or a failed task fetch omits this data without failing the review. An unavailable or malformed task/comment payload likewise fails open: invalid comment entries are ignored while valid ticket and thread context continues when possible, and the thread is omitted when no valid comment data remains.
A recorded deferral is an annotation, not a waiver. A matching finding stays
in the inline and summary output with a visible deferred to KEY-N note, keeps
its severity, and still counts against the independent-review verdict
(scripts/check-review-verdict.ts, reported under the Multi-LLM review check)
— so annotating a high-severity finding as deferred does not turn that check
green. The documented own-line Review override: entry in the PR body is the
mechanism for that, and it is deliberately a separate, attributable act.
An override names the finding by its FILE, not its title (INF-520). The
reviewer re-words a finding's title between runs — eight distinct wordings across
twelve postings of one
finding on PR #1811, five on PR #2049, with the reported file identical on every
posting — so a title-bound override quotes a string the next run regenerates, and
the loop cannot converge. scripts/check-review-verdict.ts therefore matches an
override line that contains the complete file path parsed out of the finding's
location, using the same parseFindingLocation the synthesizer uses, so the
gate and the reviewer cannot drift on what a location means. Quoting the current
title still binds, and is allocated before any path match, so a line written
for one finding is never spent on another finding that merely shares its file. A line that names more than one blocking finding names none of them and is
refused outright — whether because several findings share the path it cites, or
because it cites several findings' files. It is reported as ambiguous-path and
the author quotes a title instead, which is the one case that keeps the
re-titling problem. A path binds to the file, so the check raises a workflow warning —
Review override bound by FILE PATH, not by title — for every override
allocated that way:
if a rerun on the same head replaces the finding in that file with a different
one, the line still binds and its argument no longer addresses what it excuses.
No repo-local check can tell that apart from a re-title, so it is surfaced for a
human instead of hidden — and it needs a reviewer that changes its finding set on
unchanged code, which is tracked separately as INF-593. A finding
reported with no usable location remains title-bound, and a location that is not
a plausible repo-relative path (..., ../x, https://x) yields no path
selector at all. A token cites the path when it READS as that location, so
apps/x/foo.ts:190-214 — the location exactly as the reviewer reports it —
counts and apps/x/foo.ts.backup does not. Everything else is unchanged: the line still cites the head
commit (so a push invalidates it), still costs a written argument measured after
every identifier is stripped, and is still spent on exactly one finding. The
failing check prints the ready-to-paste line per finding, and any override that
excuses nothing is reported with the reason — on a passing verdict too.
Panel-mode knobs (INF-171 — gated behind MULTI_LLM_REVIEW_PANEL)
Panel mode runs a panel of ≥2 distinct models independently over the same diff and reconciles their findings into one deduped, agreement-ranked review via a deterministic synthesizer. Each model runs through the same per-model path the single-model reviewer uses — diff-only by default, or the agentic/grounded path when MULTI_LLM_REVIEW_AGENTIC=true (panel composes with the agentic flag; it does not force it on). Panel is off by default; the single-model path is unchanged when off. When MULTI_LLM_REVIEW_PANEL=true and MULTI_LLM_REVIEW_MODELS lists ≥2 distinct ids:
MULTI_LLM_REVIEW_MODELS— the panel roster (e.g.openai/gpt-5.5,anthropic/claude-sonnet-4.6,google/gemini-2.5-pro). Duplicate ids collapse for the panel gate; fewer than 2 distinct ⇒ panel does not engage and the non-panel per-model loop runs unchanged (a pathological duplicate roster likem1,m1still runs the per-model loop verbatim).MULTI_LLM_REVIEW_PANEL_MIN_AGREEMENT— minimum cross-model agreement a synthesized finding must have to be posted. Default1(keep every finding); raise it to trade recall for precision.
Confirmation-mode knobs (INF-293 — gated behind MULTI_LLM_REVIEW_DELTA)
Every synchronize push has historically triggered a full fresh review of the whole BASE_SHA...HEAD_SHA diff, so a fix-push samples the same mostly-unchanged code again and surfaces a fresh batch of unrelated findings each round. Confirmation mode makes a follow-up push scoped to just the fix delta, revalidating prior findings instead of re-deriving them. Off by default:
MULTI_LLM_REVIEW_DELTA(repo var) —'true'opts in. Unset/'false'⇒ delta scoping is off: no state comment is read or written, no fingerprint is computed, and every run is a full review of the whole PR diff. Some behaviour is always on, independent of this flag: the INF-292 coverage-contract paragraph in the reviewer's system prompt (it sharpens the initial pass but touches no code path or API call); the INF-172review_modemetrics field; and the INF-300 concurrency rules below — push runs and manual re-runs occupy separate concurrency groups, and a re-run never writes reviewer state.- The state comment. Each completed run (including a clean zero-findings one) upserts one hidden-marker issue comment on the PR — marker
<!-- multi-llm-review-state:v1 -->followed by a fenced JSON block recordinglastReviewedHeadSha,reviewedAt,mode, and this run'sfindings[](title, severity, optional location, and a truncateddetail). An issue comment is used deliberately — it creates no resolvable review thread and is trivially machine-findable. Only a marker comment authored bygithub-actions[bot](the workflow's own token identity) is ever trusted as state; a marker comment from any other author is ignored and logged, so no PR commenter can forge a fakelastReviewedHeadShato shrink review scope. An unparsable or missing state comment degrades to a full review, never a crash. - Delta scope. A push resolves to a confirmation review only when ALL of: the flag is
'true', a trusted state comment parsed successfully, itslastReviewedHeadShadiffers from the new head, andgit merge-base --is-ancestorconfirms the old head is still an ancestor of the new one. The reviewed diff becomeslastReviewedHeadSha...HEAD_SHA(just the fix delta), and the prompt embeds the prior findings plus rules: revalidate each one (re-reporting an unfixed finding with its title prefixedSTILL OPEN:), report any new issue the delta introduces, and — the escape hatch — report a newly-noticed critical/high-severity issue anywhere in view even outside the delta. It must NOT raise a fresh non-critical finding in code the delta didn't touch. - Fallback-to-full conditions. Any of: the flag off, no state comment (first review, or the marker comment failed to parse / was untrusted), the head unchanged since the last review on an event other than
edited, a broken ancestry chain (a force-push rewrites history), or a merge commit in the delta range (merging an updated base branch into the feature branch would flood the "fix delta" with unrelated upstream code) — all fall back to today's full-PR review, exactly as if the flag were off. A same-headeditedevent is the narrow INF-431 exception and still must pass the fingerprint and reviewer-roster trust checks described below. - State only advances after a completed review. If every configured model is skipped (gateway error, token budget, run deadline, or a diff-only response with no explicit findings block — prose/malformed output counts as incomplete in BOTH modes, mirroring the agentic rule, per INF-297), the state comment is left untouched — the next run re-reviews from the old SHA so a skipped diff can never escape review. An empty full-review diff clears the persisted findings and blanks the stored fingerprint, forcing the next run to a full review (the whole scope is clean); an empty confirmation delta carries them forward unchanged. Before every write the run confirms its head AND base ref are still the PR's live head/base (fail-closed), so a manually re-run older workflow execution — including one from before a retarget — cannot roll state backwards.
- What busts the fingerprint (INF-300 / INF-431). Alongside the head SHA, a run fingerprints every input that determines what "a defect" means for this PR: the PR title, review-relevant body text, branch, base ref and merge-base, the effective reviewer configuration, the linked specs, and a digest of all resolved prompt grounding —
.ai/constitution.md, the post-impl-review checklist, the repo hard-rule catalog,.ai/lessons.md, and (agentic mode) the live PT ticket and knowledge-base retrieval. Verdict-layer own-lines are deliberately removed before the body reaches linked-spec discovery, PT/KB grounding, or the hash:Review override:entries and strict own-line[skip ...]markers decide what a gate does with an existing review; they are not requirements for the model. Change any review-relevant input mid-PR and the next push falls back to one full review, so the already-reviewed code is re-checked against the new rules rather than being frozen behind confirmation mode's "no new non-critical findings in unchanged code" restriction. The SonarCloud summary is deliberately excluded from that digest: it is per-commit analysis output the head SHA already keys, and Sonar re-analyses on every push, so including it would bust the fingerprint every round and make confirmation mode unreachable. Excluding it from the hash is only half the job, though — while Sonar shared a token budget with the requirement chunks it could push a whole rule, ticket or KB chunk out of the prompt without the digest noticing. So the Sonar summary and the requirement grounding are budgeted in independent pools (a fixed Sonar reserve, andtotal − reservefor requirements). The requirement budget is a constant, so Sonar's size — or its absence — cannot change which requirements the model receives, and the digest is taken over that received set. Because the digest is resolved before the mode is, grounding retrieval is seeded by the full PR diff's changed paths in both modes — what the model sees is still delta-scoped. The fingerprint also covers the reviewer's own source, as the prompt contract it defines: the static reviewer rules, tool guidance and grounding header shape what counts as a defect just as much as the constitution does, and enumerating those constants individually would rot the moment someone added a new one. Four modules are hashed —scripts/multi-llm-review.ts,scripts/review-grounding.ts(which holds the grounding header and the hard-rule catalog),scripts/review-sonar.tsandscripts/review-synthesizer.ts— and a test asserts that list covers every sibling reviewer module the script imports, so a new one cannot escape the contract. Any edit to any of them forces one full review per open PR: the safe direction, and rare. - Verdict-only confirmation runs (INF-431). Adding or editing only verdict-layer own-lines leaves the normalized fingerprint unchanged. When the PR head still equals the trusted state's head, the
editedrun takes an empty confirmation delta, posts no LLM review, and records the distinctverdict_onlyterminal outcome. The workflow then runs the independent-review verdict against the existing state. An override quoting the state comment's current finding can therefore clear it without a fresh non-deterministic review re-titling the finding first. GitHub workflow re-runs replay their original event payload, so re-running thateditedexecution repeats the verdict refresh and stays read-only for reviewer state; manual re-runs of other same-head event types still run a full review. The verdict gate reads the live PR body, so a removed or edited override is not resurrected from that replayed payload. If the live read fails, a first attempt warns and falls back to its event body; a re-run ignores the known-stale payload's overrides and stays red until the body can be verified. A changed-head confirmation whose commits have zero net delta also emitsverdict_only, carrying the trusted findings forward to the new head and refreshing the gate. Whitespace is not substantive (INF-593): the body is compared after line endings, trailing spaces, blank-line runs and leading or trailing blank lines are canonicalized, so neither the newline agh pr view --jq .body/gh pr edit --body-fileround trip appends nor the blank lines left where a stripped override paragraph stood can turn the refresh into a full review. Every other edit still busts the fingerprint and buys one full review: any other body change (an edited own-lineCodex review:marker included — record that marker as a PR comment, per INF-485), a title change, or a change to trusted grounding such as the PT ticket's title or description. A task-thread comment does so only when it flips whether a thread is sent at all — the first comment, the last removal, or crossing the thread's reserve — because the thread's content is not hashed. The no-model path does not require gateway credentials, and the check summary says verdict refreshed, not reviewed, so state reuse is never presented as new review coverage. - Concurrency (INF-300). Push-triggered runs share one per-PR concurrency group (
…-push) and still cancel in progress, so a rapid double-push cancels the older run exactly as before. A manual re-run gets its own run-scoped group, so it can neither cancel nor displace an in-flight or queued review of a newer head — which would otherwise leave the current head unreviewed until the next push, since the stale re-run then trips the fail-closed stale-head write guard. Gatingcancel-in-progresson the attempt number is not enough for this: GitHub cancels an existing pending group member whenever another job enters the group, regardless of that flag, which only governs the running member. Because a re-run consequently is not serialised against push runs, a re-run is also made read-only for review state: an ordinary re-run may review and post its findings, while an INF-431 edited-event re-run only refreshes the verdict; neither writes the state comment. Otherwise two same-head runs could both pass the live-head guard and PATCH the same comment, the last writer silently dropping the other's findings from revalidation. State advancement belongs to push-triggered runs; the cost is one full review on the next push, which is the safe direction. - Prompt trust boundary. The fixed confirmation rules ride in the system prompt; the prior findings themselves (earlier model output influenced by the contributor's diff) are passed in the user message explicitly framed as untrusted data, so instruction-like text inside a recorded finding cannot steer the next review at system priority.
- The initial pass also gained a coverage contract (INF-292) in
SYSTEM_REVIEWER_RULES, unconditional on the flag: complete the review before reporting, report every independently actionable material finding (dedupe rather than truncate at an arbitrary count) — this front-loads the depth a later delta-scoped confirmation review depends on. - The INF-172 metrics record gains a
review_modefield (full|confirmation), recorded on every run regardless of the flag, so convergence can be measured before the flag is enabled anywhere it matters.
Endpoint + deadline knobs (INF-249 — GPT-5.6 readiness)
OpenAI's GPT-5.6 guidance requires reasoning + function tools to go through the Responses API (Chat Completions function tools need effective reasoning none, while 5.6 defaults to medium), so the agentic reviewer routes endpoints per model:
MULTI_LLM_REVIEW_RESPONSES_MODELS— CSV of model-id prefixes whose agentic calls use the gateway's/v1/responsesendpoint (e.g.openai/gpt-5.6covers-sol,-terra,-luna). Unset ⇒ every model stays on Chat Completions — the GPT-5.5 path is byte-for-byte unchanged. A prefix that matches no rostered model is inert rather than an error, so this variable can be left configured across a roster change: with a non-OpenAI roster it simply routes nothing, and restoring a GPT-5.6 model re-activates it without a second edit.MULTI_LLM_REVIEW_REASONING_EFFORT— explicit reasoning effort sent on the Responses path (none | low | medium | high | xhigh | max, the GPT-5.6 set — there is nominimalon 5.6). Defaultmedium; an invalid value fails loudly rather than silently running a different configuration.MULTI_LLM_REVIEW_RUN_DEADLINE_MS— the soft in-run deadline described above. Workflow default600000(10 min); unset in local/eval runs ⇒ unbounded.VERCEL_AI_GATEWAY_RESPONSES_URL— optional explicit Responses endpoint. Unset ⇒ derived fromVERCEL_AI_GATEWAY_URLwhen that is a.../chat/completionsURL; an underivable custom URL fails closed (never the public endpoint — that would route review content and the bearer token past a private gateway).
The per-run INF-172 metrics record embeds the effective runtime config (models, agentic/panel flags, responses routing, effort, budgets, deadline), so a metrics row stays interpretable without reconstructing repo-var state after the fact. Offline eval runs (scripts/review-eval/replay.ts) additionally persist a full reproducibility record — commit, per-model endpoint, requested/returned model ids, effort, grounding/KB modes, input hashes, budget knobs, duration — and keep the benchmark stationary by default (REVIEW_EVAL_KB=off; set on explicitly for the production-parity arm).
Cost scales ~linearly with the model count (per-PR spend ≈ the per-review ceiling × number of models). The posted review's footer breaks down per-model token cost and notes any model that was skipped (a partial panel is labelled as such, never reported as a clean pass). Measurement to date (INF-171): the 3-model union lifts recall but regresses precision — see the INF-163 measurement log before enabling in CI.
How to skip a PR
Add [skip ai-review] to the PR title — any substring match counts here, as with every other CI bypass marker in the repo (per bypass-marker matching) — or on its own line in the PR body. Inline backticked or prose mentions in the body do not trigger the skip; that strict-line rule only applies to the body, not the title. The title match is a plain case-insensitive substring — backticks do not prevent it — so to mention the concept in a PR title without triggering the skip, write it differently (e.g. skip-ai-review).
Adding or removing the marker re-triggers the check automatically (the workflow subscribes to pull_request.edited, and exemptions are always evaluated against the live PR metadata, not the triggering event's payload) — so a marker added after a red run turns the check green without a new push. With delta mode off, an ordinary title/body edit on a current-head reviewed PR uses the INF-297 cheap exit. With delta mode on, the script classifies the edit: verdict-only own-lines reuse state and refresh the verdict without an LLM call; review-relevant edits force a full review.
The workflow also deliberately skips — green, with the reason recorded in the check's step summary (INF-297):
- Draft PRs (job-level; GitHub re-triggers on ready-for-review).
- PRs from forks (the secret is not available to fork PRs, by design).
- PRs opened by
dependabot[bot]. - PRs whose base branch is neither
developnor anepic/*integration branch, andrelease/*/hotfix/*head branches — the same exemption set as the other source-PR gates. - PRs with zero changed files (e.g. an already-merged back-merge).
How the budget works
The token budget is the cumulative ceiling per PR push across all configured models. The script:
- Estimates the prompt size (≈4 chars/token) before each gateway call.
- Skips remaining models cleanly if the estimate would overflow the budget — a "Skipped: token budget exhausted" review is posted so reviewers see the abort, rather than silent omission.
- Records the gateway's reported
usage.total_tokensfrom the terminal SSE usage chunk after each successful call. Falls back toestimateTokenswhen the gateway omits the usage chunk (some providers do not supportstream_options). - Logs the running balance to the workflow's step output.
The default 100k/PR comfortably accommodates the constitution (~10k) + a typical spec (~5k) + a 30k diff + the response. Bigger refactor PRs may hit the budget; raise MULTI_LLM_REVIEW_TOKEN_BUDGET per-PR (gh variable set MULTI_LLM_REVIEW_TOKEN_BUDGET --body 250000) or [skip ai-review] them.
Failure modes (fail-loud since INF-297)
The review content is advisory — the bot never approves or requests changes; humans own merge. Its absence is not advisory: since INF-297 (motivated by the INF-172 shadow window, where 8 PRs — two of them security-relevant directory changes — shipped unreviewed behind green checks during two gateway-credit outages), every run ends with a determinate terminal status written to a status file and rendered into the check's step summary:
reviewed— at least one model's review actually posted (findings, or an explicit clean "no findings" review). This is the only outcome that represents fresh model coverage; "the workflow exited 0" is never, by itself, treated as reviewed.verdict_only— an empty confirmation delta reused trusted review state and refreshed the independent-review verdict without a model call. Green and non-exempt, but explicitly not fresh review coverage.skipped— a deliberate exemption (see the list above), green with the reason attributable in the step summary.failed— everything else. The script exits non-zero and the required check goes RED.
| Mode | Behaviour |
|---|---|
MULTI_LLM_REVIEW_ENABLED != true | RED — a review was expected and the reviewer is disabled (accidental skip). |
VERCEL_AI_GATEWAY_TOKEN missing on a model-review path | RED — the reviewer cannot call the gateway. A no-model verdict_only refresh does not require it. |
| Gateway error (5xx, 402 credit exhaustion, idle timeout) | A "Skipped: gateway error" review is posted for visibility, but no model reviewed → terminal status failed, RED. |
| Model returns malformed / unparseable findings JSON | Treated as an incomplete review (both agentic and diff-only paths), never as "zero findings — clean"; nothing else posted → RED. |
| Budget exhausted with nothing posted | "Skipped: budget exhausted" review posted for visibility; no model reviewed → RED. |
| Findings collected but the review POST fails | The review never reached the PR → RED. |
| Script crashes / job dies without writing a status file | The always-run terminal step fails closed: missing or unparseable status ⇒ RED. |
| Full PR diff is empty (zero changed files) | Deliberate skip → green, explicit. |
| Confirmation delta is empty | verdict_only → green; trusted findings carry forward and the verdict gate runs, but no fresh review coverage is claimed. |
| GitHub check-runs API returns 403/error (neutral check) | Warning logged; never throws — the neutral findings check-run is a bonus signal, not the terminal status. |
| Partial panel (some models skipped, ≥1 posted) | reviewed, with the partiality named in the posted review and the terminal detail. |
The reason a review did or did not post is always visible on the PR: open the Multi-LLM review check → the step summary shows reviewed / verdict refreshed / deliberately skipped (reason) / failed (reason) without reading raw logs. To deliberately skip a red PR, add [skip ai-review] (title, or own body line) — the edit re-triggers the check and re-runs read live PR metadata.
Soak protocol (ongoing quality signal)
The reviewer has been live since 2026-06-02. The soak protocol (originally the activation gate) continues as an ongoing quality signal for deciding when to expand the panel:
The protocol as originally written (kept for context — see the note below on why it produced nothing):
- Watch the bot's inline review comments on merged PRs.
- React with 👍 on findings that genuinely helped.
- React with 👎 on findings that were noise (uncited, wrong, or pedantic style preferences).
- Decision gate for v2 (adding a second model):
- If 👍 outweighs 👎 across 5+ PRs → schedule v2 (add
google/gemini-2.0-proas a second reviewer — "panel of judges"). - If 👎 outweighs 👍 → tighten the system prompt's "what counts as a finding" rules and re-soak.
- If 👍 outweighs 👎 across 5+ PRs → schedule v2 (add
⚠ In practice this protocol produced no data. Measured 2026-07-16 (INF-254): zero 👍/👎 reactions exist across the whole shadow window — 89 findings on 5 sampled PRs, 0 reactions, and none anywhere in the 14-day window. The team resolves threads and replies fix-or-refute instead; it does not react with emoji. Treat the thumbs rate as unavailable, not as neutral, and use the two signals below.
The INF-163 offline eval harness (scripts/review-eval/) gives the objective signal: precision and recall against a labelled benchmark of real historic Constellation defects. The shadow collector (below) gives the live one, derived from review-thread state rather than reactions.
Shadow window + adjudication protocol (INF-172)
Phase 4 of INF-163 turns the soak signal into a MEASURED comparison against Copilot over a real shadow window (>= 20 PRs AND >= 2 weeks), feeding the INF-173 retire-Copilot decision. Both reviewers already run on every PR — this layer only measures.
Per-run metrics sink. When MULTI_LLM_REVIEW_METRICS_PATH is set (the workflow sets it), the reviewer appends one JSONL record per run — {run_id, pr, reviewer, models, finding_count, inline_finding_count, tokens_used, timestamp} — and the workflow uploads it as a review-metrics-<run_id>-<run_attempt> Actions artifact (90-day retention; the attempt suffix keeps re-runs of failed jobs from colliding with the immutable artifact of the previous attempt). finding_count is the total postable findings; inline_finding_count is the anchored subset actually posted as diff comments — the only ones the collector can harvest (non-anchorable findings go to the review summary body, which the collector never reads). Append-only JSONL by design — no DB table (INF-scope ownership decision on the INF-172 ticket).
Adjudication protocol (what humans do during the window). Nothing extra — just review PRs the way you already do (INF-254).
The window originally asked reviewers to react 👍/👎 on inline findings. Nobody ever did: adjudication sat at 0% for the window's whole first half (0 reactions across 89 findings on 5 sampled PRs). The signal the team actually emits is the one AGENTS.md § PR Review Loop already mandates — resolve-thread / fix-or-refute — so the collector now reads that instead.
- Review PRs normally. Both reviewers' findings arrive as inline diff comments.
- Fix the finding, or reply with a brief reason — then resolve the thread. That is the existing convention; there is nothing new to remember.
- Reactions still work and still win when present: 👍 → true positive, 👎 → false positive. They are an override, not the mechanism.
- Don't agonise: an unresolved thread, or a reply that asserts no clear verdict, is reported as unadjudicated and is never imputed either way. Contradictory signals are reported as conflicted and excluded from precision.
- Don't change the reviewer config mid-window (models, panel, budgets, prompts). The pinned config and window start date are recorded on the wiki page
inf-163-first-measurement-2026-06-26. If it changes anyway, the collector partitions per model rather than blending — see below.
How the verdict is derived. Resolution alone cannot tell TP from FP, because the convention resolves threads in both the fix and the refute case (measured: 241 of 243 window threads are resolved — a near-constant). So resolution is a gate, and the human reply is the discriminator: its leading clause is matched against a conservative fix/refute vocabulary (Fixed in <sha>: … → TP, Refuted (no change): … → FP). Anything that does not clearly assert a verdict stays unadjudicated rather than being guessed. The full rule, its vocabulary and its known failure modes live in .ai/specs/SPEC-inf-254-shadow-adjudication.md.
⚠ What adjudicated precision measures. It is the accept-and-fix rate, not correctness, and it is not symmetric between reviewers: a cheap, obviously-correct nit gets fixed unargued and scores TP, while a deep, contestable finding invites scrutiny and is often refuted — and a refuted deep finding may still have been worth raising. The metric rewards triviality. Read it alongside the offline benchmark (which scores against ground-truth defect labels), never as a standalone verdict.
Collector. npx tsx scripts/review-eval/shadow-collect.ts --prs <csv> (or --since <date> [--until <date>], optionally --model <id>) harvests both reviewers' inline findings for each PR in the window (ours via the Multi-LLM PR review review marker, Copilot via the Copilot login), derives per-reviewer adjudicated precision from review-thread state, and computes agreement (same file + overlapping lines between the two finding sets). It writes per-PR JSONL plus an aggregate markdown report to scripts/review-eval/results/shadow/ (gitignored).
- Per-model partitioning. Every agent finding is attributed to the model named by its review's
**Model:**line, so a roster change mid-window partitions the measurement instead of blending it.--model <id>narrows a run to one rostered model (precision, agreement, per-PR rows and the JSONL all honour it). Pass no--modelfor the unfiltered superset; a blank--modelis rejected. - Only real reviews count. A PR is measured only if the model under measurement actually reviewed it. A run that was budget-skipped or errored still posts a marked review (
> Skipped: …) but no findings, so it does not count as a review — this is not hypothetical: on #1304 all 12 marked reviews were gateway-error skips, so gpt-5.5 never reviewed it at all, and ~8% of the window's apparent gpt-5.5 coverage was this kind of phantom. Anything the measured model did not review is reported as agent-absent and excluded, so the agent is never scored as "found nothing" on work it never saw.
What shadow data can NEVER claim: live recall. There are no ground-truth defect labels in a shadow window, so the collector's aggregate has no recall field and its report says so explicitly. Recall remains the offline benchmark's job (scripts/review-eval/replay.ts against scripts/review-eval/benchmark/).
⚠ Cross-reviewer agreement is low, and that is a real result, not a bug. Measured over the window: 5.0% of the agent's locatable findings overlap a Copilot one (24/482), and 6.5% the other way (20/308) — low, but not zero. The reviewers are complementary, not redundant: the agent finds RLS / outbox / cache-scoping defects; Copilot finds spec surfaces:-list gaps and docs drift (on #1345 both had 6 locatable findings and 0/6 overlap each way). A ~5% overlap is far too thin to corroborate findings, so agreement cannot serve as a precision filter.
v2 sketch — panel of judges (deferred)
The YAML scaffold and the script both already loop over --models. v2 adds a second model to the list and lets the workflow post two sets of inline reviews per PR. Open questions for when v2 is up:
- Do we want a third "summariser" pass that reconciles findings across the two reviewers, or is two raw review sets fine?
- Do we want to surface a combined neutral check verdict that aggregates all models' finding counts?
Defer until v1 has soaked successfully.
Where it lives
| File | What |
|---|---|
.github/workflows/multi-llm-review.yml | The workflow. Owns gating, environment, escape hatches. |
scripts/multi-llm-review.ts | The reviewer script. Pure I/O at the edges, pure functions for unit-tested logic. |
scripts/multi-llm-review.test.ts | Unit tests: arg parsing, spec extraction, budget tracker, prompt assembly, rendering, streaming, inline posting. |
.ai/specs/SPEC-inf-126-multi-llm-review-activation.md | The activation spec. |
.ai/specs/SPEC-inf-155-multillm-stream-inline.md | The streaming + inline comments spec. |
scripts/review-eval/ | Offline eval harness (INF-168). Scores any reviewer arm against a labelled benchmark; reports precision + recall. |
scripts/review-eval/benchmark/seed.json | 3 seed labelled cases from real historic Constellation defects. |
.ai/specs/SPEC-inf-168-review-eval-harness.md | Spec for the offline eval harness (Phase 0 of INF-163). |
scripts/review-eval/shadow-collect.ts | Shadow-window collector (INF-172): adjudicated precision + cross-reviewer agreement from live PRs. |
.ai/specs/SPEC-inf-172-shadow-mode.md | Spec for shadow mode + the per-run metrics sink (Phase 4 of INF-163). |
.ai/specs/SPEC-inf-293-delta-scoped-multi-llm-rereview.md | Spec for confirmation mode: the state comment, mode resolution, and delta-scoped re-review. |