Review workflow

Reading a review

Check titles

Each admitted reviewer has a neutral GitHub progress check tied to its run and exact PR state. The reviewer name, version and executor identify whose result you are reading. Execution and publication are separate:

Wide tables: scroll horizontally to see more columns.

Status Meaning
Review queued Accepted, waiting to run.
Review in progress Execution started.
Review published Review publication is confirmed.
Review execution complete; … Execution finished; the rest of the title names publication status, such as pending, held or not confirmed.
Review unfinished A model attempt started without a complete result.
Review not started No model attempt launched.
Review superseded A newer head, base or PR state made the run no longer current.
Review cancelled Execution or authorization was withdrawn.
Review blocked A durable problem prevented check delivery.

Older runs may say Review complete. Read the separate publication facts before assuming a completed execution delivered a review. Historical GitHub comments and recorded notices keep their original wording; a newer web page does not rewrite them. GitHub's own “Completed in …” prefix is not proof that model work succeeded.

Notices: why a reviewer has not started

A neutral notice explains a refused or waiting reviewer without starting model work. Common causes are:

Wide tables: scroll horizontally to see more columns.

Title Cause and next step
Reviewer selection refused The named reviewer is unknown, disabled or unavailable. Request the configured enabled reviewers or use their exact identity.
Review refused This repository is not permitted for the selected subscription executor. The Codex login path also requires an explicitly listed private repository.
Review paused The maintainer paused new work; queued work can start after the pause lifts.
Review held API accounting is uncertain or the spending cap is exhausted. The maintainer must inspect the relevant budget records.
Review waiting Runtime, GitHub API access or a subscription usage window is temporarily unavailable. A known reset is shown when recorded.

A notice is not a credential-health check or a statement about every account's plan allowance. Limits and budget describes recovery.

The review's layout

The review identifies its reviewer, head, base, run and verified finding count. Sections appear when relevant:

  • Cost — recorded subscription observations and any separately metered API usage; observations are not invoices.
  • Additional verified findings — confirmed findings beyond the inline limit.
  • Unverified suspicions — candidates that no verifier confirmed; they do not count as findings.
  • Pre-existing observations — issues judged to predate the change.
  • Scope notes — what a stage could not read, check or execute.

Up to five findings are posted inline. Each gives its severity and title, claim, impact, evidence, verification and suggested change. An incomplete run can retain suspicions even when publication is held or unavailable; retained evidence does not establish successful GitHub delivery.

Severity (P0–P3)

Severity states how much a finding matters. The current prompt describes P1 as leading a builder to a wrong or impossible result, P2 as requiring a decision before building, and P3 as misleading wording. P0 is the higher severity tier in the finding schema. These are finding priorities, not model confidence scores. A fuller public legend remains tracked in #112.

Verified vs. unverified

The producer records candidates as suspicions. A fresh verifier can confirm, reject or reclassify them. Only confirmed candidates count as verified findings. Unconfirmed suspicions remain separately labeled, including when a run ends before verification. A verification label is agent-produced evidence, not a proof that the whole codebase is correct.

The aggregate verdict (Approve / Request changes)

Where an installation exposes the repository's verdict setting, Comments keeps individual reviewer findings without an aggregate verdict. Rules can add a separate deterministic Approve or Request changes review; selecting Rules is not itself an approval.

Request changes can follow verified introduced P0/P1 defects. Approve requires complete, valid evidence from the selected reviewers and required checks for the current commit, with the publication and accounting checks satisfied. Pending, incomplete, unavailable or withheld evidence cannot be treated as a clean run. Subscription execution must not be assumed to satisfy the offline proof needed for Approve. The two-subscription logins path is not included in the present Rules facts contract; do not assume the same aggregate support for every executor.

The source baseline contains the evidence and aggregate publication paths. Whether Rules is enabled in a specific installation depends on its supported schema, configured authority and accepted rollout. The UI setting does not enable those prerequisites. Changing the mode applies to newly admitted work and does not retroactively approve old reviews. Progress checks stay neutral, and Sigma never merges the PR.

Cost lines and the run page

API cost estimate records known and held amounts from Sigma's API ledger. Known means settled in that ledger; held means reserved and unsettled. These estimates are token-derived, not invoices. A continuation decision can accept a specific uncertain exposure without settling its charge or completing its run.

Observed subscription usage comes from recorded CLI usage and may be valued at API list price. It is not billed by Sigma, is not part of its API cap and does not describe the provider's subscription price. Producer and verifier observations are separate. An unavailable stage estimate means unknown, not zero; a recorded numeric observation must not be mistaken for a complete financial total.

The run page adds stage details, scope notes, resources and separate delivery status. It currently requires access as this Sigma's connected owner; PR authors can read the GitHub review and check without entering that workspace. Recorded Requested timestamp shows the recorded queue time in the browser's time zone, with UTC in the no-JavaScript fallback. Provider reset prose currently stays in UTC. Historical results preserve their recorded execution and usage rather than today's repository settings.