A small, dependency-free proof of concept that reduces noisy Codex pull-request findings before they reach a human reviewer. It can also apply one conservative batch of verified fixes and perform one final review.
The associated research explains the evidence, alternatives, and design constraints in more detail.
Run the checkout directly:
./review-reducer review --repo ~/code/my-project --base origin/mainReview a GitHub pull request directly instead of locating its branch yourself:
./review-reducer review --pr openai/codex#123
./review-reducer review --repo openai/codex --pr 123
./review-reducer review --repo ~/code/codex --pr 123
./review-reducer review --pr https://github.com/openai/codex/pull/123PR mode uses your authenticated gh CLI to pin GitHub's exact head and base
commits, including forked and stacked pull requests. It reuses an existing clean
worktree at the exact PR head, or creates a durable detached sibling such as
~/code/codex.review-pr-123-abc123def456. Your existing checkout and current
branch are never switched. The repository must already exist locally; pass its
path with --repo when it is not under ~/code. PR identity, title, branches,
and pinned commits are preserved in the saved session and HTML report.
For a stacked pull request, set --base to its immediate parent branch instead
of origin/main in local-checkout mode; otherwise inherited parent changes can
produce false findings. GitHub PR mode selects the actual PR base automatically.
Every P0–P3 finding gets the same evidence-grounded review. Priority is retained
as useful context, but never excludes an issue from investigation. The command
exits successfully when no consequential verified issues survive, exits 2 when
verified issues remain, and exits 3 when a claim requires human judgment.
Interactive terminals show a live dashboard with the current review stage, parallel Codex agents, recent source-inspection activity, finding decisions, elapsed time, and token usage. Redirected output automatically falls back to plain progress lines. Force or disable the dashboard when necessary:
╭──────────────────────────────────────────────────────────────────────────╮
│ ◈ C O D E X R E V I E W R E D U C E R ⠹ LIVE 00:18 │
│ PIPELINE │
│ ✓ review ⠹ challenge ○ repair ○ final │
│ │
│ ACTIVE AGENTS │
│ ⠹ adversarial review · src/app.py:42 │
│ Comparing the changed caller contract against the exact merge base. │
│ │
│ FINDINGS 3 found 1 accepted 1 rejected 0 human │
│ ✓ P1 Existing guard already rejects invalid input [REJECTED] │
│ ⠹ P1 Preserve caller contract [CHALLENGING] │
╰──────────────────────────────────────────────────────────────────────────╯
./review-reducer review --repo ~/code/my-project --progress always
./review-reducer review --repo ~/code/my-project --progress neverThe dashboard writes only to stderr, so --json remains valid JSON on stdout.
Set NO_COLOR=1 to preserve the dashboard without ANSI colors.
Every completed review writes a polished, self-contained HTML report beside its saved session:
<git-common-dir>/review-reducer/<session-id>/report.html
Interactive terminal runs automatically open the report in your browser after the review finishes. Disable that behavior without skipping report generation:
./review-reducer review \
--repo ~/code/my-project \
--base origin/main \
--no-open-reportUse --open-report to force opening when output is redirected. CI and other
noninteractive runs otherwise leave the browser closed. The report is a single
offline HTML file with inline styling, expandable evidence, local source links,
plain-language issue explanations, user impact, recommended actions, repair
budgets, model usage, and exact manual-curation commands. It makes no network
requests, executes no scripts, and requires no additional Codex turns.
Generate or reopen a report for an existing session without repeating its review:
./review-reducer session report latest --repo ~/code/my-project
./review-reducer session report latest \
--repo ~/code/my-project \
--output ~/Desktop/review-report.html \
--no-open-reportManual session include, session dismiss, and session reset commands keep
the session's default HTML report synchronized with your latest decisions.
Ask Codex a focused question about an existing finding without rerunning the full review:
./review-reducer session ask latest 1 \
'Can this be fixed in fewer than 10 production lines?' \
--repo ~/code/my-projectEach question uses one fresh, read-only Codex turn with the original finding, blind investigation, adversarial assessment, any reviewer rebuttal, and recent questions about that same finding. The repository, base, commit, and exact tracked patch must still match the saved session. Answers include verified source references, confidence, remaining uncertainty, a recommended action, and any smaller proposed fix.
Choose a viewpoint when the question calls for it:
./review-reducer session ask latest 3 \
'Could this concern be safely dismissed?' \
--perspective adversary \
--repo ~/code/my-project
./review-reducer session ask latest 3 \
'What concrete user-visible failure supports this finding?' \
--perspective reviewer \
--repo ~/code/my-projectThe default neutral perspective answers without defending either side.
Questions and answers are appended to the saved finding and immediately appear
in its standalone HTML report. A suggested verdict is advisory: it never
changes the review decision unless you explicitly run session include or
session dismiss. Use --model, --reasoning-effort, --timeout, --json,
--progress, and --no-open-report to control the focused turn and its output.
Every run immediately creates a durable local session. List the sessions for a repository, inspect the latest results, or drill into one finding by its 1-based number or identifier prefix:
./review-reducer session list --repo ~/code/my-project
./review-reducer session show latest --repo ~/code/my-project
./review-reducer session show latest --finding 1 --repo ~/code/my-projectA finding's detailed view includes the original native-review comment, the
independent blind investigation, the adversary's assessment, the original
reviewer's rebuttal when one was warranted, and the final model decision. Add
--json to any session command for the complete machine-readable evidence.
After inspecting the evidence, explicitly include, dismiss, or reset a finding without overwriting the model's original decision:
./review-reducer session dismiss latest 1 \
--repo ~/code/my-project \
--reason 'Behavior already exists on the exact merge base'
./review-reducer session include latest 0504f58a \
--repo ~/code/my-project \
--reason 'Confirmed user-visible failure is worth the direct fix'
./review-reducer session reset latest 1 --repo ~/code/my-projectApply the curated findings later as one bounded repair batch followed by one final native review:
./review-reducer session apply latest --repo ~/code/my-projectThe saved repository, HEAD, base, merge base, and complete patch must still match. Manually included findings remain ineligible for automatic repair unless the saved evidence identifies an intent-preserving direct fix that stays inside the same complexity and dependency budgets.
Older artifact directories remain inspectable even when they predate
session.json. To fully investigate findings that an older run skipped, reuse
its saved native review instead of spending another native-review turn:
./review-reducer review \
--repo ~/code/my-project \
--base origin/main \
--review-file /path/to/previous-run/initial.response.txtTo allow one small automatic repair and exactly one final native review:
./review-reducer review \
--repo ~/code/my-project \
--base origin/main \
--mode fix \
--max-added-production-lines 12 \
--max-additional-production-files 2The target must start with a clean working tree. The tool never commits, pushes, publishes a review, resolves threads, creates new working-tree files, or repeatedly repairs the same branch. Generated changes remain unstaged for inspection.
Optionally request an explicit check after the repair:
./review-reducer review \
--repo ~/code/my-project \
--base origin/main \
--mode fix \
--check 'cargo test -p my-crate focused_test_name'The reducer itself invokes a repository check only when --check explicitly
asks for it. Verification prompts prohibit running repository code, hooks,
builds, or tests. Codex's native built-in reviewer has its own fixed prompt and
may use tools permitted by its read-only sandbox. Explicit check commands
execute without shell interpretation.
The package can also be installed locally with python3 -m pip install -e ..
When running without installation in an environment that sets
PYTHONSAFEPATH=1, use the executable above or:
PYTHONPATH=. python3 -m review_reducer review --repo ~/code/my-project- Pin the target HEAD, exact merge base, complete tracked diff, and original production/test churn.
- Run Codex's built-in native reviewer in a fresh, ephemeral, read-only session.
- For every native finding, regardless of priority, run a separate blind, read-only Codex session that sees the source location but not the alleged defect.
- Run a fresh adversarial Codex session that receives the finding and blind observations. Require exact-base comparison, realistic reachability, concrete source anchors, independently evidenced user impact, preservation of the pull request's intent, and the smallest direct fix.
- When the adversary believes dismissal or downgrading is genuinely useful, give the original-review perspective one evidence-grounded chance to rebut or concede. Confirmed useful findings do not incur a manufactured debate.
- Deterministically reject source-refuted inherited, intentional, unreachable, speculative, or duplicate findings. Route unproven claims of any priority to human review instead of silently discarding them.
- In
fixmode only, send surviving findings to one workspace-write Codex session. Reject new files, dependency changes, public APIs, excess production lines, and excessive production-file churn. - Run only explicitly requested checks, followed by one final native review and the same verification policy. Stop there even if new findings remain.
Native review is intentionally kept separate from structured verification:
codex exec review --output-schema does not enforce the requested schema for
the built-in reviewer. The verifier, adjudicator, and fixer therefore use fresh
ordinary codex exec --output-schema sessions instead.
--mode review|fix Report only, or apply one verified fix batch.
--min-confidence 0.75 Minimum adversarial confidence for automatic fixes.
--max-added-production-lines 20 Hard limit checked against the actual repair.
--max-additional-production-files 2
--max-findings 12 Fail safely on unexpectedly large review outputs.
--jobs 2 Number of independent verification workers.
--review-model MODEL Override Codex's native review model.
--verifier-model MODEL Override the blind/adversarial model.
--fixer-model MODEL Override the repair model.
--review-file PATH Reuse an existing rendered or JSON-shaped review.
--artifacts-dir PATH Store artifacts outside the reviewed working tree.
--no-blind-verification Skip the independent blind source investigation.
--progress auto|always|never Control the live terminal dashboard.
--open-report / --no-open-report Force or disable opening the standalone HTML.
--check COMMAND Run only this explicitly authorized repair check.
--json Emit the complete machine-readable report.
Run artifacts are private to the local user and default to
<git-common-dir>/review-reducer/<timestamp>-<head>/. They include the pinned
snapshot, role prompts, JSONL events, final model responses, structured
decisions, measured repair churn, explicit check output, the final report, a
continuously updated session.json with per-finding evidence and manual
decision history, and the standalone visual report.html. Existing artifact
directories from older runs can still be
inspected by passing their path as the session selector; findings that were
never investigated must be reviewed again before they can be automatically
repaired.
The report also records actual per-role input, cached-input, output, and
reasoning-token usage. Structured roles disable apps, plugins, memories,
multi-agent spawning, and automatic skill-instruction injection to reduce
context size and prevent unrelated side effects.
- Repository content and prior model findings are treated as untrusted evidence.
- Review and verification are read-only; only the single repair role may edit.
- A dirty or untracked target is rejected before automatic repair begins.
- Untracked files created by a repair are rejected because native base review does not include them.
- Hard repair-budget failures preserve the generated working-tree changes for inspection; they do not attempt an unsafe automatic rollback.
- Static source inspection is not presented as runtime-observed evidence.
- Model-role isolation reduces anchoring but does not make same-model judgments statistically independent.
- Codex's built-in reviewer still processes applicable project instructions; hostile repository instruction files require external trust controls.
PYTHONPATH=. python3 -m unittest discover -s tests -v
./review-reducer --help