Dual-Rival: Why Single-Reviewer AI Reviews Miss Half the Defects
I ran two AI reviewers on the same code review. Zero overlap in findings, five of six rounds. Here's the asymmetric-coverage property—and how to run Dual-Rival correctly.

I ran two AI reviewers on the same set of proposed amendments. Concurrently. Independently. Codex came back with four findings. Claude came back with nine.
I expected overlap. Maybe a couple of duplicate flags. Maybe a third of the list shared.
I opened both outputs side by side and read them line by line, looking for anywhere the two reviewers had caught the same thing.
Zero.
Codex's four findings were all citation and line-mapping errors. The evidence pointed to transcript line 2248 when the actual content sat at lines 230-238. An amendment attributed a quote to session S386 when the transcript confirmed S387. Inventory item numbers off by one. Mechanical work. Does the quote actually live where you said it lives?
Claude's nine findings were all architectural. The premise-check gate required self-detection by the session that was already frame-accepted. Two clusters shared an under-correcting costs more principle, raising the question of intentional reinforcement vs redundancy. Distinguishability claims of 20-30% overlap without a shown test methodology. Structural probing. Does the amendment discharge its cluster? Do the numbered sequences make logical sense?
Same amendment set. Same review pass. Different defect classes. Disjoint coverage.
This isn't the first time I've seen it. The pattern has validated across six independent rounds in 2026: S327, S350, S387, S388 R3, S388 R6, S390 Phase 3 R2.
Six runs, six different work types. Zero overlap five times, one overlap once.
The implication is uncomfortable if you've been running single-reviewer AI code review: you're systematically missing half the defect class. Not randomly half. Specifically the half your reviewer's cognitive bias doesn't probe. Codex misses architecture. Claude misses literal citation drift. Neither catches both unless you run both.
I'll show you how I know, then what the fix looks like. The adapter pattern covered non-monotonic model evolution. Directive vs narrative covered rule form. The recursive governance vulnerability covered governance recursion. This piece is about review structure.
What each AI reviewer actually does well
Imagine an X-ray on a light box. Two doctors walk up to read it.
One reads structure. Heart silhouette. Lung fields. The shape of the mediastinum. She's looking at the picture.
The other reads counts and measurements. Rib spacings. Foreign object positions. Cardiothoracic ratio. She's looking at the literal numbers.
Same X-ray. Two readings. Different findings.
The structural reader can swear nothing's misaligned while the measurement reader catches a swallowed coin two ribs in. The measurement reader can sign off on every literal count while the structural reader notices the mediastinum is shifted three centimeters off-center. Each is strong where the other is weak. That's what Codex and Claude do on adversarial review.
Same dynamic plays out on the mat. Two coaches watching the same roll see different rolls. One coach watches grip-fighting micro-mechanics -- where the hands are, who has the dominant grip, whether the framing is stable. The other watches positional dominance -- who has the angle, who is loading whose hips, where the scramble is going if it goes. The grip coach calls a sweep three seconds before it lands because he saw the grip-break that set it up. The position coach explains, after the fact, why the sweep was inevitable from where the player had taken the angle. Same roll. Two readings. Different findings.
Codex reviews like the measurement reader. Does the quote appear on the line you cited? Does the reference number check out? Does the session ID match the transcript? Mechanical. Literal. Doesn't try to read the architectural picture.
Claude reviews like the structure reader. Does the amendment discharge the pattern it claims to address? Does the mitigation contradict its own evidence? Does the framework hold together? Doesn't always check exact line numbers.
Whichever reviewer you've chosen is systematically missing the defect class the other would catch.
Six rounds of Dual-Rival: one validated property
One round of zero-overlap is anecdotal. Two is suggestive. Six is a property. The honest version: it's a property in my stack across six runs. Run it on yours. If you get overlap rates above 10-15% across three independent runs, your two reviewers aren't asymmetric enough. Pick a different pair.
S327 was the first time I noticed it. Codex found five BLOCKERs on a batch of skill edits, all empirical corrections with exact line references. Claude found different architecture issues. My notes read "cross-model value confirmed again."
S350 was my structural-gates work on PROTOCOLS.md. Codex found cross-file rule-text inconsistencies; Claude found conceptual gaps in how the gates would fire during real workflows. No overlap.
S387 was my adapter-fix amendment work. Four citation errors vs. nine architectural issues. Zero overlap. This is the run I led the article with.
S388 R3: I ran a Rival on a refinement plan, mid-arc. Codex flagged a BLOCKER on a rule-body mechanical inconsistency (a tri-split in detection signals). Claude-side flagged new WATCHes on detection-signal observability and mitigation-template scope. Different findings.
S388 R6: my final re-verify on the 100% STRICT state. Codex 9.4/10 YES with one residual WATCH on verbatim quote traceability. Claude-side 9.7/10 YES with no new findings. The WATCH Codex caught was architecturally invisible to Claude-side, and that gap turned into a methodology refinement I'll get to.
S390 Phase 3 R2 was the most interesting one. An 11-item batched amendment to the playbook and adapter family. Round 1 surfaced 12 distinct findings between the two reviewers, with exactly one overlap: both flagged a section-numbering contradiction between BCP §2.2 and §2.4. That one overlap is worth naming. It's the defect class where the two reviewers' specialties intersect. Numbering consistency is mechanical (Codex) and architectural (Claude) at the same time. Overlap fires when a defect straddles both modes.
Zero-overlap is the default. Overlap is diagnostic.
At this point the asymmetric coverage isn't a hypothesis. It's a validated methodology property. I have guesses about mechanism: different training paradigms, different alignment objectives, different tokenizations. The pattern stands regardless of which guess is correct.
Why zero-overlap is the feature
Here's the part that took me a while to internalize. I kept reading zero-overlap as a problem to solve. Two reviewers should agree more, I thought.
Wrong frame. The disagreement IS the coverage.
If Reviewer A and Reviewer B overlap 80%, you're paying 2× cost for 1.2× coverage. Bad trade.
If they overlap 0%, you're paying 2× cost for 2× coverage. Every finding is unique.
This is also why running the same reviewer twice doesn't help. Same model, same input, mostly same findings. The structural gain from two reviewers comes from their cognitive asymmetry. The less your reviewers resemble each other, the better.
Run them independently. Compose at synthesis.
For a while I thought the order mattered. I dispatch the Codex call first, so I assumed Codex-then-Claude was a sequence — that the second reviewer was reading the first's findings and building on them.
It isn't a sequence. It's a coincidence of dispatch.
Both reviewers run concurrently. Neither reads the other's output. Codex probes citations while Claude-side probes architecture, in the same window, with no channel between them. The "Codex first" I kept writing in my notes meant only that Codex's call left the gate first. It never meant Claude waited on Codex's answer.
That distinction is the whole methodology.
The independence is what protects the asymmetry. If Claude-side read Codex's findings before forming its own, you'd contaminate exactly the asymmetry you ran two reviewers to get. Claude would anchor on what Codex already flagged — pruning its architectural pass to avoid the citation ground Codex covered. You'd converge the two readings toward each other. You'd pay 2x cost and watch the coverage degrade back toward 1x. Reviewer independence isn't a nice-to-have. It's what keeps the asymmetry intact — and the asymmetry is the coverage.
So nothing composes at the reviewer layer. The reviewers don't talk. They can't — talking is the failure mode.
Composition happens one layer up. After both reviewers return their findings, the team classifies across both channels — and that's where an architectural issue gets read against a citation issue, where a defect that straddles both modes gets named as the overlap it is. The S390 section-numbering finding wasn't caught because one reviewer read the other. It was caught because two independent reviewers both surfaced it, and synthesis saw the collision.
Get the layers right:
- Reviewers: concurrent, independent. No reads across the channel.
- Synthesis: sequential, integrative. This is where the two findings sets compose.
I had this backwards in my own notes for a stretch of 2026 — I wrote "order matters" when what mattered was that the orders never touched. S327 was the first run. I've run them independent every time since. The dispatch order is bookkeeping. The independence is the product.
When single-Rival is the call (or the fallback)
I don't run Dual-Rival on everything. My CLAUDE.md risk classification is the decision rule:
- COMPETITION (irreversible: production deploys, governance edits, external sends): I require Dual-Rival.
- SPARRING (security, auth, financial, credential-adjacent): single-Rival acceptable; Dual-Rival when stakes warrant.
- DRILLING / ROUTINE: single-Rival usually sufficient.
The tradeoff pivots at COMPETITION. Below that line, single-Rival's gap is acceptable. At and above, missing half the defect class is an operational risk.
Sometimes a reviewer just isn't available. My S387 work hit this when Codex MCP disconnected mid-session. The correct response isn't to pretend single-Rival covers what Dual-Rival covers. Run the available reviewer with discipline, document the coverage gap explicitly, and either defer or proceed with the gap on record. The mistake would be silent acceptance.
The team isn't a fourth reviewer
S390 Phase 3 was an 11-item batched amendment. I tried running everyone in parallel: Codex plus three Claude-side domain agents (EA, software-dev, research) all reviewing as independent reviewers. I had to back up twice during that session before the structure landed.
The actual flow is three roles, not four reviewers:
- Reviewers produce findings. Codex and Claude-side Rival. Two reviewers, concurrent and independent.
- Team classifies findings. EA, software-dev, research. Per the S181 Finding Analysis Protocol, they classify Codex's and Rival's findings as ALREADY ADDRESSED / RESTATEMENT / PARTIAL / NEW. They're the Classification Gate, not a second review pass.
- Session synthesizes TEAM RECOMMENDATION. I read the classifications and present a synthesized recommendation. Not raw findings for the human to adjudicate.
Running all four in parallel collapses roles 1 and 2. The team agents end up producing their own findings instead of classifying the reviewers'.
A four-reviewer pattern IS correct, but for a different activity. A panel reviewing the same artifact in parallel is the right shape for ratification of a consolidated position. Dual-Rival is REVIEW with three roles. Four-reviewer panels are RATIFICATION with four readings. Different work.
Argument vs artifact: the Round 6 refinement
S388 Round 6 showed me a property I hadn't named before. Two Rival reviewers were reviewing the same audit artifacts. Claude-side 9.7/10 YES, closed cleanly. Codex 9.4/10 YES but flagged one WATCH: the audit defensibility would improve with direct quote snippets for the two re-maps.
Claude-side had reviewed the audit notes for argument soundness. Did the re-mapping rationale make sense? Yes on every check.
Codex had reviewed the same notes for artifact self-containment. Did the artifact contain enough evidence to verify the argument independently of outside context? No. The audit notes claimed a verbatim trigger match but didn't surface the quoted trigger text, leaving a future auditor to trust the claim instead of verify it.
Two distinct checks with distinct blind spots:
- Argument soundness: given full institutional memory, does the reasoning hold?
- Artifact self-containment: without institutional memory, can a future auditor verify the reasoning from the artifact alone?
Claude-side leans architectural and is strong at soundness. Weaker at self-containment because the is the argument valid check and the does the artifact carry evidence to support it check feel like the same check but aren't.
Codex leans literal and is strong at self-containment. Weaker at soundness because it doesn't probe whether the rule attribution is architecturally correct, just whether the quoted evidence exists.
Practical takeaway: run both checks explicitly. Not just does the Rival think this is a good argument. Also could a six-months-from-now auditor verify this from the artifact alone. I now run both every time. Round 6 was the moment the distinction got named.
Zero new findings at Round 2 is a different kind of pass
One more refinement. The Round 2 verdict pattern carries more information than threshold-met-or-not. I used to read R2 as a single signal: did the score clear threshold? Yes or no. Done.
Three outcomes scale differently:
- R2 threshold met AND new findings introduced. Pass, but fixes cascaded. F-META recurrence risk is live. R3 warranted regardless of score.
- R2 threshold met AND one or two minor cosmetic findings. Pass. Non-blocking.
- R2 threshold met AND zero new findings. Pass plus confirmation. Fixes were clean at the architectural level. Strongest positive signal a Round 2 can carry.
S390 Phase 3 R2 hit zero-new-findings on Codex (9.4/10 YES, 0 new) and near-zero on Claude-side (9.2/10 YES, 1 resolved WATCH + 2 cosmetic). Single-round scoring can't carry that signal. Iterative Dual-Rival can.
Track NEW FINDINGS count across rounds. Zero NEW at R2 is strong. Three-plus NEW with threshold cleared warrants R3 for F-META recurrence check.
Dual-Rival vs Evaluator-Optimizer
These get conflated. Anthropic's Evaluator-Optimizer pattern pairs a generator with a reviewer that iterates on feedback. Dual-Rival uses two disjoint reviewers on a finished artifact. Same direction (adversarial review as discipline). Different mechanism. Different failure mode addressed.
The two compose. My stack uses Evaluator-Optimizer-style iteration inside SASSY (my content production methodology) and Dual-Rival as the final pass before COMPETITION-tier content or governance ships.
Cost and coverage
I track the cost of every Rival cycle I run. Dual-Rival is roughly 2× the cost of single-Rival. Two invocations, two token budgets, two cycles.
For COMPETITION-tier work, the math works: 2× cost, 2× coverage. Clean ROI.
For DRILLING and ROUTINE, the math doesn't: stakes are recoverable, single-Rival's coverage is acceptable, the second invocation doesn't earn its keep.
Don't Dual-Rival everything. Don't single-Rival COMPETITION-tier. The tier distinction is the decision rule.
Closing
If your AI review stack only uses one model, you aren't getting the coverage you think you're getting. I wasn't, for a stretch of 2026 before the second reviewer earned its slot.
The fix is structural, not algorithmic. Pick two reviewers from different training paradigms. Run them concurrently — neither reading the other. Compose their findings at team synthesis, not at the reviewer layer. Categorize your review flows by tier. Dual-Rival the COMPETITION ones.
Six independent rounds have validated the asymmetric-coverage property across 2026. It isn't hypothesis. It's methodology. If you're maintaining governance or code review at scale, this is the pattern that catches both halves of the defect space.
Try both. Watch the zero-overlap. That's what I'd do.
Piece 4 of 6 in the Opus 4.7 series. Shape vs fidelity covers the sovereign vocabulary that makes cross-model review infrastructure possible.