ACCESS
Technology

I Asked a Competing Model to Tear Apart My System

I built an AI-assisted operating system on Claude. Then I packaged 55 files and asked Codex to audit it cold. The diagnosis was sharp. The prescription was 6x what the patient needed. Here's what happened — and what it teaches about cross-model review.

George Samuelson
George Samuelson
·9 min read·Technology
I Asked a Competing Model to Tear Apart My System

I built a system on Claude — it's past 1,500 sessions now. When I ran this audit, 195 sessions in, I packaged the whole thing: 55 files of architecture, governance, routing tables, and operational standards. An AI-assisted operating system for everything I do — the things I run, the ones I'm part of, the ones I don't talk about.

Then I packaged all of it and handed it to Codex.

"Review this cold. Five dimensions. Tell me where it breaks."

Here is what happened.

Why Cross-Model Review Matters

The same model that builds your system shares your blind spots. That is not a design flaw — it is a structural reality. The model that helped create your routing tables cannot objectively assess your routing tables. The model that wrote your governance standards has internalized the assumptions baked into those standards.

This is the same principle I teach in jiu-jitsu: you cannot be your own training partner. You need someone who does not know your game, does not share your tendencies, and has no relationship with your ego.

In AI-assisted work, that means a different model. Different training data. Different tendencies. Different failure patterns. I have been using cross-model adversarial review — what I call the Rival — for over a year. But this was the biggest test yet: a full system audit, not a single artifact review.

The Setup: 55 Files, 5 Dimensions

I built an audit package. Not a casual "look at this file." A structured, comprehensive package:

  • 55 files covering the complete system: CLAUDE.md (the operating constitution), all agent definitions, all skills, governance standards, routing tables, session transcripts, the adversarial review playbook itself
  • 5 audit dimensions: Structural Integrity, Operational Consistency, Security Posture, Scalability, Self-Audit Integrity
  • Clear instructions: Review cold. No relationship. No history. Classify findings as CRITICAL, HIGH, MEDIUM, or LOW. End with a verdict: "Am I Safe? YES" or "Am I Safe? NO"

The package was designed to give Codex everything it needed to understand the system without any of the context that might soften its assessment.

Cold review. That is the point.

The Verdict: "Am I Safe? NO"

Codex came back with 24 findings.

SeverityCount
Critical0
HIGH11
MEDIUM12
LOW1

Verdict: "Am I Safe? NO."

No Critical findings. That matters. The system was not fundamentally broken. But 11 HIGH findings meant real issues existed — issues I had not caught through self-review.

Here is what Codex actually found that was worth finding.

What Codex Got Right

The diagnosis was good. Genuinely good. Four clusters stood out:

1. Routing contradictions. The legacy adversarial-reviewer-agent was still in routing tables even though the Rival had replaced it. Two paths to the same function. Not dangerous, but messy — the kind of drift that causes confusion over time.

2. Enforcement gaps. Some governance instructions said "MUST" and "MANDATORY" but had no structural enforcement mechanism. They relied on the model reading the instruction and deciding to comply. Codex flagged this correctly — instruction-only enforcement is the weakest class of control.

3. F-code count drift. The playbook referenced "10 F-codes" but the actual catalog had grown to 14. Documentation had not kept pace with the taxonomy.

4. Classification rules needed tightening. The anti-suppression rule for finding classification was implicit, not explicit. Codex correctly identified that classification should be additive only — you can escalate a finding's severity, never downgrade it.

These were real issues. Not theoretical concerns about hypothetical futures. Real drift, real gaps, real fixes needed.

What Codex Proposed

Here is where it gets interesting.

Codex did not just identify problems. It prescribed a complete treatment plan:

  • 44 controls (C-001 through C-044) across 5 implementation phases
  • 12 microtask packs estimated at approximately 9 hours of work
  • A 30-day rollout plan with weekly milestones
  • A "Better Us" operating system including:
    • Operating principles
    • A behavioral contract (Codex commitments and George commitments)
    • Daily, weekly, and monthly review rhythms
    • A three-tier decision framework
    • A scoreboard with tracked metrics
  • Evidence blocks with SHA-256 hashes for audit integrity

The prescription was thorough, sophisticated, and — for a 1-person operation — disproportionate.

The Meta-Insight

Here is what I realized watching this unfold: audit models do what audit models do.

If you ask a model to find problems, it will find problems. If you give it permission to prescribe solutions, it will prescribe comprehensive solutions. It does not have organizational context. It does not know you are one person running the whole operation yourself. It sees a system, identifies gaps, and proposes enterprise-grade remediation because that is what thoroughness looks like from the outside.

This is not a criticism of Codex. The audit was good. The findings were real. But the gap between "here are your problems" and "here is the proportional fix" is where human judgment lives.

No model fills that gap for you.

The Team's Assessment

I brought the audit results to my team — a software development specialist and a strategic advisor. Both assessed independently. Both converged on the same conclusion.

The diagnosis is good. The prescription is over-engineered 6x.

Here is what the team surfaced:

The evidence blocks are theater. Codex proposed SHA-256 hashes on evidence blocks and status gates like "No task execution until status: PASS." Sounds robust. But Claude Code has no middleware. There is no enforcement layer that reads a hash, verifies it, and gates execution. The hash is generated by the session. The hash is verified by the session. Self-referential verification. Same enforcement class as writing "MUST" in capital letters — instruction-only, dressed as structural.

This is F-DESCRIBE from my own taxonomy: describing enforcement without building enforcement.

The scale concerns are premature. Multiple findings flagged "10x scale bottleneck" risks. I have 1 user. Me. A lot of moving parts, but one person at the controls. Solving for 10x scale when you are at 1x is not engineering — it is speculation.

The operating system is bureaucracy for a team that does not exist. Weekly scoreboard reviews. Monthly safety posture assessments. A behavioral contract between George and Codex. These are governance structures for an organization. I am not an organization. I am a person with good infrastructure.

What We Actually Did

Five fixes. Two hours of work. Cherry-picked from 44 controls.

1. Remove the legacy agent from routing tables. The old adversarial-reviewer-agent was still listed. Removed. Clean routing.

2. Add a note clarifying the Rival's architecture. The Rival is not the agent-ownership pattern used by the rest of the system. It is a two-participant model: session plus Codex. The documentation now says that explicitly.

3. Fix the F-code count. The playbook said 10. The catalog has 14. Updated.

4. Add the anti-suppression rule. Classification is additive only. You can escalate a finding. You cannot downgrade one. Now explicit.

5. Restrict rival fast to non-security artifacts. Quick reviews should not be used on security-related work. The speed tradeoff is acceptable for content; it is not acceptable for security.

Five fixes. The other 39 proposed controls were rejected. Not because they were wrong in theory — many were sound engineering practices. They were wrong in context.

What This Teaches

Cross-model audit works. I would do it again tomorrow. But here is what the experience clarifies:

The audit is not the answer. The audit is the question. Codex gave me a list of problems. My job — the human's job — was to evaluate which problems are real, which are proportional, and which are solutions looking for problems to justify their existence.

Models audit maximally. An AI auditor does not have a "proportional response" setting. It finds everything it can find and proposes the most comprehensive fix it can imagine. That is correct behavior for an auditor. It is the human's job to apply organizational context.

Cross-model review catches what self-review cannot. The routing contradictions, the F-code count drift, the implicit rules that should have been explicit — I missed all of these in my own reviews. A fresh model with no history and no relationship found them in one pass. Cold review works.

The prescription reveals the prescriber. Codex's 44-control plan tells you exactly how Codex thinks about remediation. Understanding that is as valuable as the findings themselves. Audit models default to comprehensiveness. Your job is to filter for relevance.

Your AI Should Audit Your AI

If you are building anything serious with AI — and "serious" just means "you depend on it" — you should have a different model reviewing the work periodically.

Not the same model. Not the same session. A structurally independent reviewer that does not share your history, your assumptions, or your blind spots.

Package your system. Define your dimensions. Ask the hard question: "Am I Safe?™"

Then apply judgment to the answer.

The model tells you what is wrong. You decide what matters.


Your AI should audit your AI. If you want to learn how to set up cross-model review for your own systems, reach out.

Part of the Human-AI Workflow Failures series. Earlier pieces cover phantom completions, the F-code taxonomy, and the Am I Safe? framework. Continue with 44 Controls for a 1-Person Operation — why I said no to 39 of 44 proposed controls, and the framework for proportional governance.


George Samuelson is a technologist with a B.S. in Computer Science from The College of New Jersey and winner of NJ Ad Club's "Jersey's Best Marketing Under 40." He has built digital systems for multi-million dollar businesses and hospital networks across 10 locations. He consults on AI implementation, automation strategy, and scaling solo operations.

ai-auditcross-model-reviewsystemscodexclaudeadversarial-reviewframeworks
Share: