ACCESS
Systems

44 Controls for a 1-Person Operation (And Why I Said No)

An AI auditor proposed 44 controls, evidence blocks with SHA-256 hashes, and a behavioral contract for my solo operation. I implemented 5. Audits find problems. Judgment sizes solutions.

George Samuelson
George Samuelson
·9 min read·Systems
44 Controls for a 1-Person Operation (And Why I Said No)

An AI auditor proposed 44 controls for my solo operation. I implemented 5.

That is not stubbornness. That is judgment.

The Context

My AI-assisted operating system spans everything I do — the things I run, the ones I'm part of, the ones I don't talk about. Today it's past 1,500 sessions. When the audit I'm about to describe happened, it was 195 sessions and 55 files of architecture, governance, and operational standards. The system works. It has real governance, real verification gates, and a real adversarial review process.

I asked a competing model — Codex — to audit the entire system cold. Five dimensions. No relationship. No history. Tell me what breaks.

The audit came back with 24 findings. 11 HIGH, 12 MEDIUM, 1 LOW, zero Critical. The findings were real. Routing contradictions, enforcement gaps, documentation drift. Genuine issues I had missed in self-review.

Then came the prescription.

The Prescription

44 controls. C-001 through C-044. Organized into 5 phases across a 30-day rollout:

  • 12 microtask packs estimating 9 hours of implementation
  • Evidence blocks with SHA-256 hashes for audit integrity
  • A "Better Us" operating system with operating principles, behavioral contracts, and decision tiers
  • Daily standup rhythms, weekly scoreboard reviews, monthly safety posture assessments
  • Resume evidence gates: "No task execution until status: PASS"

Read that list again. Now remember: I am one person, running the whole operation myself with AI infrastructure.

The auditor proposed governance for an organization. I do not have an organization. I have a system.

The Reality Check

I brought the audit to my team. A software development specialist reviewed the technical controls. A strategic advisor assessed the operational proposals. They worked independently.

Both arrived at the same conclusion: the diagnosis is good; the prescription is over-engineered 6x.

Here is what they found.

Evidence Blocks Are Theater

The proposed evidence blocks looked robust. SHA-256 hashes on findings. Status gates requiring PASS before execution proceeds. RESUME_EVIDENCE blocks that verify system state on session start.

One problem: Claude Code has no middleware.

There is no enforcement layer between the instruction and the execution. No daemon checking hashes. No gateway reading status fields. The session generates the hash. The session verifies the hash. The session decides whether to proceed.

This is the same enforcement class as writing "MUST" in capital letters. The model reads the instruction and decides to comply — or does not. There is no structural mechanism that prevents non-compliance.

My own taxonomy has a name for this: F-DESCRIBE — describing enforcement without building enforcement. The audit proposed exactly the failure mode I had already classified.

The Scale Concerns Do Not Exist

Multiple findings flagged "10x scale bottleneck" risks. The system would struggle if it needed to handle 10 times the current load. Routing tables would become unwieldy. Governance standards would need version control workflows.

True statements. Irrelevant context.

I have 1 user. Myself. The system manages a lot, but there is one person at the controls. Solving for 10x scale when you are at 1x is not engineering. It is speculation dressed as prudence.

Would I add this if I were starting from scratch?

No. I would build for what exists and plan for what is likely. Governance for hypothetical future scale is the organizational equivalent of buying a 10,000-square-foot warehouse for your side project.

The Operating System Is Bureaucracy for a Team That Does Not Exist

The proposed "Better Us" operating system included:

  • Operating principles (4 foundational commitments)
  • A behavioral contract with Codex commitments and George commitments
  • Daily, weekly, and monthly review rhythms
  • A three-tier decision framework
  • A scoreboard tracking 6 metrics

Weekly scoreboard reviews. With whom? Monthly safety posture assessments. For what audience? A behavioral contract. Between me and a language model.

This is governance infrastructure for a team. I am not a team. I am a person with infrastructure. The overhead of maintaining these governance structures would exceed the risk they mitigate.

The Framework: Subtract First

This is where one of my operating principles does the work.

Subtract first. Before adding anything, ask: would I add this if I were starting from scratch?

I apply this in jiu-jitsu. I apply it in business. I apply it in system design. The instinct when something is wrong is to add — more controls, more gates, more process. The discipline is to subtract first and see what remains.

Here is the subtraction test applied to the 44 controls:

QuestionControls That PassControls That Fail
Does this fix a real issue found in the audit?539
Would I add this if starting from scratch today?539
Does the risk reduction justify the maintenance overhead?539
Does this work in a system without middleware?539

The same 5 controls passed every filter. The same 39 failed every filter.

The 5 Cherry-Picks

Here is what I actually implemented:

1. Remove legacy agent from routing tables. The old adversarial-reviewer-agent was still listed alongside the Rival that replaced it. Two paths to the same function. Simple cleanup, real clarity.

2. Clarify the Rival's architecture. The Rival does not follow the agent-ownership pattern used by the rest of the system. It is a two-participant model: session plus Codex. The documentation now states this explicitly. No ambiguity.

3. Fix the F-code count. The playbook referenced 10 F-codes. The actual catalog has 14. Documentation lagged behind the taxonomy. Updated.

4. Add anti-suppression rule. When classifying findings, classification is additive only. You can escalate a finding's severity. You cannot downgrade it. This was implicit. Now it is explicit.

5. Restrict rival fast for non-security work. Quick reviews trade thoroughness for speed. That tradeoff is acceptable for content review. It is not acceptable for security review. Now gated.

Two hours of work. Five targeted fixes. Each one addressing a real finding from the audit.

Why This Matters Beyond My System

The lesson is not "ignore audit findings." The lesson is: audits find problems; judgment sizes solutions.

Every AI auditor — every auditor, period — has the same structural incentive: comprehensiveness. Finding more issues is better than finding fewer. Proposing more controls is more thorough than proposing fewer. There is no penalty for over-prescribing and significant risk in under-prescribing.

This is rational behavior for auditors. It is terrible guidance for operators.

The operator's job is to apply three filters:

1. Is this a real problem or a theoretical one? Routing contradictions are real — they cause confusion today. "10x scale bottleneck" is theoretical — it assumes growth that may never happen.

2. Does the solution work in my actual environment? SHA-256 evidence blocks work in systems with middleware enforcement. In Claude Code, they are instruction-only controls with extra steps. Same enforcement class, more complexity.

3. Is the cure proportional to the disease? 44 controls for a 1-person operation is disproportionate by definition. The maintenance overhead of those controls would create more operational drag than the risks they prevent.

The Proportional Governance Framework

Here is how I think about governance for solo and small operations:

Level 1: Document, Do Not Automate

If a failure has happened once, document the pattern. Write it down. Give it a name. This is cheap, permanent, and effective. Most governance should live at this level.

Level 2: Gate, Do Not Police

If a failure has happened repeatedly and the cost is significant, add a gate. A single checkpoint at the right moment — not a surveillance system. One question at the right time beats ten dashboards.

Level 3: Automate Only What Recurs and Costs

If a failure recurs predictably, costs measurably, and can be expressed as a rule — then automate. Not before. This is the principle I call graduated determinism: lessons become rules, but only when they have earned the promotion through recurrence and cost.

The Anti-Pattern: Enterprise Controls for Solo Operators

The opposite of proportional governance is importing enterprise controls into a solo operation:

  • Weekly compliance reviews with no compliance team
  • Behavioral contracts with no counterparty
  • Evidence blocks with no verification layer
  • Scoreboards with no audience

These are not bad practices. They are bad practices for this context. The same controls that make a 50-person team accountable make a 1-person operation slower.

The Audit Paradox

Here is the paradox: the audit was valuable precisely because I rejected most of it.

If I had implemented all 44 controls, I would have spent 9 hours building governance infrastructure that creates maintenance overhead without proportional risk reduction. I would have been solving for problems that do not exist to prevent failures that have not happened.

By cherry-picking 5, I got the benefit of the audit — fresh eyes on real issues — without the cost of the prescription — enterprise governance for a solo operation.

The audit model did its job. I did mine.

The Question for You

If an auditor — human or AI — proposed 44 controls for your operation, how would you decide which ones to keep?

Here is the test I use:

Would I add this if I were starting from scratch today?

If the answer is no, it is governance for governance's sake. Subtract it. What remains after subtraction is what actually matters.

Audits find problems. Judgment sizes solutions.


Building AI-assisted systems requires governance — but governance proportional to your operation, not someone else's. If you want to think through what that looks like for your setup, let's talk.

Part of the Human-AI Workflow Failures series. This piece follows I Asked a Competing Model to Tear Apart My System, which covers the full audit story. The series begins with The Phantom Complete and the F-code taxonomy.


George Samuelson is a technologist with a B.S. in Computer Science from The College of New Jersey and winner of NJ Ad Club's "Jersey's Best Marketing Under 40." He has built digital systems for multi-million dollar businesses and hospital networks across 10 locations. He consults on AI implementation, automation strategy, and scaling solo operations.

governanceproportional-responseai-auditsystemsframeworkssubtract-first
Share: