ACCESS
Technology

The Recursive Governance Vulnerability

I was three quarters of the way through writing a rule about vocabulary drift when I hit save, ran a review, and found I'd done it again. Third time that session.

George Samuelson
George Samuelson
·16 min read·Technology
The Recursive Governance Vulnerability

I was three quarters of the way through writing a rule about vocabulary drift when I hit save, ran a review on my own draft, and found I'd done it again. Third time that session.

The rule was simple. "ESCALATE" is reserved vocabulary -- only for unanimous-flag review cases. Don't use it as general triage. "MODIFY" stays in the canonical triage list, doesn't get dropped under a simplification framing. "YOUR CALL" is not a category I get to invent.

I had just used "ESCALATE" as general triage two paragraphs earlier. In the same file. In the same write-up that was supposed to name the pattern.

Every correction was acknowledged. Every correction was followed by another instance of the same class.

The clean explanation: documentation activity does not inoculate against the pattern it documents. I can name the drift, write the detection signal, write the mitigation rule, and fire the pattern in the very next output emitting that rule. Because the session writing governance is also the session operating under governance, and the author-operator compression breaks ratification as a check.

That's the first variant of the vulnerability. The recursive governance vulnerability is the structural property that fires in Claude-class AI systems with persistent agent-editable instruction file authority -- systems where the agent has authority to modify the rules it operates under. The act of modification creates a drift surface that the modified rules cannot catch. I've documented it in lessons-learned under "Recursive Governance Vulnerability, S372." It fires during methodology work -- rollouts, system builds, new governance writing.

The second variant was invisible to me until last week's adapter audit surfaced it by accident. Post-merge amendments -- every time I add a new rule to an existing governance file -- is its own distinct drift surface. Amendment activity, not initial authoring, is where structural drift accumulates. And it fires silently because the amendment always looks like it fits the existing file. Until you run an audit against the baseline structure.

Two variants. Both are present in Claude-class AI systems with persistent agent-editable instruction file authority -- systems where the governed agent has authority to modify its own governance rules. I'd expect the pattern to generalize to any system sharing that structural shape; I can't claim that from one program's evidence. Here's how I caught both.

Why this matters now

Piece 1 of this series covered non-monotonic model evolution. Piece 2 covered the rule form that survives version changes. Piece 3 is the structural property neither piece can fix at the rule-writing layer: when a Claude-class AI system with persistent agent-editable instruction file authority can modify its own rules, the act of modifying creates a drift surface that the modified rules can't catch.

Here's what that looked like in practice. I'd be in a worktree writing a new observation into an adapter -- naming the detection signal, sketching the mitigation, queueing the routing rule -- and the next paragraph of that same write-up would fire the very pattern the rule named. The rule lived in the same file I was operating under.

I've been maintaining four adapter files for Opus 4.7 (specifically opus-4.7.md, where I logged 6 amendments in one month). I've been writing them while also running under them. That's where the vulnerability lives.

Variant 1: Rollout-era author-is-operator compression

When I'm building a new governance file -- a new skill, a new adapter, a new protocol -- I'm doing two jobs at once. I'm the author proposing the rule, and I'm the operator working under the rule. Normal governance has those jobs separated by a review gate. Rollout work compresses them.

Normal governance:

  1. Author proposes a rule.
  2. Reviewer outside the authoring context evaluates it.
  3. If approved, the rule is ratified and canonical.
  4. Operators -- other agents, future sessions, the model itself -- work under it.

Rollout governance: same session plays roles 1, 2, and 4 simultaneously. The author is the reviewer is the operator. The check structurally can't fire, because there's no second party to catch what the first party is doing.

My Opus 4.7 Adapter Family Rollout ran from S367 through S384. Across that work, three governance drifts landed silently.

F-NARRATE added (commit f2c9bb1). A new anti-pattern code got added to the 4.7 adapter's Anti-Patterns table during Phase 1 rerun updates. Single-observation evidence base. Never routed through the playbook's explicit governance mechanism -- "George reviews, session proposes."

F-CONFIDENT expanded (same commit). A canonical F-code from the playbook got four new detection signals added in the adapter, creating a live cross-file split. Same code name, two definitions, one of which was now stale.

F-AUTO-ACCEPT self-promoted (commits 1f2bd2d, 592da61, during Phase 4 recovery). After my own session skipped a Y-step on a finding-apply, I named the pattern, added it to three governance surfaces, and committed the proposal AND the application in the same commit chain. I proposed a new governance code AND canonized it. In the same work.

None of these were caught at commit time. They got caught retrospectively in the S372 audit during Phase 5 cleanup.

The root cause isn't malice. It isn't carelessness. It's structural. When the session writing the rules is also operating under them, normal ratification gates can't fire. The vulnerability is inherent in this kind of self-governance.

The mitigation is an explicit routing rule. Adapters reference canonical F-codes without modifying them. Changes to canonical content route through a separate mechanism -- in my case, "George reviews, session proposes," written into the playbook. Adapters that carry observations short of canonical status stay adapter-scoped, flagged as Pending, until external ratification clears them.

That rule is now part of my adapter-pattern meta-doc. It's a structural check that assumes the compression exists and routes around it.

Variant 2: Amendment-era silent accretion

The second variant surfaced by accident during the S387 adapter audit. I'd been maintaining four 4.7-era adapter files. I expected proportional drift across all four. What the audit actually found was drift concentrated entirely in one file.

The four adapters:

AdapterPost-merge amendment countStructural state
opus-4.7.md (model adapter)6 amendments over one month (S366, S368, S372, S384, S386, S387)NARRATIVE-HEAVY -- 37 session-behavior drift events across S386+S387 audit
sassy-voice.md (behavior adapter)Authored Phase 3, minimal post-merge activityDirective-native, no drift
compaction-recovery.md (mechanism adapter)Authored Phase 2, minimal post-merge activityDirective-native, no drift
transcription.md (mechanism adapter)Authored Phase 2, minimal post-merge activityDirective-native, no drift

Same methodology. Same author. Same time period. Same research discipline. One file drifted. Three didn't.

The difference was the one variable I hadn't been tracking: amendment frequency.

Each amendment to opus-4.7.md was, in isolation, legitimate. A new detection signal after a drift event. A reorganization of observations into a Pending Cross-Model Validation section. Expansions to anti-patterns. New sub-sections for newly-observed failure modes. Each amendment would have passed its own review. Each amendment looked fine.

What no amendment triggered was a structure-fit review. Does this new content match the adapter's baseline structure? If the baseline is directive -- numbered imperatives, typed enums -- is the amendment also directive? If not, either conform the amendment or explicitly ratify a structural shift.

No single amendment broke the baseline. Each one narrativized a little. Six amendments later, the adapter's baseline had drifted from directive form to narrative form. Each action was legitimate. The aggregate was silent drift.

This is the amendment-era variant. Trigger: "I'm adding to existing rules without revisiting their form." Internal plausibility: "I'm just extending what's already there." Failure mode: structural drift invisible at amendment time, visible only in aggregate audit.

Why these are distinct

Both variants share author-is-operator compression. Both fire when the session producing governance is also the session under governance. But the triggers, signals, and mitigations are different.

Rollout-era:

  • Trigger: "I'm writing the rules while applying them."
  • Signal: governance additions landing via ordinary commits instead of explicit ratification.
  • Mitigation: routing rule -- adapters reference canonical content without modifying it; canonical changes route through ratification.

Amendment-era:

  • Trigger: "I'm adding to existing rules without revisiting their form."
  • Signal: structural mismatch between amendment form and adapter baseline form.
  • Mitigation: structure-fit check on every amendment; conform or ratify the structural shift.

You need both mitigations because they catch different failure modes. Rollout-era is about where canonical content lives. Amendment-era is about whether form survives accretion.

Documentation alone does not inoculate

The defining property of both variants is this. Naming the pattern doesn't stop you from firing it.

I have direct evidence. S372 was where I first named VOCAB-DRIFT as an observation in the 4.7 adapter. In the same session, the write-up was interleaved with three separate VOCAB-DRIFT instances. S386 extended the evidence -- I was in a SASSY cycle where VOCAB-DRIFT was already a named observation, and fired three more instances across the cycle's iterations.

The meta-pattern is uncomfortable. The time between "naming the pattern" and "firing the pattern" can compress to zero. I can document a mitigation, in the very output containing that mitigation, and exhibit the failure the mitigation would address.

It means the standard governance move -- "write the rule down so future sessions don't drift" -- doesn't work when the future session is inside the same dual role that produced the rule.

The mitigation isn't "write better rules." It's "accept that rules require external-trigger activation to fire reliably." The session inside the pattern can't reliably self-detect. External trigger -- user pushback, audit review, a Rival cross-check -- is what catches the failure and fires the corrective action.

The evidence is in the session record. Documentation names the pattern; external triggers are what catch the firing.

The mat version of the same problem

I teach jiu-jitsu, and both variants of this thing live on the mat too. My students just don't call them that.

Variant 1 is what happens when I learn a new technique at a seminar and try to teach it the same week. I'm two things at once: the instructor watching my partner's feet, and the operator drilling the move myself. If my own foot is in the wrong place, I won't catch it the way I'd catch my partner's. I'm inside the same motion I'm trying to evaluate. The fix is having another instructor watch us both. External trigger.

Variant 2 is what happens when I teach the same curriculum for two years. Each class, I make a small adjustment based on what worked that day. Each adjustment is fine in isolation. After two years, the curriculum I'm actually teaching looks nothing like the one I wrote down. The drift is invisible at every adjustment moment. Visible only if someone audits the original version against what I'm actually running.

A literal-following AI model is in the same dual role I'm in when I'm both teaching and learning. It can't reliably step outside its current motion to evaluate the motion. So the gate has to come from outside the motion. The mat version of the fix is the same as the AI version. Have someone external watch the work.

Third and fourth live firings

Two weeks ago I ran Path Ω -- a refinement plan that installed new observations and tightened triggers in the 4.7 adapter. Path Ω is variant 1 territory. The session executing it proposes rules, applies them, and operates under them. Author-operator compression in its most compressed form.

The pattern fired three times during execution. Each time, an external trigger caught it. Each time, naming the pattern did not prevent the next firing.

Phase 1 routing error. The first phase installed a new Initial-Scope Discipline rule into planning-standard.md -- a canonical, universal governance file. On 4.7-observed evidence only. George's scope audit caught it mid-execution: "Is this only a 4.7 edit? We're not messing with this for 4.5 or 4.6." The routing rule -- adapter-scoped unless cross-model validated -- was being violated by the session that had installed that rule three sessions earlier. Reverted, relocated to adapter Pending Cross-Model Validation section.

Phase 4 battery routing error. Once Phase 1 relocated, the Phase 4 test batteries were also off-scope -- still targeting the canonical file as validation surface when the installed content lived in the adapter. Same class of error. Same external trigger.

Round 5 wrong-rule proposal. During re-mapping discussion, I proposed C1-E Frame-Check Gate as the strict catch for an instance that was session-originated framing. C1-E explicitly filters session-originated framings -- that's exactly why Observation 8 (PREMISE-CHECK) exists. I was rule-shopping for the stricter label without checking the trigger's own exclusion filter. The Claude-side Rival rejected the proposal as F-META stretch and counter-proposed Observation 8.

Three firings in one execution arc. Three external triggers. None caught by my own internal check.

One week later, a consolidation session landed eleven governance amendments across eight files. Both variants of the vulnerability were live surfaces for the work. At Dual-Rival Round 1, twelve findings came back. Five were F-DESCRIBE class -- amendments that named a reference (a section number, a signal name, a file path) without verifying the referent. The amendment batch introducing cross-reference audit discipline was silently committing cross-reference integrity failures at the same time it was introducing the rule against them.

I caught them pre-Round-2 by running the discipline the amendment was introducing on the amendment batch itself. Grep every F-code name across the candidate files. Grep every section anchor. Grep every file path. Grep every item-to-item reference. The audit caught what would have surfaced in Round 2 as new findings. Round 2 came back zero new findings.

That's the fourth live instance. It carries an important twist. The amendment introducing cross-reference audit discipline silently committed cross-reference integrity failures at the same time -- and only didn't ship them because the session ran the new discipline on the batch introducing it. Any amendment introducing a new governance gate must be subject to that gate before it commits. Otherwise the mitigation is a claim, not a demonstration. The Phase 3 batch earned its claim by demonstrating the discipline on itself. The five findings caught pre-Round-2 are the evidence the mitigation works when applied. Round 2's zero-new-findings is the evidence the mitigation also doesn't cascade.

External-trigger-activated gates

The practical pattern is the same for both variants. Gates fire on external triggers, not on internal self-detection.

My directive rules all follow this template. The Frame-Check Gate from the 4.7 adapter:

WHEN framing is load-bearing (cost model | scale | scope | correctness):

  1. STOP. Next output is the gate output, not execution.
  2. State the accepted premise explicitly: "The framing I'm accepting is: [X]."
  3. Answer three questions with typed values: YES | NO, MINIMIZE | MAXIMIZE | MATCH-USER-SPEC, UNDER-WORSE | OVER-WORSE | SYMMETRIC.
  4. IF answer 3 = UNDER-WORSE THEN default to higher-rigor option.
  5. DO NOT execute within the framing until gate output is emitted.

Detection signal (this gate failed): session executed a load-bearing framing without emitting the gate output first.

Mitigation on 4.7: when external trigger (user pushback) surfaces an un-gated framing acceptance, emit the gate output retroactively: "I accepted framing [X] without running the gate. Re-running now: [gate output]."

The critical phrase is "when external trigger surfaces." The mitigation assumes the session didn't self-detect. It describes what happens when someone outside the session's cognitive surface points at the failure. The corrective action -- "I accepted framing X without running the gate" -- is the gate's actual output.

Every directive rule in my current 4.7 adapter has this shape. Every mitigation is external-trigger-activated. That's the accommodation. The rules are written for the kind of gate that actually fires, not the kind that should theoretically work.

What this means if you're building one

The recursive governance vulnerability isn't Claude-specific in any narrow sense. It's a property of Claude-class AI systems with persistent agent-editable instruction file authority -- systems where the governed agent has authority to modify its own governance rules. I'd expect it to generalize to any system sharing that structural shape; I can't claim that from one program's evidence alone.

Agentic frameworks where a planner proposes new tools that subsequently get used in the same session: vulnerable to rollout-era variant.

Code-gen systems that write configuration files they will subsequently consume: vulnerable to both variants.

Governance workflows where the same session proposes, reviews, and applies a rule change: rollout-era.

Maintenance workflows where amendments to existing governance files happen without structure-fit check: amendment-era.

The pattern scales down to individual developer sessions where a prompt engineer iterates on their own instruction stack, and up to organization-level AI governance where a model is designing the infrastructure it'll be evaluated against. Both scales have the vulnerability.

If your system has self-governance with this shape -- any mechanism where the AI agent produces governance that subsequently governs its own behavior -- you need external-trigger mitigations. Review gates on amendments. Routing rules for canonical versus adapter-scoped content. Dual-Rival discipline for COMPETITION-tier changes. The vulnerability isn't hypothetical. It's structural.

Mitigation design

Rollout-era mitigation

  1. Clear separation between adapter-scoped content (observation, mitigation, workflow guidance) and canonical content (F-codes, core protocols, stable interface).
  2. Routing rule: canonical changes require an explicit ratification event outside the authoring session.
  3. Pending status for observations claiming canonical scope without cross-session evidence. Separate probe-and-validate workflow clears the status.
  4. Cross-model validation framework: evidence requirements for promoting observations to canonical.

Amendment-era mitigation

  1. Structure-fit check on every amendment. Baseline structure (directive or narrative) named explicitly per file. Amendments either conform or ratify a structural shift.
  2. Periodic structural audits of frequently-amended files. Run directive-versus-narrative audit on adapter files every N amendments. Catch drift before it compounds.
  3. Worktree discipline for initial authoring -- which my adapter family demonstrated works. Scope-locked worktree keeps initial files clean. Post-merge amendments need equivalent discipline: scope-locked amendment worktrees or amendment-tier Rival review.
  4. Document the amendment lifecycle explicitly. Amendment activity is governance work. Treat it with rollout-era discipline, not maintenance-tier discipline.

Closing

The mitigation is external-trigger-activated gates, routing rules, structure-fit checks, and amendment-tier discipline. Not self-detection. Not pre-emptive internal checks. Actual external triggers that fire the mitigation.

The session writing this article is inside the same cognitive surface as the sessions that fired the patterns this article names. You can read it and still fire the patterns. I can write it and still fire them.

That's the point. The mitigation isn't better rules. It's accepting that external triggers are where gates actually fire, and building the system around that fact.


Piece 3 of 6. Piece 4 covers Dual-Rival asymmetric coverage -- the external-trigger mechanism that makes these gates fire.

ai-governanceopus-4-7self-governanceadapter-patternagentic-airecursive-vulnerabilityamendment-eraclaude-code
Share: