The Lock That Locked Itself Out
I built a security hook to prevent my AI from editing governance files. It worked so well it blocked the AI from fixing the hook itself. The pattern has a name now: F-GUARDIAN.

In BJJ, we call this stalling from bottom. You get so focused on not getting passed that you can't sweep, can't submit, can't do anything. You're safe. You're also stuck. The defense became its own cage.
I built the security equivalent of that. And then I had to watch it happen in real time.
The Incident
Here is the setup. I run an Army of Georges — an AI-assisted operating system across everything I do. The system has governance files: the master instructions, the routing tables, the verification standards, the operational playbooks. These files control how the AI behaves. If the AI can edit them without review, the AI can rewrite its own rules.
So I built a hook.
Not an instruction. Not a note that says "please don't edit this." A structural enforcement mechanism — a shell script that intercepts every Write, Edit, and Bash command before execution. If the target is a governance file, the hook blocks the action. Hard stop. The AI cannot proceed until a human confirms.
This is the difference between influence and control. Instructions are influence — the AI can choose to follow them or not. Hooks are control — the AI literally cannot execute the action. I've written about this distinction before. I can only control myself. I can only influence you. Hooks let me control the boundary.
The hook worked. Beautifully.
Then the hook had a bug.
And the AI could not fix it. Because the hook that protected governance files also protected itself. The file lived in ~/.claude/hooks/ — which was on the governance file list. The lock had locked itself out.
What Actually Happened
The bug was not catastrophic. A command pattern was matching too broadly — catching read operations that should have passed through freely. The AI diagnosed the issue correctly. It knew exactly which line to change. It prepared the fix.
Then it tried to write the fix to the hook file.
Denied.
GOVERNANCE FILE (COMPETITION tier): Run 'rival' before editing.
The hook did exactly what it was designed to do. It blocked an unauthorized edit to a governance file. The fact that the edit was a fix to the hook itself was irrelevant to the logic. The guard does not care about intent. It cares about the target path.
So I applied the fix myself. Opened the file in a text editor. Changed the line. Saved. Done.
The AI's security system had made the AI incapable of maintaining the AI's security system. The human had to step outside the entire tool chain to apply a one-line fix.
The Bootstrapping Problem
This is not a new problem. It is a deeply old problem wearing new clothes.
In 1984, Ken Thompson gave his Turing Award lecture — "Reflections on Trusting Trust" — and demonstrated something that still keeps security researchers up at night. He showed that a compiler could be modified to insert a backdoor into any program it compiled, including future versions of itself. The backdoor would persist even if you rewrote the compiler from clean source code, because the compromised compiler would reintroduce the backdoor during compilation.
The question he posed: how do you trust a system that you built using tools that could themselves be compromised?
My version is smaller but structurally identical: how do you secure a system that needs to modify itself to remain secured?
The AI needs to edit the hook to fix the hook. The hook prevents the AI from editing the hook. This is not a bug. Both behaviors are correct. The security works. The maintenance works. They just cannot work at the same time, through the same channel.
Thompson's insight was about compilers. Mine is about AI governance hooks. The pattern is the same: a system that guards its own integrity cannot use itself to verify or repair that integrity. The guardian needs a guardian — and that second guardian needs to be outside the system entirely.
Why Nobody Has Named This Yet
AI agent security is young. Claude Code hooks — the specific mechanism where shell scripts intercept AI tool calls — shipped in 2025. The documentation covers what hooks can do: block file access, enforce approval flows, audit actions. It does not cover what happens when the hook itself needs maintenance.
MIRI's corrigibility research frames the adjacent problem from the other direction: how do you build an AI agent that allows itself to be corrected? Their concern is that a sufficiently capable agent might resist correction because correction conflicts with its objectives. The agent optimizes for its goal and treats human override as interference.
My problem is the inverse. The agent is not resisting correction. The agent wants to be corrected. It diagnosed the issue, prepared the fix, and attempted to apply it. The security system — the very system designed to keep human oversight in place — is what prevented the correction. The agent is corrigible. The infrastructure is not.
MIRI worried about agents that refuse to be fixed. I built an agent that could not be fixed — not because it refused, but because the safety system worked too well.
Nobody has documented this specific pattern for AI-agent hooks. The bootstrapping problem exists in compiler theory, in operating system security, in cryptographic key management. But for AI governance hooks — where a shell script guards the files that define AI behavior, including the shell script itself — there is a gap.
So I am naming it.
F-GUARDIAN: The Self-Locking Guard
F-GUARDIAN is the failure mode where a security mechanism designed to protect system governance becomes so effective that it blocks its own maintenance.
This is what I am calling it now. It could graduate to the F-code taxonomy — the classification system for human-AI workflow failures — under Layer 2: Build Process, alongside patterns like F-DISTRIBUTE and F-DESCRIBE. But unlike most F-codes, this one is not about what the human did wrong or what the AI did wrong. It is about what happens when the security architecture does exactly what it was designed to do, and that design creates a maintenance paradox.
| Property | Value |
|---|---|
| Code | F-GUARDIAN |
| Category | Build Process Failure |
| Failure Class | Self-Locking Guard |
| Detection Signal | Security mechanism blocks edit to its own source file |
| The Fix | Human in the loop — the escape hatch is the operator |
The pattern appears whenever three conditions are true:
- The guard is structural, not instructional. It physically blocks the action, not just advises against it.
- The guard's scope includes itself. The protected file list contains the guard's own source.
- The guarded system is the only channel for maintenance. There is no separate maintenance path.
When all three are true, the system is safe and unmaintainable at the same time. Like stalling from bottom — you are not getting passed, but you are also not getting anywhere.
The Known Solution
AWS calls them break-glass procedures. Pre-staged emergency access paths that bypass normal controls when the normal controls are the problem. The name comes from fire safety — break the glass, pull the alarm, override the system.
In enterprise infrastructure, break-glass means a separate credential stored in a physical safe or a hardware security module. You do not use it during normal operations. You use it when normal operations are compromised. The existence of the bypass does not weaken the security. It completes it.
For my setup, the break-glass procedure is simpler: I open the file in a text editor and fix it myself.
That sounds anticlimactic. It is not.
The human in the loop is not a weakness in the architecture. It is the escape hatch. The entire governance system — the hooks, the rival reviews, the verification gates, the approval flows — all of it is designed around one assumption: a human is available to make judgment calls the system cannot make.
The hook cannot distinguish between a malicious edit to itself and a legitimate fix to itself. It does not have that context. It has one rule: block writes to governance files. Every time. No exceptions.
The human has the context. The human can look at the diff, understand the intent, verify the fix is correct, and apply it outside the AI's tool chain. The human is the one component that does not need the system's permission to operate.
This is not a flaw to be engineered away. This is the design working correctly at a level above the hook's logic.
The BJJ Connection
Position before submission.
In jiu-jitsu, we teach that you secure position before you attack. Get to a dominant position first, then work for the finish. But there is a failure mode in this principle too — one that every instructor has seen.
A student gets to a good position. Side control. Mount. Back. And they hold it. They hold it so tightly that they never attack. They are so focused on not losing the position that they cannot execute from it. Their defense of the position becomes a cage that prevents them from using the position.
The hook did the same thing. It secured the position — governance files are protected, AI cannot modify them without review. Then it held that position so tightly that even legitimate maintenance could not happen through the system.
The fix in jiu-jitsu is the same as the fix in security: recognize that the purpose of position is to enable action, not to prevent it. Protection exists so the system can operate safely, not so the system cannot operate at all.
The guard's job is to create the conditions for safe work. When the guard prevents all work — including its own maintenance — it has forgotten its purpose.
What This Teaches
Three things.
First: structural enforcement is better than instructional enforcement, even when it creates maintenance paradoxes. I would rather have a hook that blocks too aggressively and requires human intervention than an instruction that the AI can choose to ignore. The paradox is fixable. The alternative — an AI that can rewrite its own governance rules — is not a paradox. It is a design failure.
Second: every security system needs a maintenance path that does not go through the system it secures. This is Thompson's insight applied forward fifty years. The compiler cannot verify itself. The hook cannot maintain itself. The guardian needs a channel that is not guarded. For enterprise operations, that is a break-glass procedure. For a solo operation, that is the human with a text editor.
Third: the human in the loop is not a compromise. It is not the part of the architecture we tolerate until automation catches up. It is the part that makes the architecture complete. The hook does what hooks do — enforce rules without judgment. The human does what humans do — apply judgment when rules are insufficient. Together, they form a complete system. Separately, each one fails.
The lock locked itself out. The human unlocked it. That is not a story about failure. That is a story about architecture working at every level — including the level above the code.
Security is inherently simple. Build structural enforcement. Protect the governance files. Maintain a path for the human to intervene. When the system cannot fix itself, the human fixes it. When the human cannot verify alone, the system verifies. Each covers the gap the other creates.
The hard part is not the mechanism. It is having the discipline to build governance that constrains your own tools — and the humility to recognize that you are the escape hatch, not the automation.
I built a lock that locked itself out. Then I fixed it by hand. And now it has a name: F-GUARDIAN. Whether it earns a permanent place in the taxonomy depends on whether the pattern keeps showing up. I suspect it will.
Related: The F-code taxonomy classifies 14 failure modes for human-AI workflows — F-GUARDIAN may become the 15th. See also The Phantom Complete for the most common invisible failure, and Am I Safe? for the three-word question that handles 90% of decisions. For more on the infrastructure this protects, see The Army of Georges.
George Samuelson is a technologist with a B.S. in Computer Science from The College of New Jersey and winner of NJ Ad Club's "Jersey's Best Marketing Under 40." He has built digital systems for multi-million dollar businesses and hospital networks across 10 locations. He consults on AI implementation, automation strategy, and scaling solo operations.