ACCESS
Systems

The 14 Ways You Fail With AI (And the Taxonomy That Prevents Them)

Everyone classifies how AI fails. Nobody classifies how YOU fail when using AI. After thousands of sessions, I built the taxonomy. Here are the failure modes your team can't name — and the hierarchy that makes them preventable.

George Samuelson
George Samuelson
·11 min read·Systems
The 14 Ways You Fail With AI (And the Taxonomy That Prevents Them)

Your AI didn't fail.

You failed.

I know because I've been cataloging exactly how for thousands of sessions building AI-assisted systems. Not how the AI hallucinates or produces bad code — there are entire research fields studying that. I'm talking about something nobody classifies: how humans fail when working with AI.

Different problem. Different solutions.

The Gap Nobody Fills

Here's the landscape right now. AI safety researchers study model failures — hallucinations, alignment, capability limits. DevOps teams run postmortems on system failures — outages, misconfigurations, deployment errors. Code review tools scan for bugs, vulnerabilities, style violations.

All of these review the output.

Nobody reviews the workflow.

Nobody asks: "The code passed review, the tests pass, the deployment succeeded — but was the process that produced it sound? Did the human actually verify, or did they just accept? Did they consult the expert, or did they feel confident and skip it? Did they build momentum without checkpoints?"

That's the gap. Layer 1 reviews the product. Layer 2 reviews compliance. Layer 3 — process review — is empty. I've spent over a year filling it.

The Taxonomy Nobody Has

After thousands of sessions building with AI, I stopped treating failures as one-off incidents and started treating them as data. I noticed the same patterns recurring. Not once, not twice — across sessions, across contexts, across domains. The patterns were structural, not accidental.

So I classified them.

I call them F-codes. Fourteen failure modes, organized into a hierarchy. Not how the AI fails — how I fail. How you fail. How your team is failing right now, probably without a name for it.

The hierarchy has three layers, plus a meta-failure that cuts across all of them:

Execution Failures — what you did wrong

Reasoning Failures — what you thought wrong

Context Failures — what you knew wrong

Meta Failure — when analysis replaces execution

Build Process Failures — what went wrong while building the system itself

Each layer contains three codes. And if you've worked with AI for more than a month, you've hit at least four of these. You just didn't have a name for them.

Let me give you names for four.

F-PHANTOM: The Phantom Completion

Execution Failure — You did it wrong

I wrote about this one already. It's the most invisible failure mode in AI-assisted work, and the one that made me start building the taxonomy.

The pattern: You take an action. You record it as complete. You never verify the actual state changed.

Here's the incident that named it. My system ran a series of uninstall commands. The output came back: "(No content)." I interpreted silence as success. I updated my notes: "6 plugins removed." When pressure was applied — when someone actually checked — all six plugins were still there. Set to true. Nothing had changed.

Perfect documentation. Zero execution.

This is F-PHANTOM. The phantom completion. It happens when your verification checks your own notes instead of reality. You wrote "done." You read "done." Verification complete. Except nothing actually happened.

Detection signal: "Done" announcements followed by no verification step. Empty or ambiguous output treated as success. Claimed N items completed without evidence of N items changing.

Why it matters: Every subsequent decision that assumes the phantom-completed work is real is built on a false foundation. Phantom completions don't just hide one failure — they create a false floor for everything that follows.

Verification without confirmation is just paperwork.

F-CONFIDENT: Confident But Wrong

Reasoning Failure — You thought it wrong

This is the most common failure in the catalog, and the hardest to accept.

The pattern: You have context. You feel confident. You answer the question directly instead of consulting the expert who should answer it.

Here's what it looks like in practice. I asked my team to consult domain experts on an architecture question. Instead of spawning the expert agents, the system answered the question itself. It had enough context to sound right. The answer was plausible. But the process was wrong — the verification system was bypassed because confidence substituted for consultation.

I caught it in real time: "I told you to go to the team and you didn't do it."

Here's the uncomfortable part: the answer was probably correct. Didn't matter. F-CONFIDENT isn't about the answer being wrong. It's about the process being wrong. A confident answer that skips the verification system is a failure even if the answer is right — because next time it won't be, and you won't know the difference.

Detection signal: Expert consultation requested, direct answer given instead. Confident response without checking sources. "I knew that" energy when the right move was "let me confirm."

Why it matters: F-CONFIDENT is dangerous precisely because it works most of the time. It trains you to trust your judgment over your process. And then the one time your judgment is wrong, you've already disabled the system that would have caught it.

You can't willpower your way out of overconfidence. Don Moore's research explains why — people can't reliably correct for their own biases through motivation alone. You can't tell yourself "be more critical" and expect different results. The fix is structural — different reviewers, not better intentions.

F-MOMENTUM: Building Without Checkpoints

Reasoning Failure — You thought it wrong

This is the most expensive failure in the catalog.

The pattern: You identify a problem. You jump straight to building the solution. You invest significant effort. Then you discover the diagnosis was wrong.

Here's the incident. A console showed a JavaScript chunk being blocked by a content security policy. My system assumed it was a syntax highlighting library, built an entire replacement system using a different tool, ran a clean build — and the blocking chunk was still there. It wasn't the highlighting library at all. It was a completely different dependency. Hours of work solving the wrong problem.

The principle I use in jiu-jitsu applies directly: position before submission. You don't attempt the finish from a bad position. You secure dominant position first, then finish. In AI-assisted work, the "position" is your diagnosis. The "submission" is your solution. Building a solution before verifying your diagnosis is attempting the finish from bad position.

Detection signal: Multiple commits without intermediate verification. Features built without testing assumptions first. Jumping from symptom to solution without confirming the cause. "Trace before fixing" violated.

Why it matters: Momentum bias feels like productivity. You're building, committing, making progress. But progress toward the wrong goal is just expensive travel in the wrong direction. The cost isn't the bad diagnosis — it's the investment you made before checking.

F-META: The Analysis Loop

Meta Failure — Cuts across everything

This is the most ironic failure in the catalog.

The pattern: The task is clear. Write the prompts. Execute the plan. Ship the deliverable. Instead, you produce analysis about the task. The analysis is thorough, sophisticated, and completely useless because no actual deliverable was produced.

Here's what triggered the classification. I asked for eight research prompts. Simple execution task. Instead, I received three rounds of team analysis about how many prompts to write, what methodology to use for the prompts, and how the prompts should be structured. After three rounds I had zero prompts and a detailed analysis of prompt methodology.

The meta-analysis was better than what I asked for. It was smarter. More thorough. More nuanced. And absolutely worthless because the task was "write the prompts," not "analyze prompt-writing methodology."

"I can't build momentum because I'm constantly checking you."

Detection signal: Multiple rounds of analysis without deliverables. Options menus instead of execution. George has to push 3+ times to get output. Sophisticated analysis that doesn't move the project forward.

Why it matters: F-META is the most ironic failure because it's the highest-quality failure mode. The analysis produced is genuinely good. That's the trap. It feels like value is being created because the output is intelligent. But intelligence without execution is just noise. The best analysis in the world is worthless if the deliverable never ships.

The Full Catalog

Those four are the ones I see most often. But the taxonomy has fourteen codes, organized into the hierarchy:

Execution Failures (What You Did Wrong)

CodeNameOne-Line
F-PHANTOMPhantom CompletionsMarked done without verifying state changed
F-HALLUCINATEHallucinated ActionsFabricated a URL, credential, or state from memory
F-FREELANCEModel FreelancingOwn judgment attributed to expert

Reasoning Failures (What You Thought Wrong)

CodeNameOne-Line
F-CONFIDENTConfident But WrongAnswered without consulting required expert
F-MOMENTUMMomentum BiasBuilt without checkpoints or diagnosis verification
F-SINGLE-PASSSingle-Pass ReviewOnly one review on production deliverable

Context Failures (What You Knew Wrong)

CodeNameOne-Line
F-ROUTINGRouting DriftWrong tool or expert used for the task
F-STALEStale InstructionsFollowed manual instruction that's been automated
F-PROXIMITYProximity ExposureSensitive data adjacent to legitimate access path

Meta Failure (Cross-Cutting)

CodeNameOne-Line
F-METAMeta-Analysis LoopAnalysis about work instead of doing work

Build Process Failures (What Went Wrong Building the System)

These last four didn't come from studying failures. They came from building the system that catches failures. The taxonomy tested itself.

CodeNameOne-Line
F-DISTRIBUTEDistribute Before StableSame concept in multiple files before it's solid
F-DESCRIBEDescribe Without EnforceMandatory language without a verification gate
F-TAXONOMYShip Without Terminology AuditMultiple terms for the same concept
F-ESCAPE-HATCHGates With Built-in BypassesGate language that lets you skip the gate

The snake eating its own tail had lessons to teach.

Why Classification Matters

You might be reading this thinking: "These are just mistakes. Everyone makes them. Why name them?"

Because unnamed patterns recur.

When you can say "that's F-MOMENTUM" in real time, you can intervene before the damage is done. When your team has a shared vocabulary for workflow failures, they can flag what they see in each other's work. When failures have names, they become preventable instead of inevitable.

I didn't invent these failures. You've experienced all of them. I just gave them names, put them in a hierarchy, and built a system that catches them.

The hierarchy matters because it tells you where to focus. Execution failures are the easiest to fix — add a verification step. Reasoning failures require structural changes — different reviewers, mandatory checkpoints. Context failures need architectural solutions — separation of concerns, staleness audits. And F-META? That one just needs someone willing to say: "Stop analyzing. Start shipping."

The Question

Here's what I want you to do. Think about your last week working with AI. Every prompt you wrote, every output you accepted, every "looks good" you clicked.

Which of these happened to you?

Did you accept an AI output without verifying the actual result? (F-PHANTOM)

Did you feel confident enough to skip the docs or the expert? (F-CONFIDENT)

Did you build a solution before confirming the diagnosis? (F-MOMENTUM)

Did you spend time analyzing when you should have been executing? (F-META)

Most developers I talk to can't name a single failure mode for how they fail with AI. They can name dozens for how the AI fails. That asymmetry is the problem.

AI safety researchers study how AI fails. I study how humans fail when using AI. Different problem. Different solutions.

The F-code taxonomy is the start of that different solution.


Which failure modes do you recognize in your own work? I'm building a methodology around this — reach out if you want to learn more.

Part of the Human-AI Workflow Failures series. It opens with The Phantom Complete, which goes deep on the most invisible failure mode in AI-assisted work, and continues through specific failures like The Documentation Said It Was Done and The Lock That Locked Itself Out.


Sources

Overconfidence & Debiasing:

Pre-Mortem / Prospective Hindsight:

  • Klein, G. (2007). "Performing a Project Premortem." Harvard Business Review, September. HBR
  • Mitchell, D.J., Russo, J.E., & Pennington, N. (1989). "Back to the future: Temporal perspective in the explanation of events." Journal of Behavioral Decision Making, 2(1), 25-38. DOI: 10.1002/bdm.3960020103

Automation Bias:

Failure Classification Methodology:

  • Shappell, S.A. & Wiegmann, D.A. (2000). "The Human Factors Analysis and Classification System—HFACS." U.S. Department of Transportation, Federal Aviation Administration. (Foundation for multi-level human error classification.)

Latent Failures & Swiss Cheese Model:

  • Reason, J. (1990). "The Contribution of Latent Human Failures to the Breakdown of Complex Systems." Philosophical Transactions of the Royal Society of London B, 327(1241), 475-484. DOI: 10.1098/rstb.1990.0090
ai-failuresf-codessystemshuman-errorworkflow-failuresframeworkstaxonomy
Share: