ACCESS
Technology

Shape vs Fidelity: The Vocabulary That Makes Cross-Model Review Work

My content-strategist agent started rejecting good work after the model changed. Two words fixed the routing: shape and fidelity. The vocabulary that turns cross-model review from judgment calls into rule application.

George Samuelson
George Samuelson
·13 min read·Technology
Shape vs Fidelity: The Vocabulary That Makes Cross-Model Review Work

Opus 4.7 shipped. Within a couple of days my content-strategist agent was rejecting good work.

It flagged a clean draft as "missing warmth." Flagged the next one as "too direct, add a bridge." The drafts were fine. The agent was scoring them against a rubric written for Opus 4.6's prose, and 4.6 was no longer the model writing them.

The agent wasn't broken. It was calibrated to a model that had just stopped being the default.

Every operator hits this the first time a model they depend on gets replaced. The reviewer keeps its old eyes.

The obvious fix is "update the rubric." That's two problems wearing one coat. You have to work out what actually changed from 4.6 to 4.7. And you have to work out what should never change, no matter which model is writing. Those are different questions. One is about packaging. One is about substance. My rubric, before April 20, couldn't tell them apart.

Then two words fixed it.

Shape. Fidelity.

The adapter pattern covered why models don't improve in a straight line. Directive vs narrative covered the form rules take. The recursive governance vulnerability covered governance that modifies its own rules. Dual-Rival covered review structure. This piece is about the words you review with.

The problem with no vocabulary

Pre-April-20, a reviewer agent would flag something like "this feels a bit curt, maybe warm it up." I had no principled way to route that.

Sometimes the agent was right. The draft was genuinely curt, missing the teaching rhythm that makes it sound like me.

Sometimes the agent was wrong. The draft was correctly direct for 4.7's native voice, and "warm it up" was a 4.6 instinct. Taking the note would have broken the voice, not saved it.

Same sentence from the reviewer. Opposite correct responses. Nothing in my rubric told me which was which.

The reviewers weren't low quality. They were technically correct by their rubric every single time. The rubric was the problem. It measured voice against a 4.6 baseline and couldn't separate "the voice drifted off the author" from "the voice shifted to match the new model." The first is a real defect. The second is just a model being itself.

With no words for those two cases, every finding routed by gut. Some sessions I accepted packaging feedback and damaged a piece. Some sessions I rejected real substance gaps because I'd been burned accepting packaging feedback the week before. The rubric wasn't measuring the decision I actually had to make.

Two terms fixed the routing.

The two definitions

Shape is model-dependent packaging: the prose-surface choices a model makes that shift every time the model version changes. I measure it at the output. Here's what moves:

  • Words per section
  • "You already know this" bridges per section
  • Bullets vs inline italics for a list
  • Bold and emoji per 200 words
  • Opener style
  • Closing-line length
  • How often the prose pivots rhetorically

Each of those shifts per Opus version. I measured them. Phase 3 testing put the 4.5, 4.6, and 4.7 values in a table. Same author. Three different native packaging patterns.

Shape drifts when models change because shape is the thing models change.

Fidelity is cross-model substance: the properties of the work that must hold regardless of which model drafted it. These don't move when the model moves:

  • Does this sound like the author
  • Is the argument complete
  • Are the sovereign terms preserved word-for-word
  • Is it factually accurate
  • Is the bookend architecture intact
  • Did the register stay out of corporate-generic
  • Is the teaching still embedded in the example

A well-reviewed piece holds its fidelity no matter which Opus version drafted it. These are the lines I won't let a model move. If fidelity drops between versions, something is wrong that has nothing to do with shape.

What the vocabulary does to routing

Once the two words exist, every finding routes one of two ways.

Shape-class: session authority. My read on native voice for the current model beats the reviewer's 4.6-trained rubric. If the reviewer says "add a warming bridge" and the draft is already in 4.7 native shape, I reject the note. Not because the reviewer is wrong by its rubric. Because the rubric is aimed at a model that isn't writing anymore.

Fidelity-class: reviewer authority. The reviewer's read on substance beats my drafting intuition. If the reviewer says "there's a causal step missing between paragraph four and five," I accept and fix. Argument completeness is an analytical question. The reviewer is better at it than my gut is mid-draft.

Session authority over shape. Reviewer authority over fidelity. Every finding has a clear owner.

Voice looks like the exception that breaks this. It sits on the fidelity list ("does this sound like the author"), yet the reviewer's "add warmth" note gets rejected as shape. Both hold. The fidelity question is whether it sounds like the author at all. The shape question is whether it sounds like him the way the previous model did. The first is the reviewer's call. The second is mine.

Before the vocabulary, every finding was a judgment call. After, it's rule application.

Why this unlocked per-model calibration

The vocabulary made a specific piece of infrastructure possible: per-model scoring overrides. My SASSY/Voice adapter now carries a table of six ways a 4.6-trained rubric miscalibrates on 4.7 output. M1 through M6.

A few rows:

M1. 4.6 rubric: "dense prose with several teaching bridges signals care." 4.7 override: score whether the teaching move happened, not how many sentences it took. Expect a one-sentence teaching move on 4.7 where 4.6 used three.

M4. 4.6 rubric: "final draft 10-20% longer than the first means feedback got integrated." 4.7 override: length delta is not a fidelity signal. Score the final prose, not the direction it moved. Final can be shorter than first and still be better.

M5. 4.6 rubric: "moderate bold and emoji means intentional emphasis." 4.7 override: bold and emoji are optional on 4.7. Emphasis lands in sentence structure, not typographic markers.

Every M-row is a shape-class override. The fidelity scoring (voice, argument, sovereign terms) doesn't get a row, because it doesn't change between models.

That's the tell that the vocabulary is doing real work. Without it, M1–M6 is a flat list of rubric tweaks with no organizing principle. With it, the structure is obvious: shape gets overrides, fidelity doesn't. The vocabulary organized the infrastructure before it wrote a line of it.

F-NARRATE: the sharpest test of the distinction

The tightest test of shape-vs-fidelity ran on a single pattern. We call it F-NARRATE.

On 4.7, push the effort level to XHIGH and add a "think step-by-step" steering phrase, and the model starts narrating its own process before doing the work. "I have the file. Now I will identify the issue. First I will examine..." Phase 1 testing caught it in roughly a third of XHIGH runs: 1 of 3 in a small N=3 battery. Small sample. I'm calling it directional, not settled.

The question was simple, and the answer mattered. Is F-NARRATE shape or fidelity?

If it's fidelity, it's a cross-model failure mode. It canonizes in the Rival playbook and every review on every model has to catch it. If it's shape, it stays in the 4.7 adapter as a 4.7-specific note, and the 4.6 adapter says out loud that it does not apply there.

So we ran the probe. Same XHIGH-equivalent steering, on 4.6, in Phase 3 testing (S389). Result: zero preamble narration. 4.6 doesn't even expose XHIGH the way 4.7's five-level system does. The trigger itself appears 4.7-specific on current evidence.

F-NARRATE is shape.

There's a sibling pattern that looks almost identical: F-META, where the model produces analysis about the work instead of the work. Same surface symptom: words where output should be. Different disposition entirely.

F-NARRATEF-META
LocationPreamble-bounded: orientation prose before the actionAnywhere: interleaved or post-action summary drift
TriggerXHIGH + step-by-step steering (4.7-specific)General process drift, any model
ConsequenceDelays execution; the work still landsDisplaces execution; the work may not land
Intervention"Stop announcing, do it" / drop the effort level"Stop analyzing, produce" / process correction
SeverityRecoverableMore degraded

F-NARRATE is shape. F-META is fidelity. Tell them apart or you mis-treat the symptom: apply F-META's fix to a 4.7 F-NARRATE problem and the disease is still there.

That's not a stylistic point. It's an enforcement one. The 4.7 adapter carries F-NARRATE with an explicit "do not apply below 4.7" line. The 4.6 adapter carries the mirror: an explicit note that F-NARRATE does not apply there, pointing back to the 4.7 adapter for scope. The research is what supports that 4.6's equivalent drift is F-META, not F-NARRATE. Don't import the other model's mitigation. And there's a hard rule under it: F-NARRATE must stay scoped to the 4.7 adapter, and the cross-reference audit checks for exactly that. Not because the N=1 probe settles it, but because the rule is cheap and the cross-model mis-application it prevents is not. Directional-until-promoted is the same bounded-claim discipline the rest of this series runs on.

Without the vocabulary, F-NARRATE and F-META collapse into one blurry "the model is stalling" concept. A 4.6 session inherits F-NARRATE's fix: lower the effort level, use HIGH not XHIGH. It does nothing, because 4.6 doesn't have XHIGH in that form. You treated the wrong disease because you never named the two diseases apart.

With the vocabulary, the scope is clean. Shape patterns stay scoped to their model. Fidelity patterns canonize across all of them. Same surface, opposite disposition, and the words are the only thing holding the line.

Naming a thing makes it cite-able

Here's the part that isn't obvious until you've lived it.

Once "shape vs fidelity" is a defined term, it becomes cite-able. That changes what a spawn prompt can say.

My content-strategist agent's spawn composition has this line in it: score fidelity (does it sound like George, is the argument complete), not shape (length, bullet density, bold count). The reviewer knows what's in each bucket because the buckets are defined.

Without the vocabulary, that instruction has to be an enumeration. Don't flag bullet density, don't flag bold count, don't flag word count, do flag missing argument beats, do flag stripped jargon. And the enumeration still misses every case you didn't list.

Named terms become references. References make rules short. Short rules scale. Enumerated rules don't.

Once the term is named and used consistently, it stops being nomenclature and becomes an asset. "Shape vs fidelity" is in my teaching vocabulary now. It's cite-able in content; this article uses it that way. It's teachable in a course. It's the shared language in any conversation with my team, or with another operator working cross-model review.

Vocabulary is an asset.

This generalizes past content

The distinction came out of my content review. It didn't stay there. I started seeing it everywhere I evaluate model output.

Code review. Shape is style: naming, comment density, formatting. Fidelity is correctness, security, the algorithm you chose. Style drifts per model. Correctness doesn't.

Data analysis. Shape is how the analysis is narrated: tables or prose, terse or expansive. Fidelity is whether the conclusion is right. Narration drifts per model. The conclusion is model-independent, or it's wrong.

Response drafting. Shape is tone, length, formality, hedge density. Fidelity is whether you answered the question and whether the facts are true.

Any time you evaluate AI output across model versions or providers, the shape/fidelity axis is already there, whether you've named it or not. Rubrics that don't split the two miscalibrate at every model change. Rubrics that do can absorb a model swap with shape-class overrides and never touch fidelity scoring.

The pattern is domain-general. The decomposition is reusable the moment you introduce it.

How to build it

Here's the pattern I run. If you're standing up cross-model evaluation, it's five steps.

1. Decompose the rubric you have. For every criterion, ask one question: does this measure something that drifts per model, or something model-independent? Tag each one. You'll get two lists. The shape list is longer than you expect. Mine was, every time I've run it. Most "this feels off" criteria are shape.

2. Build per-model shape overrides. For each shape criterion, measure its native value on each model you care about. Store it as a per-model table. That's the M1–M6 pattern. Reviewing model X's output? Use model X's shape values.

3. Freeze the fidelity criteria. Don't version them. Same fidelity rubric for every model. If the fidelity rubric needs to change, that's rubric evolution, a separate operation from a model migration. Know which one you're doing.

4. Instrument the spawn prompts. Tell the reviewer agent explicitly: score fidelity, not shape; see the fingerprint reference for per-model values. Now it knows which decision path it's on.

5. Route triage by class. Findings come back, you classify each one. Shape-class: session authority, probably reject if the review is 4.6-calibrated against 4.7. Fidelity-class: reviewer authority, usually accept.

Pre-vocabulary, that's five judgment calls per finding. Post-vocabulary, it's rule application.

The vocabulary is the infrastructure

One principle worth stating flat out: the words you use to talk about AI output are part of the governance, not commentary on it.

No vocabulary, no clean routing. If you can't name the difference between shape and fidelity, your rubric can't route by it, because the category doesn't exist to route into.

Defining the words is what makes the build routable. Writing the definitions down, using them the same way every time, citing them in agent instructions: that is infrastructure work. It lives next to the config files, not in a glossary nobody reads.

Consistent words scale. Once the terms are stable, every new rule references them instead of re-deriving them. Review protocols, spawn prompts, training material, all composed from the same primitives.

Sovereign vocabulary is the one piece of infrastructure you build once and never write twice.

Closing

Two words moved cross-model review from judgment call to rule application. M1–M6 overrides became designable. Spawn prompts got a citeable routing rule. F-NARRATE and F-META stopped contaminating each other across models. None of that needed new code. It needed the distinction named.

If you're running AI evaluation at any scale, introduce the shape-vs-fidelity split, or whatever names fit your domain. Name it. Define it. Use it the same way every time. Watch the rubrics get cleaner and the model migrations stop hurting.

Vocabulary is an asset. Treat it that way.


Piece 5 of 6 in the Opus 4.7 series. Piece 6 — the amendment-drift surface (capstone) covers amendment activity as the drift surface: the place all of these patterns aggregate. Every piece ties back to the structural-fit principle from the adapter pattern.

shape-vs-fidelitycross-model-reviewai-governanceopus-4-7adapter-patternf-narratef-metamodel-migrationai-methodologysassy
Share: