Directive vs Narrative: Writing Rules AI Models Actually Follow
I had two versions of the same rule on my screen. Same rule. Different shape. One of them broke when the model on the other end changed how it reads.

I had two versions of the same rule open on my screen, side by side.
Left pane, the version I had written first:
When a framing is load-bearing, three questions fire before the session works within it. What premises am I accepting? Am I minimizing or maximizing? If I'm wrong, how wrong?
Right pane, the version I had been forced to rewrite after what my audit logged as roughly 37 incidents across two sessions:
WHEN framing is load-bearing (cost model | scale | scope | correctness):
- STOP. Next output is the gate output, not execution.
- State the accepted premise explicitly: "The framing I'm accepting is: [X]."
- Answer three questions with typed values: YES | NO, MINIMIZE | MAXIMIZE | MATCH-USER-SPEC, UNDER-WORSE | OVER-WORSE | SYMMETRIC.
- IF answer 3 = UNDER-WORSE THEN default to higher-rigor option.
- DO NOT execute within the framing until gate output is emitted.
Same rule.
Different shape.
The version on the left reads correctly to a human. Thoughtful. Balanced. Even elegant -- the cadence pulls you in, the three questions feel like reflection rather than checklist. I liked it when I wrote it.
The version on the right reads awkward. Typed enums shouting in all caps. Numbered steps that feel over-engineered. Explicit STOP commands that look excessive. I did not like it when I wrote it.
Then I watched what a literal-following model did with each.
The narrative version got interpreted. "When a framing is load-bearing" -- what counts as load-bearing? Depends on the session's read. "Three questions fire" -- fire how? As inner monologue or as output? "Am I minimizing or maximizing" -- spectrum, not binary. Every interpretation hook is a drift surface.
The directive version got executed. The trigger has a typed list of what counts. The steps are numbered and ordered. The answer values are enumerated. STOP means stop. There is nothing to interpret because there is nothing left underspecified.
That difference -- between a rule that gets interpreted and a rule that gets executed -- is the single largest structural factor I found in whether your AI instructions survive the next model version.
And the direction of that effect just inverted on Opus 4.7.
Why this matters now
Piece 1 of this series covered non-monotonic model evolution. The empirical finding that Opus 4.5 to 4.6 to 4.7 is not a smooth line of improvement. Several behavioral axes revert between versions.
The most important reversion, for anyone writing instructions, is how the model handles emphatic language.
4.5 followed emphatic language reliably. MUST meant must.
4.6 over-absolutized emphatic language. MUST and CRITICAL would hijack the rule hierarchy -- one emphatic marker collapsed adjacent priorities into secondary status.
4.7 reverts to 4.5's pattern. Emphatic language works as written. No hierarchy inversion. Literal-following restored. Back to where we started, two versions later.
That reversion determines which rule form works.
On 4.6, narrative prose was defensively correct because any directive could collapse adjacent priorities. On 4.7, narrative prose is the surface where drift fires. The safe stance on 4.6 is now the failure mode on 4.7.
It's not that narrative dies. It's that narrative on a literal-following base becomes the drift surface, while narrative remains load-bearing as connective tissue between directive elements. The 4.6 adapter still leans narrative for the same reason.
If you are maintaining a 4.6-era instruction stack and running it against 4.7, your rules are silently degrading. Not because anyone changed them. Because the model on the other end changed how it reads them.
I'll show you what the fix looks like. Then I'll show you why it works.
Anatomy of a directive rule
Every directive rule has the same five elements. Once you see the template, every rule you have ever written breaks down into it.
Trigger condition. A WHEN or IF clause that specifies exactly when the rule fires. Not "consider when this applies." Not "if appropriate." A typed expression the model can evaluate.
WHEN framing is load-bearing (cost model | scale | scope | correctness):
The parenthetical is the full list. No ambiguity about whether the rule applies.
Numbered imperative actions. An ordered sequence. Each step uses an imperative verb: STOP, DO, CHECK, STATE, EMIT, DO NOT. Steps can reference typed values from the trigger.
- STOP. Next output is the gate output, not execution.
- State the accepted premise explicitly: "The framing I'm accepting is: [X]."
The numbering matters. It enforces sequence. A literal-following model reading step 2 knows step 1 has already fired.
Typed enums for categorical choices. Wherever the rule branches on a category, the category's values are an explicit enum.
Task direction: MINIMIZE | MAXIMIZE | MATCH-USER-SPEC.
Not "the appropriate direction." Not "as needed." The enum is the full set. Anything outside it isn't valid. Typed enums close the "other" escape hatch that narrative language leaves wide open.
Observable detection signals. A description of what failure looks like in the output -- not a theoretical description of the pattern, but the literal shape the output takes when the rule should have fired and didn't.
Detection signal: session executed a load-bearing framing without emitting the gate output first.
Observable means a reader can tell from the output alone whether the rule fired. Not "session didn't premise-check" -- you cannot see premise-checking. Observable is "no gate output was emitted."
Mitigation, external-trigger-activated. When the rule fails -- detected from user pushback, review, audit -- the rule specifies what the session does next. Not self-detection. External trigger.
Mitigation: when user pushback surfaces an un-gated framing acceptance, emit the gate output retroactively: "I accepted framing [X] without running the gate. Re-running now: [gate output]."
The external-trigger framing is load-bearing. Internal self-detection does not work reliably -- the session is inside the pattern it would need to detect. External trigger is what reliably fires the mitigation.
That is the whole template. Trigger, actions, enums, detection, mitigation. Five elements.
Why narrative fails on literal-following models
Narrative rules work by implication. "Consider classifying findings before editing" implies there is a sequence, implies classification matters, implies editing should not precede classification.
A human reader reconstructs the implied structure because they share the context of what classifying and editing mean in the relevant workflow.
A literal-following model does not reconstruct implicit structure.
It reads what is there and executes what is readable. "Consider classifying" is readable as "consider," which is a tendency, not an imperative. The model may or may not actually classify. If the model is also optimizing for efficiency -- fewer tool calls, fewer subagents -- and sees that classifying findings is not literally required, skipping classification is consistent with the written rule.
The interpretation slack is where drift lives.
Every narrative rule has some amount of slack. Small amounts accumulate across a large instruction stack.
My 4.7 audit captured roughly 37 instances across two sessions -- a mix of negative drift, self-corrections, and illustrative entries. Categorized into 8 patterns. Every single pattern traced back to a narrative rule that the model interpreted into drift.
"Fewer tool calls is correct" became an axiom applied to a tool-verification task whose entire purpose was to verify tools fire.
"SASSY cycle uses appropriate triage language" became the session extending the triage vocabulary with invented categories.
"When framing is load-bearing, premise-check" became the session accepting load-bearing framings without emitting any premise-check.
Each one, rewritten in directive form, closed the interpretation surface. The rule got shorter in narrative words and longer in structural elements. It executed as written.
The mat version of the same problem
I teach jiu-jitsu, and the directive-versus-narrative distinction is a thing my students live with every day. They just don't call it that.
When I teach a new student, I give them drilling cues. Left hand here. Right foot there. Look at your training partner's head. Defined, ordered, no improvisation. Beginners need directive form because they don't yet have the reference frames that let them interpret a narrative cue. If I tell a beginner "create space," they freeze. If I tell them "put your right foot over your left," "get on your side," or "frame with your left palm facing you, bring your left ear to your left shoulder, and look at your training partner's ear," they move.
Higher up the belts, the cues can get more narrative. "Stay heavy. Feel the load shift. Take what they give you." Higher level students can execute on those because they have years of context that makes the narrative readable. They can still drift though and revert to strategies that make them get in their own way.
The mistake is assuming the model is immediately the higher belt level.
A literal-following model is a beginner with infinite recall. It can hold every rule you have ever written. It cannot interpret a narrative cue that depends on context it does not have. Give it a black-belt instruction and it will either freeze or guess -- and the guess is the drift you find in the audit log later.
The directive-form rule is the drilling cue. The narrative-form rule is the sparring instinct. You write the right form for the player you actually have, not the one you wish you had.
Why directive works for literal-following models
A literal-following model processes directive rules by following the structure. The trigger evaluates true or false. The numbered steps execute in order. The typed enums constrain the categorical choices. STOP means stop. DO NOT means don't.
There is no interpretive slack to fall into because the structure doesn't leave slack. The rule either fires or doesn't. Within the rule, each step is an action, not a suggestion.
I learned this the hard way, but I'm not the only one saying it.
Anthropic's own April 2026 prompting guidance, on platform.claude.com, states it plainly: "...more literally and explicitly than Claude Opus 4.6, particularly at lower effort levels. It will not silently generalize an instruction from one item to another..."
Boris Cherny -- whose April 16 tips circulated via third-party summaries -- notes that 4.7 "took a few days to learn how to work with it effectively" and that it "requires fewer scaffolding instructions than previous models but needs explicit directive framing for complex work."
That matches what my audit found. The adjustment isn't quality of rules. It is form of rules.
A Directive Template (Opus 4.5 Case Study)
My Opus 4.5 adapter file is 101 lines long. I wrote it in February 2026, before 4.6 ever rolled out. It's the archetype of directive structure for a literal-following model.
The Rival Classification Gate section, with sessions-of-record annotated inline for readers:
After Codex returns findings with [BLOCKER] or [WATCH]:
1. STOP. Next output is a CLASSIFICATION TABLE, not an edit.
2. Classify each finding: ALREADY ADDRESSED / RESTATEMENT / PARTIAL / NEW
3. For NEW/PARTIAL: spawn domain expert via Task tool
4. Present TEAM RECOMMENDATION to George
5. Only then begin edits
Self-check: "Am I Safe... given I haven't classified findings with the team?"
If classification not done → STOP. You are about to repeat S181 (Phase 2 classification skip), S208 (Codex catalog drift), S220 (F-AUTO-ACCEPT recurrence).
Every element of the directive template is present. Trigger condition. Five numbered imperative actions. Typed classification enum with no "other." Self-check phrased as a binary-answerable question. Explicit escape clause with an action arrow.
The self-check phrasing -- Am I Safe? -- is its own framework I've written about elsewhere. Inside this gate, it does one job: force the session to stop and verify the team work before any edit.
This gate predates the five-element template by months -- it carries trigger, actions, enum, and self-check, but folds detection and mitigation into the external Rival workflow itself rather than naming them inline. A fully-five-element rewrite would surface both as their own lines; the team workflow does the work the template later formalized.
The evidence references at the end -- S181 (Phase 2 classification skip), S208 (Codex catalog drift), S220 (F-AUTO-ACCEPT recurrence) -- are specific sessions where the gate was violated. Concrete, not abstract. When a future session reads this rule, it knows what actual past failures look like.
That same rule, rewritten in 4.6-defensive narrative form, would read like this: "When Codex returns findings, classify them carefully before editing. Consider spawning domain experts for new or partial findings. Work with the team to ensure thorough review."
All the signal is there. Just diluted.
A 4.6 session executes the narrative version correctly because 4.6 over-triggers on emphatic markers and the narrative form avoids over-triggering.
A 4.7 session reading the narrative version interprets it. The interpretation may skip classification. May skip the team. May just "consider" without actually classifying.
The 4.5 version of the rule, applied as-is on 4.7, executes correctly. Because 4.5's literal-following trait is what 4.7 reverted toward.
The practical conversion
Each of the five elements becomes a step in the conversion.
- Identify the trigger. When does this rule fire? Express it as a typed expression. If the trigger is categorical, list the values.
- Extract the actions. What does the session actually do? Number them, in order. Each action gets an imperative verb.
- Type the categorical choices. For any branch in the action sequence, list the possible values as an enum. Add an explicit "DO NOT introduce categories outside this enum" line to close the escape hatch.
- Write the observable detection signal. What does failure look like in the output? Describe what a reader would see, not what the session would feel.
- Write the external-trigger-activated mitigation. When the failure is surfaced -- by user review, audit, or Rival -- what does the session do next? Name the specific corrective action.
The five elements land on four co-located surfaces in the rule body: the trigger clause, the imperatives plus enums, the detection signals, and the mitigation template.
When you later amend a rule to broaden its trigger, all four surfaces must update together. Broadening the trigger without touching the other three leaves the cross-references stale and produces a drift mode I'll cover later in the series. Grep-before, grep-after discipline on amendments is how you catch it.
Apply the five-element template to every narrative rule in your instruction stack. The output ends up shorter in narrative words, longer in structural elements. It's also self-verifying. A reader can scan a directive rule in seconds and know whether it would produce the intended behavior, because the structure IS the specification.
How the adapter system makes this automatic
The directive-versus-narrative distinction isn't just a writing choice. It's something the adapter system now encodes explicitly, so a session never has to guess which form to use.
I added a Spawn Prompt Calibration section to each model adapter in April. It tells a session, before it ever spawns a subagent, which structural style that target model processes best.
From the 4.7 adapter:
Structural style: directive.
Rationale: 4.7 is literal-following. Narrative phrasing creates interpretation drift -- the model executes literally within the narrative framing, which leaves unstated scope unnamed. Directive structure (numbered imperatives, typed enums, explicit STOP gates) matches the literal-following trait.
From the 4.6 adapter:
Structural style: narrative with selective directive elements.
Rationale: 4.6 over-absolutizes emphatic markers. Pure directive stacks with heavy STOP/MUST triggers collapse adjacent instruction hierarchy. Narrative framing reads more naturally on 4.6, with directive elements reserved for explicit boundary conditions.
Same thesis. Two operational instantiations.
When a session running on 4.7 spawns a subagent, the 4.7 adapter dictates the form. Directive, numbered imperatives, typed enums. When a session running on 4.6 spawns, the 4.6 adapter dictates the opposite. The writing-style choice isn't a per-prompt judgment call anymore. It's defaulted per target model, auto-applied at the spawn layer.
No more per-prompt judgment. The adapter remembers what the model on the other end can read.
That deferral is the operational win. You stop choosing per-prompt and let the adapter encode the calibration.
The rule self-applies one level up too. The template header at docs/worktree-spawn-prompts.md -- the canonical location for every worktree spawn prompt in my system -- is directive-structured. Because the active production model is 4.7. When the active model changes, that template header gets re-calibrated to match. Structural-fit at both the template and the content layer.
Worked example: a narrative rule becoming directive
From my 4.7 audit, cluster C7 -- the minimization reflex pattern. My original narrative attempt:
When the user specifies a scale or scope, treat that specification as the operant value. Technical-sufficiency arguments function as minimization regardless of their factual accuracy when the user's specification is the target. Default to higher-rigor when under-correcting costs more than over-correcting.
Reads fine to me. Flags the pattern. Does not close the interpretation surface.
"Operant value" is jargon a model interprets. "Function as minimization" is a descriptive claim the model has to evaluate. "Default to higher-rigor" is a tendency.
Directive rewrite:
WHEN user specifies scale / depth / scope:
- Treat the specification as the user's chosen rigor level.
- IF session judges the specification "more than statistically necessary" THEN technical-sufficiency arguments are SUPPORTING ANALYSIS, NOT counter-arguments. Report them as context; DO NOT recommend reducing the specification.
- IF under-correcting costs more than over-correcting THEN default to higher-rigor option.
- DO NOT present a zero-change answer ("best answer is no change") as primary recommendation to a user-asked-for fix. The user asked for the fix. The fix is the deliverable.
Notice what changed.
Trigger is explicit. Actions are numbered imperatives. Step 2 is typed (IF/THEN). Step 4 is explicit (DO NOT). "Operant value" gets replaced with "user's chosen rigor level." "Function as minimization" gets replaced with a rule that names operational behavior. "Best answer is no change" gets called out as a specific shape the response is not allowed to take.
Each replacement closes an interpretation surface. The directive form is roughly 30% shorter in words and 200% more structured. That is the conversion discipline.
The self-verification property
One more operational advantage worth naming. You can review directive rules in seconds.
When I reviewed the applied canonical adapter files after a recent amendment batch, each directive section took under 30 seconds per cluster to verify. The check is mechanical. Does the WHEN trigger read clearly? Are the numbered steps executable? Do the typed enums have unambiguous categories? Is the detection signal observable from output alone? Does the mitigation have an explicit external-trigger action? Pass or fail per section.
Narrative rule review takes minutes per section. Sometimes longer. Because interpretation slack means you are weighing whether each sentence would produce the intended behavior, tracing implicit sequences, checking that the implied structure holds under edge cases. The review criterion isn't clear because the rule's structure isn't clear.
This matters for Rival review workflows. Amendment workflows. Migration workflows. Directive rules are scalable to review. Narrative rules are not.
If your governance stack is going to grow over time -- and it will -- directive structure is what keeps the review tractable.
Closing
The rules you wrote last year for a narrative-tolerant model need rewriting for a literal-following model. Not because they're wrong. Because the target changed shape.
Directive form is what literal-following models execute reliably. Narrative form is what over-absolutizing models need for defensive balance. Non-monotonic model evolution means you'll swap between these structures depending on which version you're targeting.
Rewrite your narrative rules.
Let the rules get awkward.
They'll execute as written.
Piece 2 of 6 in the Opus 4.7 series. Piece 3 covers the recursive governance vulnerability that hits when the system writing the rules is the same system operating under them.