01
Real instructions read like specifications
RFC 2119 gives terms such as MUST, MUST NOT, SHOULD, and MAY defined levels of force. The vocabulary was designed for interoperable protocols, but the underlying need travels: ordinary prose has to become a constraint another system can apply.
A consequential system prompt therefore behaves less like persuasive writing and more like a policy specification. Ordering, exceptions, prohibitions, defaults, and scope matter more than tone.
02
The pattern crosses layers
The same grammar appears in tool policies, coding-agent instructions, safety boundaries, and evaluation rubrics. A first-match-wins judge is an instruction hierarchy; an acceptance gate is a prohibition on advancing without evidence; a sandbox policy turns textual authority into bounded action.
This connects language to infrastructure. Modal words do not enforce themselves—the surrounding parser, model, harness, and verifier determine whether their force survives contact with execution.
03
A measurement, not an authenticity oracle
The presence of force language does not prove that a leaked prompt is genuine, and its absence does not prove fabrication. At most it is a structural clue about whether the text resembles an operational specification.
The useful experiment is controlled: write equivalent system-prompt baselines with and without explicit modal force, then measure instruction following, conflicts, side effects, and failure under paraphrase. That would test the pattern instead of turning it into lore.
04
Work with the architecture; measure what else moves
Prompting changes what the model attends to in context. Representation or weight interventions change a different layer. Neither intervention, by itself, identifies a unique internal concept.
If changing a refusal-related behavior also changes opinionation, identity, or unrelated task performance, that coupling is part of the result. A direction is evidence about a representation under an intervention—never automatically the direction of an idea.
05
Open methods without publishing an attack manual
Merit gates should be open enough for independent researchers to test claims. That does not require publishing account-farming tactics, live exploit payloads, or instructions whose main value is operational misuse.
My line is to publish the measurement, failure boundary, and design correction. If a page could be mistaken for a how-to attack a deployed system, it does not belong in this notebook.