01
The harness is part of the model-as-used
Closed providers can tune system prompts, routers, tools, memory, classifiers, and hidden policies around a checkpoint. That harness can add capability, remove capability, or redirect the probability mass of an answer without changing the product name.
An open checkpoint exposes a different object: more inspectable and recomposable, but without the provider’s proprietary control plane. Comparing the two as if they were interchangeable weights hides the actual system boundary.
02
Inputs have geometry
Tokenization, typography, encoding, retrieved context, and instruction hierarchy affect how an input is represented. Prompt injection is one visible case of a broader fact: changing the route into a model changes the region of behavior we observe.
This is a research question, not a claim that latent representations can be read directly from surface tricks. The useful discipline is to trace the transformation stack and test behavioral consequences without pretending we have seen inside the model.
03
Benchmarking the assembled system
When providers optimize hidden harness levers for public evaluations, the benchmark may measure a product-specific instrument rather than a stable model capability. The answer is not to ignore products; it is to name and pin as much of the assembled system as possible.