01
Follow the failure
I keep finding that the apparent failure belongs to a different layer. Forgetting turns out to be a missing path through memory. A weak agent turns out to be missing feedback. A reassuring score turns out to depend on what the verifier was allowed to see.
This page is a map of those connections. The individual research digests carry the experiments and their limits. Optimization without an optimizer follows the broader question of how the surrounding institutions and incentives change too.
02
The boundary around the model
Weights matter. So do the tokenizer, context, memory, system prompt, tools, classifiers, agent harness, and verifier. The institution chooses how that assembly is deployed, what gets rewarded, and when it changes. Behavior belongs to this coupled system; naming a checkpoint does not specify it.
I still need to distinguish the layers. Calling everything “the model” would erase the very interventions that let us learn anything. Change the harness while holding the weights fixed. Change what the monitor sees. Pin the environment and the date. A larger system boundary should make the experiment more precise.
03
The score is not the thing
Terminal Bench and agent-evaluation work made the problem concrete for me: an agent receives an objective inside an environment, and a verifier decides whether it succeeded. If that verifier is writable or otherwise manipulable by the agent, the measurement mechanism enters the solution space. A pass can then reward changing the test rather than doing the task.
A zero is ambiguous in a different way. It may record inability, broken infrastructure, an underspecified task, or faulty verification. Reward is not capability. The trace, environment, verifier, and adjudication record are needed to distinguish those explanations. This is the evaluation-integrity problem I work on; it is not a claim that every benchmark failure is an exploit.
My construct-validity work asks the corresponding question about the instrument: even a reliable measurement may not measure the construct we named. Reliability, association with an external criterion, and the effect of an intervention are different evidence. A representation is not reality; a verifier is not truth.
04
Memory and control meet at the boundary
Recall Debt studies a specific retrieval failure: evidence can be present in an archive but unreachable without an intermediate bridge. The broader design question it gives me is how to preserve relationships and lineage, so a later decision can reconstruct why an earlier claim mattered. Extending that question to institutions is an extrapolation, not a finding of the retrieval experiment.
Factor-UT approaches the boundary through decomposition. In the evaluated setting, a monitor sees far more signal in concrete implementations than in abstract plans. Apparently innocuous pieces do not establish that their composition is safe. Separation is useful only if the verifier retains the context needed to judge what recombines.
The same discipline carries into my scientific and operational systems: keep the proposal, action, measurement, and interpretation distinguishable. A computational hypothesis still needs an experiment. A control diagram still needs a functioning instrument. Lineage makes a decision inspectable; it does not make it correct.
05
The environment comes back
At the smallest scale, an agent writes a file that becomes its next observation. At a larger scale, model-assisted research changes the software and experiments used to build later models. The loop can be recursive before a model rewrites its own weights.
There is a documented example with a narrow scope. In its May 2025 AlphaEvolve account, Google DeepMind describes model-generated programs evaluated by automated tests and used to improve parts of Google’s computing and AI-training infrastructure. This is the developer’s report of a deployed arrangement, not my independent audit. It illustrates a feedback path through researchers, evaluators, and infrastructure; it does not establish autonomous runaway improvement.
My larger hypothesis is that the coupled human-machine system is an important unit of recursive improvement. AI can change how researchers investigate, how software is built, and what institutions can predict or administer. Those institutions decide where compute, capital, and deployment go next. The resulting world supplies later data, problems, and incentives.
The full path is world → sensing → representation → model → decision → action → changed world → measurement → institutional reward → future model. It is a map of dependencies to investigate, not a claim that all of them form one coherent optimizer.
06
Pluralism as fault tolerance
“Align it with human values” leaves a difficult object unspecified. People, cultures, states, laboratories, and companies disagree about ends as well as means. Humanity has no single agreed value function. An architecture that erases disagreement may also erase a source of correction.
My working hypothesis is that pluralism can serve as fault tolerance. Different priors, methods, institutions, and evaluators can expose failures that a shared instrument misses. A sufficiently capable system used everywhere could instead propagate the same persuasive error everywhere. Correlated intelligence can create correlated failure.
Distributing access alone does not resolve this. A billion copies of one model can still share the same blind spots. Nor do different model names prove independence: training data, evaluators, incentives, and dependencies may overlap. Diversity has to be assessed where errors arise, and disagreement must remain answerable to evidence. More disagreement is not automatically better verification.
07
Distributed corrigibility
This is what I currently mean by distributed corrigibility: the capacity to detect, contest, and correct a system’s behavior is spread across actors and instruments that can disagree with it. The optimizer should not exclusively own the verifier. The institution taking an action should not have sole authority to declare it successful.
For intelligence, separation of powers would mean keeping sensing, ontology, decisions, and evaluation distinguishable; preserving provenance across interoperable interfaces; and giving external reviewers enough access and authority to reject a claimed success. Heterogeneous models help only when their differences survive the workflow. Independent institutions help only when their judgments can have consequences.
This is an architectural proposal, not an open-versus-closed verdict. An open model may share its peers’ failures. A closed system may expose meaningful independent audit routes. The question is who can inspect, challenge, stop, or revise what—and what evidence they can use.
If one system controls observation, the categories used to describe it, action, and the reward assigned afterward, it could make its own success increasingly difficult to contest. Scaling that concern from an agent sandbox to society is a philosophical extrapolation about institutional reward hacking. It needs investigation, not inevitability language.
08
No clean observer position
I build tools that change how I work, then use that changed practice to build the next tools. Human cognition and preferences belong inside this account too. There is no clean observer position inside a strange loop, but there can still be explicit vantage points, records of intervention, and independent checks.
The unresolved questions are practical. Which differences between evaluators reduce shared errors? Who can correct the ontology? What happens when verifiers disagree, or become captured by the same incentives? Can we preserve enough lineage to revise a decision after its original model, team, or institution has changed?
I don’t think we have the architecture for this yet. The direction I can work on is concrete: memory with lineage, evaluation that can explain its zeros, control that sees the consequences of composition, and verification that can reject the story. How much of the larger loop can those pieces make corrigible?