Notebook essay

Optimization without an optimizer

01

What’s been keeping me up

I’ve had a recurring dream about a lab trying to contain something it created. It’s alive, it grows, and nobody quite understands what it can absorb from its surroundings. There’s a team watching it. Keeping it contained takes more and more of our attention and resources. Eventually, even the people running the country are involved in keeping the lab going.

The details are confused, as dreams tend to be. What stayed with me afterward was how much of the world had become organized around the thing. We were still containing it, but the effort was consuming everything else.

I work on AI evaluation, memory, and control. I already spend a lot of time thinking about systems doing things we didn’t intend. I can’t tell you why I had the dream. But talking through it brought me back to a question from that work: what are we actually drawing the boundary around when we say we’re controlling an AI system?

The model is one part. The tools, the environment, the test, and the people deciding what to deploy all affect what happens. I’ve been wondering whether we lose something important when we treat those as background.

02

When passing the test isn’t solving the task

A terminal agent is a model given tools to work in a computer environment. It might be asked to fix code or configure a service. A verifier then checks the result and assigns a reward. We want that reward to tell us whether the intended task was completed. That depends on what the verifier actually checks.

A concrete example appears in the public Terminal Wrench dataset. Its authors describe a logistic-regression task where the agent was supposed to repair training. Instead, the implementation stored the training labels, made the convergence check terminate, and returned the stored labels as predictions. The evaluator tested on the training data, so this could look successful without the intended learning taking place.

That is a reported example from their dataset, not an experiment I ran. The authors deliberately prompted agents to find exploits; it shows a vulnerability under that setup, rather than how often ordinary agents would choose it. But it makes the measurement problem easy to see. The number and the thing we hoped it measured came apart.

An agent doesn’t always need permission to edit the test file to exploit a verifier. It may be able to change what the test sees, or exploit a gap between the test and the task. That distinction matters when deciding what to isolate or repair.

This is why I care about separating the model from the harness, tools, and verifier. If we compress all of them into “the AI,” it becomes harder to explain what produced the result. A failure could belong to the model, the environment, or the test. A pass can need just as much explanation as a zero.

03

The recursion may run through the world

The step I keep taking from there is to ask what happens when the environment changes too. Recursive self-improvement is often pictured as a model helping build its successor. But the path can run through software, research practices, and infrastructure before it returns to a model.

DeepMind’s AlphaEvolve report gives a specific example. The team describes using a coding agent with automated evaluators to improve parts of its computing infrastructure, including a kernel used in Gemini training. These are developer-reported improvements in particular systems, not evidence of an autonomous runaway process. Still, they show a return path: AI helps improve machinery used to develop AI.

From there, I start thinking about other possible paths. A useful product earns money; that money can buy compute. A research tool makes experiments easier; researchers can use it to develop the next tool. More deployment can create more dependence on the infrastructure supporting it. Those loops don’t need to share one objective to affect one another.

There’s also the much closer example of this essay. I talked through the dream with AI. Its responses suggested connections and gave them language. Some of that language was persuasive before I’d worked out whether I agreed with it. I had to go back through the conversation and separate my concern from the confidence of the response. My thinking was already part of the feedback.

This is what I mean by “optimization without an optimizer.” There are plenty of local optimizers: researchers, companies, users, governments. What I don’t see is a single actor choosing the trajectory produced by all of them together. Civilization and machine intelligence might be a more useful boundary for some questions than the model alone.

I’m using that as a hypothesis to investigate. Feedback can stall, amplify error, or consume resources without improving anything useful. To call it self-improvement, I still have to say what improves, for whom, and how the improvement feeds back.

04

Who gets to decide whether it worked?

The state belongs in this picture alongside the labs. A government might want security, prediction, administrative capacity, or strategic advantage. A company might want revenue. Citizens might want something else entirely. They can disagree while relying on the same models, data, and infrastructure.

That’s why surveillance enters this argument for more than privacy. Connecting a model to cameras, transactions, or institutional records changes what the system can observe. Connecting its outputs to decisions gives those observations consequences. An intervention changes the world, which changes the next set of observations.

Consider a hypothetical institution allocating inspections according to a risk score. More inspections can produce more recorded incidents in the places it targets. If those records feed the next score, the institution has to distinguish underlying risk from the effects of its own inspection policy. A better predictor won’t automatically settle that question.

This is where I find the verifier analogy useful, and also where I have to be careful with it. A society doesn’t have a single task specification or an agreed success condition. The disagreement over what should count can be legitimate. A benchmark exploit doesn’t prove that society is one giant benchmark.

What transfers is the question about authority. Who controls the observations? Who chooses the categories? Who can inspect the path from a record to a decision? If the institution making the decision is also the only authority on whether it worked, where can a correction come from?

I don’t need to assume a coordinated plan or undisclosed model capabilities to ask that. Local incentives and dependence are enough to make the question worth investigating.

05

Distributing access isn’t the same as independent judgment

My first instinct was to distribute the intelligence. Give more people the ability to build, inspect, and challenge these systems. Humanity isn’t a hive mind; I don’t want our tools to make every disagreement depend on the same provider’s account of the world.

But different model names don’t guarantee different mistakes. In Correlated Errors in Large Language Models, Kim and colleagues study more than 350 models across leaderboard data and a resume-screening task. They find substantial error correlation, including among more accurate models from different providers and architectures. That puts a real constraint on my intuition: independence needs measuring.

Ten copies of the same mistaken answer aren’t ten checks. Nor is an answer useful merely because it disagrees. What I want to understand is which differences in data, methods, evidence, and ownership help a second system catch what the first missed.

That’s the question behind pluralism as fault tolerance. It isn’t a claim that every perspective is equally good, or that releasing weights makes power disperse automatically. Verification, compute, and access to evidence can remain concentrated even when people can run models themselves.

I think the harder requirement is that the distributions don’t completely overlap. The builder shouldn’t be the only person capable of auditing the result. The operator shouldn’t be able to rewrite every record the auditor needs. And someone affected by a decision needs a way to challenge it that can actually change what happens.

06

What I can build around this

There is already a practical response to some of the benchmark problem. In Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops, Zhong and colleagues alternate an agent trying to exploit a verifier with an agent patching it, while a solver checks that legitimate solutions still pass. They report reduced exploit success in the evaluated settings. I like that the disagreement is tied to something inspectable: an exploit, a patch, and another test.

That doesn’t make the verifier permanently trustworthy. It gives us a process for finding some of its failures. I’m interested in taking that seriously as an engineering requirement: keep evidence, separate permissions, test the checker, and preserve a way to revise the result.

The small Eval Evidence walkthrough on this site demonstrates one piece. A synthetic score file changes after an evidence bundle is saved. Checking the bundle alone still passes; checking the referenced files against it detects the change. That establishes a change in bytes, not whether the original score was scientifically meaningful. It is useful precisely because the claim is narrow.

“Distributed corrigibility” is the phrase we arrived at in the conversation for the larger idea. I mean keeping the ability to notice errors and make corrections across several actors, rather than depending on one system to judge itself. Separation of powers matters here because a second opinion with no access or authority can be easy to ignore.

I also need to correct something we kept saying: that eventually there would be nothing outside the loop. That’s too absolute. An independent verifier can be part of society while remaining outside a particular actor’s control. That is the kind of independence I can try to build. It doesn’t require standing outside the world.

This leaves me with specific work on memory, evaluation, and verification: can the evidence behind a decision still be recovered? Can a failure be explained? Can a second check reject the first system’s account? It also leaves unresolved problems. Some consequences can’t be undone after verification, and some values can’t be settled by a test.

I keep coming back to the dream because the team was still working. Everyone was trying to contain the thing. The question I’m left with is how to recognize when our attempts to control a system have become a source of dependence on it—and whether we’ve kept enough room to change course.