01
From scores to diagnoses
Partially observed environments mix several capabilities into one outcome. The agent must see the world, infer the goal, discover hidden mechanics, and interpret feedback. A zero can come from failure at any one of those layers.
The benchmark removes or degrades World, Goal, Mechanics, and Feedback information one axis at a time. The resulting behavioral change becomes a diagnostic signature rather than another leaderboard number.
02
What survives across environments
Different environments expose different bottlenecks, and models can fail on the same axis in different ways. World degradation is consistently destructive; mechanics degradation ranges from fatal to recoverable depending on the environment and model capability.
Behavior and articulation also separate: agents can trigger or exploit a hidden mechanic without reliably stating the rule under the paper’s strict verbal filter. That is why behavioral traces matter more than confident post-hoc explanation alone.
- Capability is multidimensional, not a scalar attached to a model name.
- The environment and the model jointly determine the binding constraint.
- Prompted reflection is not evidence of behavioral adaptation.
03
Why the knockout frame travels
The same method can examine tools, memory, system prompts, demonstrations, and feedback channels. Instead of asking only whether an agent is capable, we can ask which information substrate its capability depends on.
04
Still open
The next challenge is to connect these signatures across tasks without pretending that one environment’s axes are universal. A diagnostic should remain specific enough to explain failure and general enough to compare systems.