Wensen Wu · Independent Researcher
ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a $777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.
Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026, not the main conference. This public revision is not a camera-ready submission receipt. Presentation format is not confirmed.
Our replay scores concern the public development set, not official first-exposure evaluation. ARC Prize imposes a five-times-human-baseline per-level evaluation budget; our earlier denial was incorrect. Conditional checks do not guarantee verified execution on every action. No new performance or causal reliability result is claimed.
Failure taxonomy, implementation limits, and reproduction scope · Resource accounting · Exploratory negative result
@techreport{wu2026kepler,
author = {Wensen Wu},
title = {Kepler: Auditable World Models for ARC-AGI-3},
year = {2026},
url = {https://kepler-harness.vercel.app/paper/},
note = {Public technical report, revised September 30, 2026. Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026}
}
Specify the software commit and dataset revision when reusing artifacts. No arXiv identifier or Scholar indexing is claimed.