Kepler: Auditable World Models for ARC-AGI-3

Wensen Wu · Independent Researcher

Public technical report, September 1, 2026. Revised September 30, 2026.

Abstract

ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a $777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.

Scope and revision

Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026, not the main conference. This public revision is not a camera-ready submission receipt. Presentation format is not confirmed.

Our replay scores concern the public development set, not official first-exposure evaluation. ARC Prize imposes a five-times-human-baseline per-level evaluation budget; our earlier denial was incorrect. Conditional checks do not guarantee verified execution on every action. No new performance or causal reliability result is claimed.

Failure taxonomy, implementation limits, and reproduction scope · Resource accounting · Exploratory negative result

Cite the paper

@techreport{wu2026kepler,
  author = {Wensen Wu},
  title = {Kepler: Auditable World Models for ARC-AGI-3},
  year = {2026},
  url = {https://kepler-harness.vercel.app/paper/},
  note = {Public technical report, revised September 30, 2026. Accepted to the non-archival Interpreting Agent Behavior workshop at NeurIPS 2026}
}

Specify the software commit and dataset revision when reusing artifacts. No arXiv identifier or Scholar indexing is claimed.

Author homepage · Existing author Scholar profile