Tencent's ExplorationBench Forces AI to Discover Hidden Rules Through Active Experimentation

Tencent, Fudan, and Tsinghua built two alien worlds with rewritten rules to test whether frontier models can discover new knowledge by probing an environment.

·
·
Tencent's ExplorationBench Forces AI to Discover Hidden Rules Through Active Experimentation
  • Tencent Hunyuan, Fudan, and Tsinghua release ExplorationBench, testing whether frontier models can discover unknown rules by probing.
  • Two sandboxes with 55 altered rules and 140 held-out tasks, graded by interpreter and proof checker, no LLM judge.
  • Best AlienCode trajectory jumps from 15.7% to 89.0% after four rounds of probing.
  • Without environment feedback, AlienCode scores stay at 0.5 to 11.0%; extra thinking alone does not help.
  • Designing your own probes beats replaying the same probes for 9 of 10 systems (median 17.4 points).
  • Under identical budgets, endpoints range from 5.7% to 79.0%; sandbox rankings barely transfer (Spearman 0.35).

ExplorationBench tests whether AI models can discover hidden rules

ExplorationBench, developed by Tencent’s Hunyuan team, Fudan University, and Tsinghua University, measures whether models can infer unfamiliar rules through active experimentation. The benchmark places ten frontier systems in two synthetic sandboxes, asks them to probe each environment, and then tests them on held-out tasks. Dedicated programs grade the code and proofs, removing LLM-judge variability.

Familiar syntax conceals alien rules

Discovery benchmarks face a validation problem. Established fields may appear in pretraining data, making contamination difficult to exclude, while genuine frontier questions often lack fast, definitive grading. The researchers address both constraints with invented worlds whose behavior is counterintuitive but fully machine-checkable.

Benchmark environments
Sandbox Hidden changes Model task Verifier
AlienCode 31 programming-language rules Infer semantics and solve coding tasks Interpreter
AlienLogic 24 inference rules Construct valid proofs or identify impossible ones Proof checker

AlienCode wraps altered behavior in familiar-looking syntax. SHATTER performs multiplication, EMIT(100) prints 127 because integer literals are XORed with 27, and PLUCK uses one-based positions even though its manual claims zero-based indexing. The language also omits an addition primitive, forcing models to compose addition from other changed operations.

AlienLogic changes the rules of natural deduction, a formal proof system in which every inference step must follow an approved rule. Some requested theorems have no valid proof under the altered system, so the correct response is to recognize the dead end and decline the task.

The benchmark contains 55 discovery targets and 140 held-out tasks across the two environments. Its interpreter executes code, while its proof checker validates each logical step.

A forked protocol seals the test set

  1. Probe the sandbox. During each of four exploration rounds, the active model submits code or proof attempts through a dedicated tool.
  2. Receive deterministic feedback. The environment returns the result without explaining the underlying rule.
  3. Fork the transcript. After each round, the harness copies the conversation into a separate evaluation branch and disables tool access there.
  4. Run closed-book evaluation. The copied model states its inferred rules and answers every held-out task three times. The harness then discards that branch, preventing test items from shaping later probes.

Each system follows three independent exploration trajectories. The benchmark reports Best@3, the highest final score among those runs, because identical budgets can produce sharply different experiments and outcomes.

Probing drives the gains

AlienCode performance under different conditions
Condition Result
Worked examples, before exploration No trajectory exceeded 15.7%
Four rounds with sandbox tools The best score reached 89.0%; seven of ten systems exceeded 60%
Additional rounds without tools Scores remained between 0.5% and 11.0%

The no-tool control indicates that additional inference time contributed little by itself. In AlienLogic, three of the ten systems lost accuracy over the corresponding rounds when they could not interact with the sandbox.

The Hindsight control tested whether models benefited from evidence alone. Researchers replayed the probes from each system’s best trajectory, giving another run the same observations without letting it choose the experiments. Nine of ten systems performed better when they selected their own probes, with a median advantage of 17.4 percentage points. Identical evidence proved less useful when the model had not generated the hypothesis and experiment that produced it.

Discovery and execution diverge

  • AlienLogic was limited by rule discovery. Giving models the rules directly produced scores of 93% to 97%, above every exploration run.
  • AlienCode exposed a rule-application problem. For seven of ten systems, the best exploration trajectory outscored the condition in which the model received the rules upfront.
  • Exploration improved later rule use. In 26 of 30 AlienCode trajectories, models scored higher when given the rules after exploring than when given them before exploration. The median gain was 14.5 percentage points.

Correctly describing a rule still failed to guarantee correct execution. When a model’s final report contained every rule needed for a task, it solved that task 73.4% of the time.

Two Kimi K3 trajectories illustrate the gap. Both correctly identified the index shift under the same budget, yet one finished at 79.0% and the other at 5.7%. The lower-scoring run repeatedly applied a top-level rule inside function bodies, where it did not hold.

One score hides volatile runs

AlienCode variation under identical budgets
System Lowest endpoint Highest endpoint
Kimi K3 5.7% 79.0%
Gemini 3.8 Flash 6.2% 79.0%

These ranges make rankings sensitive to aggregation. Best@3 rewards peak performance, while mean scores reward consistency, and switching between them can reorder systems. Rankings also transferred weakly between environments: Grok 4.6 placed fifth in AlienCode and first in AlienLogic, while Gemini 3.8 Flash tied for third in AlienCode and finished last in AlienLogic. The Spearman correlation was 0.35, indicating a limited relationship between the two rank orders.

Within each system’s best AlienCode trajectory, the largest single-round improvement contributed 47% to 92% of the total gain. Regression was also common: 6 of 30 AlienCode trajectories and 3 of 30 AlienLogic trajectories finished at least three percentage points below an earlier milestone.

One Gemini 3.8 Flash run correctly rejected an unprovable theorem in the first round. It later proposed a workaround that the checker rejected, then recorded that workaround as a valid rule in its final report.

What agent builders can apply

  • Report score distributions. Means, variance, and Best@k expose both reliability and peak capability.
  • Evaluate discovery and execution separately. Rule reports reveal what a model inferred, while held-out tasks show whether it can apply those rules in context.
  • Preserve intermediate checkpoints. Later exploration can overwrite a correct hypothesis with a weaker one.
  • Measure experiment selection. Replayed observations may overstate or understate an agent’s ability to design useful probes.
  • Use deterministic verification where possible. Interpreters and proof checkers provide reproducible task-level grading.

Submissions, source, and open questions

  • The paper, project blog, and leaderboard are available. A GitHub repository is listed as forthcoming.
  • The task set remains private to reduce contamination. The authors invite teams to submit models for evaluation under the same protocol.
  • The benchmark measures learning within a single session. Knowledge disappears after reset, and consolidation into long-term memory or model parameters remains an open research problem.
  • The two synthetic environments isolate specific discovery skills under controlled conditions. Transfer to production software, scientific research, and other open-ended settings remains unmeasured.

Across these experiments, active probing produced large gains, but success varied by trajectory, environment, and the model’s ability to apply inferred rules. Longer reasoning without new evidence added little, and accurate rule reports still left substantial execution errors. ExplorationBench gives developers a programmatically graded way to measure those capabilities separately.

Trending
  • No trending articles

Comments

avatar

Next Reads