In ~8 mins: Weco's 100-step run, seven persistence checks, a 20-field evidence packet, and one worked production verdict.

Weco ran 100 harness rewrites over eight unattended days.

Seven versions survived. The company says its private, fixed-cost evaluator rejected about nine in ten proposals.

That rejection rate exposes the engineering problem. A self-improving agent becomes consequential when a generated change is allowed to shape later runs.

The Persistence Gate turns that promotion decision into seven checks: Dependency, Activation, Evidence, Retention, Authority, Recovery, and Value.

“Weco's company-reported 100-step AIDE2 run. The loop retained seven successive versions”

What counts as recursive self-improvement?

Recursive self-improvement begins when an artifact produced at generation t changes how generation t+1 is produced, judged, selected, or executed.

A better answer remains a task result. A prompt, skill, memory rule, tool policy, or code patch closes a limited recursive loop when later runs load it and behave differently because of it.

“Recursive self-improvement claims differ by what changes. Weco's AIDE2 fits harness improvement, while its ignition comparison did not establish improved-improver evidence”

Four labels keep the claims separate: artifact improvement (one answer improves) and harness improvement (a persistent procedure improves). Improved-improver evidence means the new system produces better successors, while model improvement means its parameters or training process change.

Weco's AIDE2 report fits the harness row: an outer agent rewrote an inner agent's code while the task model stayed fixed. Weco also says its ignition comparison, where a discovered system took the outer-loop seat, was not statistically significant.

My read is narrower than "RSI has arrived." Current systems show that fixed models can sit inside loops that retain useful software changes, while the evaluator and promotion authority remain outside the editable system.

Why doesn't a higher benchmark score prove improvement?

A benchmark increase can follow a code change even when the proposed mechanism never ran.

Weco found exactly that in AIDE85. Its statistical anti-cheating layer contained a bug and had no effective contribution, although the selected harness scored better as a whole.

Self-Evolving Agent Harnesses makes the causal problem measurable. In one Terminal-Bench 2 analysis, a positive-delta rule would have credited four of ten mechanisms, including one whose activation beacon fired zero times.

Adaptive search makes the problem worse. In the same setting, a one-run rule credited a neutral mechanism about 60% of the time and reported a false gain of at least three percentage points about 25% of the time.

The candidate must also face simpler uses of the same budget. Rethinking the Evaluation of Harness Evolution for Agents gave each method five rollouts and found 91.8 pass@5 for sequential refinement with unit-test feedback, versus 86.2 for harness evolution.

The seven-check Persistence Gate

The Persistence Gate is a scoped release decision. A candidate can pass for offline replay, remain on hold for shadow use, and fail for production write access.

Alpha Signal

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves