How AI Agents Are Learning to Rewrite Their Own Stack

Self-Harness reports relative gains of up to 132%. EnvHarness adds up to 9.0 points and 9.8% fewer steps

·
·
·
How AI Agents Are Learning to Rewrite Their Own Stack
AuthorBy Ben Dickson
Read3 min
  • Self-Harness improved three model families on Terminal-Bench-2.0, SWE-bench Verified, and AppWorld, with relative gains reaching 132%.
  • That 132% is Qwen3.5-35B-A3B on AppWorld, from 22.5% to 52.2% overall, and all nine model-benchmark runs improved on both splits.
  • GLM-5 on AppWorld had the largest absolute gain, from 44.4% to 85.0%.
  • Across five benchmarks, EnvHarness improved held-out tasks by up to nine points while using about 9.8% fewer interaction steps.
  • The practical rule starts at the narrowest of five surfaces: skills, the harness, co-evolution, the environment, or orchestration.

Previously, developers had to optimize AI agents by hand-tuning prompts, tools, memory, and control flow to squeeze more performance.

Now researchers are turning those same components into optimization targets. Skills can rewrite themselves. Harnesses can evolve from execution traces. Models can train on the behavior their own harnesses discover.

The result is a new design question for developers: which part of the agent stack should remain fixed, and which part should learn?

Today we map the emerging self-evolving agent stack.

How AI agents are learning to optimize their own stack

Researchers are enabling AI agents to self-improve by rewriting the software and the models that power them. The goal is to remove the bottleneck of manually tuning agents with an automated loop driven by execution feedback.

The main target is the agent harness, the layer around a model that includes prompts, tools, memory, context management and control flow. These components determine what the model sees, what actions it can take and how it handles failures.

But the emerging self-improving agentic frameworks also work on fine-tuning the backbone model along with the agent, reshape training environments, and redesign how multiple agents work together.

The emerging approaches are not stages in a hierarchy, and they often overlap. But most follow a similar pattern: run tasks, collect execution traces, find recurring weaknesses, modify part of the stack, then test whether the new version is better.

Skills: Make the agent's instructions trainable

Skills are one of the easiest places to start. They are usually text files that package task-specific instructions, workflows and heuristics outside the model weights. Microsoft's SkillOpt turns the skill document itself into the object being optimized.

SkillOpt evaluates the agent on a set of tasks. It then analyzes the scored trajectories and proposes small additions, deletions or replacements to the skill file. It then tests the candidate on a held-out validation set, which contains examples that were not used to generate the edit. A suggested modification is only accepted if it improves the validation score.

Google Research's WikiSkill addresses another problem: the information needed to optimize skill files is scattered across different components. WikiSkill creates a wiki that gathers the agent's experience (both successes and failures) into a structured format.

The wiki sits as an intermediate layer between raw execution traces and the executable skills used by the agent. WikiSkills uses this information to iterate over skills and avoid rediscovering the same problems and retrying rejected fixes.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves
Trending
  • No trending articles

Comments

avatar

Next Reads