University of Washington Finds LLM Agent Swarms Copy Well but Explore Poorly

A new study tests whether populations of self-improving LLM agents actually benefit from watching each other, and finds copying crowds out exploration.

·
·
·
University of Washington Finds LLM Agent Swarms Copy Well but Explore PoorlyPRO
Read2 min
TypePaper
SubtopicMulti Agent · Fine Tuning · Rl
  • New paper defines recursive social improvement: self-interested LLM agents learning from each other's revised skills.
  • Three LLMs tested (Qwen3-14B, Ministral-3-14B, GPT-OSS-20B) across bandit and open-ended skill-evolution environments.
  • In bandits, classical UCB benefits from peers but LLMs earn less reward per token than solo learners.
  • In skill evolution, copying speeds up discovery but never beats independent learners at matched token cost.
  • Social exchanges concentrate populations around fewer discoveries, collapsing exploration diversity.
  • Experiments built on OpenEvolve; characterization study, no new algorithm proposed.

LLM agent populations copy well and explore poorly

A University of Washington preprint finds that individually rewarded LLM agents can copy and revise peers’ skills. Those exchanges failed to beat independent learners when both groups received the same total token budget. In controlled bandit tasks, social agents earned less reward per token; in open-ended skill evolution, peer observation sometimes accelerated progress while final performance remained level. For agent-swarm developers, the results show how communication costs and early convergence can erase the benefits of sharing.

Self-improving agents usually optimize their own prompts, tools, or reusable instructions. Multi-agent systems often coordinate agents around one shared objective. The study examines a different arrangement: each agent has its own task and reward, can inspect peers’ work, and must allocate limited compute among private search, observation, and execution.

A social-learning test with a real budget

The paper, available as a preprint, calls this process recursive social improvement. An agent develops a procedure, another agent copies it, and later agents revise the copied version. Useful discoveries can therefore propagate and become starting points for further search.

Each round forces an agent to choose among searching privately, observing a peer, and acting on its task. All three choices consume the same finite token budget. A token spent reading another agent’s logs or skill file cannot be spent generating an answer.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads