Sakana AI's MASS Trains One Model to Run Its Own Agent Team

Sakana AI's MASS method lets a single 27B model act as its own team, judge, and teacher, boosting score-per-token by up to 1.6x on research benchmarks.

·
·
·
Read5 min
TypeNews
  • Sakana AI and UC Berkeley released MASS, a recursive self-improvement loop using one shared model for every role.
  • Two cycles on Qwen3.6-27B hit 1.2 to 1.6x score per output token on four research benchmarks.
  • Multi-agent trajectories gave a denser training signal than single-agent ones, winning with 31% fewer tokens.
  • Training only on task-solving data also improved the model's optimizer and evaluator roles.
  • Ablations show the workflow optimizer, not the evaluator, is the real bottleneck in RSI speed.
  • Paper on arXiv and code on GitHub.

MASS trains one language model as an agent team

Sakana AI and collaborators at UC Berkeley have released a paper and open-source implementation for MASS, or Multi-Agent Self-Supervision. The method assigns one language model every role in its training loop: solving tasks, revising the agent workflow, evaluating results, and fine-tuning the shared weights on preferred trajectories. Two cycles improved an open-weights 27B model on open-ended research benchmarks without external labels or a stronger teacher during training.

Why self-improvement stalls

Open-ended research tasks often lack unit tests, gold answers, or objective reward functions. As outputs grow more complex, even domain experts may struggle to assess them consistently, leaving the model to generate much of its own training signal.

Because the same weights produce, score, and learn from the work, errors can propagate through the entire loop. A flawed assumption may appear in a solution, receive a positive evaluation, and become part of the next training round. MASS addresses this risk by dividing the work among coordinated subagents and optimizing how they exchange information.

MASS turns organization into training data

MASS assigns separate calls to the same model three jobs: an executor team solves the task, an optimizer revises the team’s workflow, and an evaluator compares completed work. Each workflow is a written plan containing four elements:

  • Roles: the responsibilities assigned to each subagent.
  • Instructions: the work each subagent must perform.
  • Contracts: the artifacts, evidence, and checks each subagent must return.
  • Hops: the order of calls and the handoffs between subagents.

Training alternates between a fixed-weights workflow search and a weight update based on the strongest trajectories:

  1. Inner loop, fixed weights: MASS maintains a champion workflow for each task. It executes a candidate workflow, collects its messages and artifacts in a workspace, asks the evaluator to compare that workspace with the champion’s, and gives the optimizer the comparison so it can propose another revision.
  2. Outer loop, updated weights: MASS runs the champion workflow 18 times per task. The current model makes pairwise comparisons among the resulting workspaces, and a Bradley-Terry model converts those comparisons into relative scores. The system retains up to 16 trajectories per task and applies LoRA, a parameter-efficient fine-tuning method, to the orchestration and subagent conversations.

Before fine-tuning, MASS removes the workflow text from the orchestrator’s prompt. The model therefore trains on the coordination decisions embedded in successful conversations without copying the written plan directly.

The gains favor research tasks

The experiments use Qwen3.6-27B in the qwen-code coding-agent environment. Training covers 12 synthetic research tasks across finance, robotics, and pharmacy, with three additional synthetic tasks held out for testing. Pairwise win rate against the base model rises from 53.9% after one cycle to 69.9% after two.

External evaluation used GPT-5.5 and Claude Opus 4.8. Their judgments were unavailable to MASS during training, separating the self-generated supervision signal from the reported evaluation.

Score per output token normalizes benchmark performance by the amount of text generated, with the base model set to 1.0. After two cycles, the results show stronger transfer to research and data-analysis benchmarks than to software engineering:

Evaluation Result after two cycles
Held-out synthetic tasks 69.9% win rate against the base model
MLR-Bench, DSBench, ScienceAgentBench, and AstaBench 1.2× to 1.6× the base model’s score per output token
Terminal-Bench 2.0 0.96× the base model’s score per output token
SWE-bench Verified 0.94× the base model’s score per output token

The training suite omitted software engineering tasks, which may explain the weaker transfer to Terminal-Bench and SWE-bench Verified.

Three results clarify the mechanism

  1. Training transfers across roles

    The model trains only on task-solving trajectories, yet its workflow optimizer and workspace evaluator also improve. Agreement between the self-evaluator and the external judges rises from 73% to 93% across cycles, helping the workflow search preserve genuine gains.

  2. Team traces use tokens efficiently

    The MASS student reaches a 68.3% win rate after processing 28.8 million supervised tokens. The strongest single-agent student reaches 64.0% after 42 million tokens, consuming 13.2 million additional tokens, or about 46% more than MASS. Orchestrator traces teach decomposition and handoffs, while subagent traces teach execution within a bounded assignment.

  3. Planning quality limits the loop

    In an ablation that kept the base model as executor, assigning GPT-5.5 to both workflow optimization and evaluation produced winning workflows for all 12 tasks within five iterations. Strengthening only the evaluator yielded a smaller improvement. For systems that can afford one stronger component, the result supports testing it first as the workflow optimizer.

Contracts absorb the task logic

Across search iterations, successful workflows move task-specific information away from role descriptions and general instructions. Contracts and hops become more detailed, specifying concrete outputs, checks, and dependencies between agents.

The evolved finance workflow illustrates the pattern. Its data specialist must emit explicit leakage checks before downstream agents can access the data, turning a broad responsibility into an enforceable handoff condition.

Where the current method works

MASS fits tasks whose output is a defensible artifact rather than an answer with an automatic pass or fail signal. Candidate applications include research reports, data-science analyses, and scientific programming projects that require evidence, intermediate artifacts, and structured review.

The reported results provide weaker support for software engineering, where both public transfer scores fall below the base model. Workflow search also remains noisy: an optimizer may discover a strong plan and still fail to improve it consistently. The authors suggest a text-optimization equivalent of a learning-rate scheduler, which would reduce the size of workflow revisions as the search converges.

The subagents share weights, so their errors remain correlated despite their separate roles. A mistaken assumption can pass through execution, evaluation, and optimization as apparent progress. The authors argue for studying these homogeneous self-improvement loops while their decisions and failure modes remain practical to audit.

Code and reproduction limits

The paper describes the method and experiments, while the repository contains the reference implementation. The core loop uses the open-weights Qwen3.6-27B model and the qwen-code environment. Reproducing the reported external evaluation also requires access to GPT-5.5 and Claude Opus 4.8.

Trending
  • No trending articles

Comments

avatar

Next Reads