Microsoft's Agent Lightning Boosts Coding AI 14.6 Points Without Rebuilding Agents

Microsoft Research open-sourced a 3,500-line RL framework that trains agents using their real deployment harnesses, boosting Qwen3.5-9B by 14.6 points on SWE-bench.

·
·
·
Read5 min
TypeNews
SubtopicAgent Frameworks · Rl
  • Microsoft Research released Agent Lightning v1.0, a 3,500-line RL framework for real agent harnesses.
  • Lifted Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified with only 6,000 samples.
  • Introduces Harnessed Agentic RL: train the same harness you deploy, no rewrites required.
  • Agents connect by pointing their model endpoint at an OpenAI-compatible LLM proxy.
  • Collocated Async RL delivers roughly 2x speedup over sync RL on fewer GPUs.
  • Runs rollouts as native Kubernetes jobs instead of paid sandboxes like Modal or E2B.

Agent Lightning trains coding agents through their production harnesses

Microsoft Research Asia has released Agent Lightning v1.0, a reinforcement learning framework that trains an agent’s underlying model while preserving the agent harness used in production. Microsoft reports that its reference coding pipeline raised Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, a gain of 14.6 percentage points from about 6,000 training samples.

SWE-bench Verified evaluates whether an agent can resolve real, human-reviewed GitHub issues. Pass@1 measures the share solved on the first submitted attempt, so the reported improvement reflects more tasks completed without retries.

The rewrite tax

Most early agentic RL systems, including verl, AReaL, and slime, assumed that the trainer controlled the entire environment loop. The model produced an action, the environment returned an observation, and the framework stored the exchange as one continuous token trajectory.

Coding agents such as mini-SWE-agent, OpenHands, OpenCode, Claude Code, and Codex already have their own context management, tool protocols, execution logic, and dependencies. Rebuilding those systems inside an RL trainer takes substantial engineering work and can introduce behavioral differences between training and deployment.

Training through the API boundary

Agent Lightning places an OpenAI-compatible LLM proxy between the agent harness and the model. Developers redirect the harness’s model endpoint to the proxy, which records each request and response while the original agent continues to manage tools, context, and environment interactions. Microsoft calls this design Harnessed Agentic RL.

Comparison of traditional agentic reinforcement learning and Harnessed Agentic RL training loops
Harnessed Agentic RL keeps the production agent loop in place and observes model calls through a proxy.

The framework treats a complete agent run as a rollout and each model call within that run as a potential training sample. This separation lets the trainer work with agents that make different numbers of model calls, use different tools, or maintain context in their own formats.

Four bookkeeping traps

A harness-owned interaction loop gives the trainer less control over tokenization, sample boundaries, and execution time. Agent Lightning addresses four resulting problems:

Problem Why it matters
Retokenization Harnesses commonly store context as text, while RL requires the token IDs sampled during rollout. Applying a chat template again can change token boundaries and corrupt the training record.
Advantage calculation One rollout can produce several samples. Sample-level baselines give extra weight to runs that happen to make more model calls.
Loss normalization Averaging loss by sample also overweights rollouts with many calls. Rollout-level normalization gives each complete run comparable influence.
Scheduling The number and length of samples remain unknown until an agent finishes, while GPU capacity and parallelism settings must be managed throughout execution.

Three services, one control plane

Agent Lightning implements its control plane in roughly 3,500 lines across three components:

Component Responsibility
API Gateway Acts as the OpenAI-compatible proxy, stores rollouts, models, and events, and links every model call to its originating rollout.
Rollout Controller Starts and manages agent executions as local processes or standard Kubernetes jobs.
Customized Trainer Uses verl for RL training and converts recorded interactions into training samples through a sample adapter.

Version 1.0 can launch agents on self-managed clusters, cloud Kubernetes services, or local infrastructure. Teams can reuse existing compute without requiring hosted sandbox services such as Modal or E2B.

One GPU pool, fewer idle cycles

Agent runtimes vary because repository size, tool use, and task difficulty differ. Synchronous RL leaves training GPUs idle while a batch waits for its slowest rollout, while conventional asynchronous designs often reserve separate GPU pools for inference and training.

Agent Lightning’s Collocated Async RL shares one GPU pool between those phases. After enough rollouts arrive, the API Gateway pauses new inference requests, allows active requests to finish, and resumes traffic after the model update. The harness does not need custom logic for this transition. Microsoft reports about a 2× end-to-end speedup over synchronous RL while using fewer GPUs than a conventional asynchronous configuration.

GPU scheduling comparison for synchronous, asynchronous, and collocated asynchronous reinforcement learning
Collocated Async RL alternates rollout generation and model updates on a shared GPU pool.

Reading the 14.6-point gain

Microsoft’s reference pipeline combines SWE-smith, mini-SWE-agent, and Qwen3.5-9B. Training on about 6,000 samples increased SWE-bench Verified Pass@1 from 41.8% to 56.4%.

The accompanying ablations found that rollout-level advantage calculation paired with rollout-level loss normalization produced higher validation reward than sample-level accounting. It also kept policy entropy, a measure of how concentrated the model’s action distribution becomes, more stable during training.

These results support the framework’s accounting and scheduling choices within the tested pipeline. They do not establish the same improvement for every model, agent harness, or task; developers should evaluate those combinations independently.

Where it fits

Agent Lightning best suits teams that already have a working agent and want to fine-tune its model with RL without maintaining a second implementation inside the trainer. A practical integration follows four steps:

  1. Keep the existing agent harness and tool loop.
  2. Redirect its model endpoint to the Agent Lightning gateway.
  3. Run rollouts locally or through Kubernetes jobs.
  4. Use the sample adapter and verl-based trainer to update the model.

Teams starting with a simple ReAct loop may find direct integration with verl or a similar trainer easier. Agent Lightning becomes more useful as the deployed harness accumulates custom context handling, tools, execution rules, and infrastructure dependencies.

The framework is available on GitHub. Microsoft’s technical report describes the algorithms, architecture, and experiments in greater detail.

Trending
  • No trending articles

Comments

avatar

Next Reads