Microsoft Research published a detailed account of Agent Lightning v1.0 on October 7, 2026, reporting that a coding-agent pipeline using Qwen3.5-9B reached 56.4% Pass@1 on SWE-bench Verified, up from 41.8%. The project repository lists the framework’s open-source release in August 2026. Its central idea is to train an agent while keeping the harness that manages the deployed agent’s interaction with its environment.

What changes when the harness owns the interaction loop

Microsoft’s Agent Lightning keeps the deployment harness in agent training

In conventional agentic reinforcement learning (RL), the training framework manages the agent’s interaction loop. Agent Lightning’s approach, called Harnessed Agentic RL, leaves that job with the deployment harness. A harness is the software that handles an agent’s environment interaction, context and control flow.

Agent Lightning places an OpenAI-compatible API proxy between the harness and the model. The proxy routes model requests and records data such as prompts, responses and log probabilities for training. If an existing harness can route its model endpoint through that proxy, it can remain in the loop rather than being rebuilt inside the training framework.

The distinction is about who manages the interaction—not whether the agent has an environment to work in. The harness continues to run the agent’s workflow; the training system uses the model-call data it captures.

The gateway, rollout controller and trainer

Agent Lightning v1.0 has three main components. The API Gateway proxies model requests and captures data. The Rollout Controller launches agent runs as local processes or Kubernetes Jobs. The Trainer handles training, integrating verl and vLLM.

This split offers a practical route for teams whose agents already use a harness: the existing setup can be retained when its model calls can be routed through the gateway. Integration still depends on that configuration; the framework does not establish compatibility with every agent without changes.

The project repository describes v1.0 as a complete refactor of earlier releases and puts its codebase at about 3,500 lines. It identifies the framework’s license as MIT.

Why this design creates training-engineering challenges

Because the harness owns the environment loop, the trainer observes sequences of model request-response pairs rather than one continuous rollout of tokens. A rollout—the agent’s run through a task—can produce a variable number of training samples. Microsoft identifies several resulting engineering challenges: retokenizing and merging samples, calculating advantages when one rollout becomes multiple samples, normalizing the training loss so rollouts that generate more samples do not receive excess weight, and scheduling variable workloads against fixed backend resources.

Those details matter because keeping a familiar agent harness changes what the training system receives and how it must organize that data. The gateway connects the two sides, but the trainer still has to make the resulting samples usable for policy updates.

Collocated Async RL shares GPUs between rollouts and updates

Agent Lightning’s Collocated Async RL approach shares GPUs between agent rollouts and model updates. When an update starts, the gateway pauses new requests and lets in-flight requests finish; rollouts resume after the update. Microsoft reports about a 2× end-to-end speedup over synchronous RL in its experiment.

That figure belongs to the reported experiment, rather than serving as a performance estimate for other models or workloads. The sequence also shows the scheduling tradeoff: the system must coordinate active agent runs with periods when the model is being updated.

Two separate SWE-bench Verified experiments

Pass@1 is the share of benchmark tasks solved on the first attempt. The two project-reported results below concern different models and training runs. Microsoft’s October 7 account describes the Qwen3.5-9B pipeline; the project repository lists the separate Qwen3.5-35B-A3B example.

Experiment and modelMetricBefore trainingAfter trainingTraining examples
Qwen3.5-9B coding-agent pipelineSWE-bench Verified Pass@141.8%56.4%About 6,000 samples
Qwen3.5-35B-A3B coding-agent exampleSWE-bench Verified47.8%61.6%1.8K examples

For the Qwen3.5-9B pipeline, Microsoft describes a workflow using SWE-smith and mini-SWE-agent, with data cleaning, environment construction, safeguards against reward hacking and RL training. The repository lists the Qwen3.5-35B-A3B result separately, with its own baseline and 1.8K examples. These are two model-specific examples on the same benchmark, not a single score for Agent Lightning as a whole.