Microsoft Research Asia has open-sourced Agent Lightning v1.0, a reinforcement-learning framework built to train agents through the same harnesses they use in deployment. For teams with existing coding or general-purpose agents, it is meant to let the training system observe and optimize real model calls without requiring a second implementation of the agent’s context, tools and execution loop.
The release formalizes that approach as “Harnessed Agentic RL.” Microsoft describes Agent Lightning v1.0 as a rebuilt, roughly 3,500-line control plane focused on real-harness integration, Kubernetes rollout execution and a reproducible coding-agent training pipeline.
Training the deployed harness
Traditional agentic RL systems generally control the interaction loop: a model produces an action, the environment returns an observation, and the framework appends that observation to the token trajectory before the next model call. Microsoft says that model is difficult to apply directly to production agent harnesses, which can have their own context-management rules, tool protocols, execution logic, dependencies, subagents and summarization.
Agent Lightning instead places an OpenAI-compatible LLM proxy between the harness and the model. The agent can continue using its own orchestration code, while the proxy associates model calls with their rollouts and records prompts, responses and log probabilities for training. Microsoft says redirecting an existing harness’s model endpoint to the proxy is usually enough to connect it to the framework.
The distinction matters because the training system may see a rollout as a series of LLM request-and-response pairs rather than one continuous token sequence. Microsoft identifies four resulting problems: retokenizing text can alter token boundaries and make adjacent calls impossible to merge cleanly; a rollout split into multiple samples can be counted repeatedly in sample-level advantage calculations; sample-count loss averaging can overweight rollouts that happen to yield more samples; and unknown sample counts and lengths complicate scheduling on fixed GPU and parallelism configurations.
A small control plane with separate execution
Agent Lightning has three core components: an API Gateway, a Rollout Controller and a Customized Trainer. The gateway stores rollouts, models and events in addition to serving as the LLM proxy. The controller starts and manages agent execution, either as local processes or standard Kubernetes jobs, keeping that work separate from the trainer.
The customized trainer, built on verl, creates rollouts, waits for them to finish, collects the resulting samples and uses a sample adapter to assemble final training inputs. Microsoft positions the relatively small codebase as a way for developers to understand, modify and extend a complete harnessed agent-RL system rather than integrate a larger reimplementation of their agent.
Scheduling rollouts and using Kubernetes
The release also introduces Collocated Async RL, a scheduling approach in which rollout work and model updates share the same GPU set. Synchronous RL can leave accelerators waiting for the slowest agent in a batch; fully asynchronous setups can improve utilization but require separate GPU pools for rollout and training. In Agent Lightning’s approach, once enough rollouts have arrived, the gateway stops accepting new requests, waits for in-flight requests to complete, performs a model update and then resumes rollouts. Microsoft says that state transition is transparent to the external harness.
In its experiments, Microsoft reports about a twofold end-to-end speedup over synchronous RL while using fewer GPUs than a conventional asynchronous arrangement. The framework also runs agents as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes or local infrastructure, rather than relying on external commercial sandbox services. That can allow organizations to use existing infrastructure for rollout execution, but it also leaves them responsible for operating and provisioning the underlying environment.
Reported coding-agent result
Microsoft evaluated the pipeline using SWE-smith, mini-SWE-agent and Qwen3.5-9B. The work covered data cleaning, environment construction, reward-hacking safeguards and RL training. With about 6,000 training samples, Microsoft reports that RL training alone increased Qwen3.5-9B’s Pass@1 score on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points.
The researchers also report that combining rollout-level advantage calculation with rollout-level loss normalization produced higher validation reward and steadier policy entropy than sample-level handling in their coding-agent experiments. Those findings directly address the distortions Microsoft says emerge when a harness turns a single rollout into a varying number of samples.
The performance and speed figures are results from Microsoft Research Asia’s specified implementation and experiments; they do not establish the same gains for other models, harnesses or workloads. Source: Microsoft Research
Definition. Agent Lightning is a reinforcement-learning framework that uses an OpenAI-compatible LLM proxy to connect deployed agent harnesses with training workflows.
| Approach | Scheduling characteristic |
|---|---|
| Synchronous RL | Can leave accelerators waiting for the slowest agent in a batch. |
| Fully asynchronous RL | Can improve utilization but requires separate GPU pools for rollout and training. |
| Collocated Async RL | Shares the same GPU set for rollouts and model updates; the gateway pauses new requests for updates. |
Key takeaways
- Agent Lightning records model calls through an OpenAI-compatible LLM proxy while leaving an agent’s orchestration code in place.
- Its API Gateway, Rollout Controller and Customized Trainer separate rollout execution from training.
- Collocated Async RL shares a GPU set between rollouts and model updates.
- Microsoft reports roughly a twofold end-to-end speedup over synchronous RL in its experiments.
- In a specified coding-agent experiment, RL training raised Qwen3.5-9B SWE-bench Verified Pass@1 from 41.8% to 56.4%.
- Microsoft says reported performance results may not apply to other models, harnesses or workloads.
FAQ
What is Agent Lightning v1.0?
It is an open-source reinforcement-learning framework from Microsoft Research Asia for training agents through their deployed harnesses.
How does Agent Lightning connect an existing agent harness to training?
It places an OpenAI-compatible LLM proxy between the harness and model, associating model calls with rollouts and recording training data.
What are Agent Lightning’s core components?
The framework includes an API Gateway, a Rollout Controller and a Customized Trainer.
What result did Microsoft report for its coding-agent experiment?
Using about 6,000 training samples, Microsoft reported an increase in Qwen3.5-9B SWE-bench Verified Pass@1 from 41.8% to 56.4%.