Microsoft open-sources Agent Lightning v1.0, a 3,500-line agentic RL framework

Microsoft open-sources Agent Lightning v1.0, a 3,500-line agentic RL framework

Researchers at Microsoft Research Asia have introduced a training paradigm they call Harnessed Agentic RL and open-sourced a fully rebuilt Agent Lightning v1.0. The idea is that whichever agent harness is used in deployment is the harness that takes part directly in reinforcement learning during training. The motivation: most agent RL systems make developers reimplement the agent inside the training framework, which is costly and means the agent being trained is not quite the agent that gets deployed. Compared with the original Agent Lightning, v1.0 puts more emphasis on staying lightweight, on integrating with real harnesses, and on a complete, reproducible training pipeline.

The post explains the limit of the traditional approach. Early RL systems such as verl, AReaL and slime assume the training framework owns the interaction loop with the environment, so a whole rollout maps onto one continuous token trajectory. Real harnesses, including mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex, each bring their own context management, tool protocols, execution logic and dependencies. Rebuilding one for training is expensive, and the rebuilt version may behave differently from the deployed one. Agent Lightning instead places an LLM proxy between the agent and the model. The agent runs as before; you point the endpoint that used to call the model API at Agent Lightning, and the training framework can observe and record the model calls.

Because the harness, not the trainer, runs the loop, the training system sees only pairs of LLM requests and responses, and one rollout can be split into a variable number of training samples. The post names four resulting challenges. Retokenization and sample merging: harnesses keep context as text, but RL needs the token IDs sampled during the rollout, and re-running text through the chat template and tokenizer can shift token boundaries, so adjacent calls cannot always be merged. Advantage calculation: subagents, summarization and retokenization split rollouts into several samples, and computing baselines at the sample level counts rollouts with more samples repeatedly. Loss normalization: averaging by sample count gives more weight to rollouts that happen to produce more samples. Training backend scheduling: sample count and length are known only after the harness finishes, while GPU counts and parallel configurations are usually fixed.

The whole framework is about 3,500 lines of code with three core components. The API Gateway stores rollouts, models and events and serves as an OpenAI-compatible LLM proxy; it links every model call to its rollout and records prompts, responses and log probabilities. The Rollout Controller starts and manages agent execution, either as local processes or as standard Kubernetes jobs, keeping execution separate from the trainer. The Customized Trainer, built on verl, creates rollouts, waits for them, collects samples and assembles final training samples through a sample adapter. For an existing harness, pointing the model endpoint at the proxy is usually enough to connect to RL training.

v1.0 also introduces Collocated Async RL, which lets rollout and model updates share the same set of GPUs. Synchronous RL waits for the slowest agent in a batch and leaves GPUs idle; fully asynchronous RL raises utilization but needs separate GPU pools for rollout and training. Once enough rollouts are collected, the API Gateway pauses new requests, waits for in-flight ones to finish, and resumes rollout after the update, all transparent to the harness. In experiments this gave about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL.

On infrastructure, the post says other Harnessed Agentic RL frameworks often host agents on commercial sandbox services such as Modal Sandbox or E2B, where cost climbs quickly with scale. Agent Lightning v1.0 runs agents as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes or local infrastructure, which the authors say lowers the cost of large rollouts and keeps the pipeline open source and reproducible.

For evidence, the researchers built a full coding-agent pipeline on SWE-smith, mini-SWE-agent and Qwen3.5-9B, covering data cleaning, environment construction, reward-hacking safeguards and RL training. The training set holds about 6,000 samples, based on an open-sourced dataset, and needs no large-scale compute. RL training alone raised Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, an absolute gain of 14.6 percentage points. The coding experiments also back the earlier analysis of advantage calculation and loss normalization: rollout-level advantage combined with rollout-level normalization reached a higher validation reward and kept policy entropy more stable than sample-level handling.

Key facts

  • Microsoft Research Asia open-sourced Agent Lightning v1.0, rebuilt around Harnessed Agentic RL: the harness used in deployment is the one that takes part in RL training.
  • The whole framework is about 3,500 lines of code, with three components: API Gateway (an OpenAI-compatible LLM proxy), Rollout Controller and a Customized Trainer built on verl.
  • In the coding-agent example, Qwen3.5-9B went from 41.8% to 56.4% Pass@1 on SWE-bench Verified, an absolute gain of 14.6 percentage points, using about 6,000 training samples.
  • Collocated Async RL lets rollout and model updates share the same GPUs and gave about a 2x end-to-end speedup over synchronous RL in experiments, with fewer GPUs than conventional asynchronous RL.
  • Agents run as standard Kubernetes jobs rather than on commercial sandboxes such as Modal Sandbox or E2B.

Why it matters

Agent quality increasingly depends on the harness around the model, yet most RL systems train a reimplemented copy of the agent rather than the one that ships. Agent Lightning v1.0 targets that gap: the harness stays unchanged and is reached through an LLM proxy, so the trained agent is the deployed agent. The post also pairs the idea with a small codebase (about 3,500 lines) meant to be small and clear enough to understand, modify and extend.

Who it affects

Teams that train or fine-tune agents built on existing harnesses. The post lists mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex as examples of harnesses with their own context management and tool protocols. It also matters to groups that already have Kubernetes clusters, self-managed, cloud or local, and want to run rollouts at scale without paying for commercial sandbox services.

How to use it

The code is open-sourced. For an existing harness, the post says that pointing the model endpoint at the Agent Lightning proxy is usually enough to connect quickly to RL training. The Rollout Controller can run agents as local processes or as standard Kubernetes jobs, and the trainer is built on verl. A complete coding-agent example is provided, built on SWE-smith, mini-SWE-agent and Qwen3.5-9B, covering data cleaning, environment construction, reward-hacking safeguards and RL training. No license is named for the open-sourced code.

How solid is it

This is a first-party blog post from the team that built the framework, so the figures are the authors' own. The headline result is specific: Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, an absolute gain of 14.6 percentage points, from about 6,000 training samples. The post also reports that rollout-level advantage with rollout-level normalization beat sample-level handling on validation reward and policy entropy stability. The 2x speedup is stated for Collocated Async RL in experiments without naming the model, task or hardware.

Risks and caveats

The post gives no comparison of the 56.4% result with other models or other training methods. No GPU counts, training time or cost figures are given for the SWE-bench experiment. The 2x speedup comes without the model, task or hardware behind it. The post itself lists four open difficulties of training with real harnesses (retokenization and sample merging, advantage calculation, loss normalization, backend scheduling), and the reported fixes are shown on coding-agent experiments.

“whichever agent harness is used in deployment is the harness that takes part directly in reinforcement learning during training”

— Microsoft Research Asia, Agent Lightning v1.0 blog post