⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

release

@microsoft⭐ 18.6k stars

Prepare and publish stable Agent Lightning releases through the repository's version bump, pull-request checks, merge, tag, PyPI trusted-publishing, and versioned-documentation workflows. Use when asked to plan, cut, verify, or explain a release; treat nightly TestPyPI builds as a separate path.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐⭐ 18.6k on GitHubGitHub ↗

Very popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

Agent Lightning was completely refactored in v1.0. For legacy releases earlier than v1.0, see this branch.

⚡ News

⚡ Key Features

  • 🪶 ~3,500 lines of code: We treat simplicity as the first principle.
  • 🧩 Train with real agent harnesses: Agents interact with the model through the Agent Lightning v1.0 proxy with ZERO changes, while keeping tools, context, control flow, and environments in the loop.
  • ☸️ Native Kubernetes support: Run agents directly as Kubernetes Jobs without relying on external sandbox services.
  • 💻 Full coding agent training example: Using only 6K training samples, an end-to-end Qwen3.5-9B workflow improves SWE-bench Verified from 41.8% to 56.4%, a gain of 14.6 percentage points. We release the full pipeline, including data cleaning, reward-hacking prevention, and training scripts. Update: We release a new coding agent training example based on Qwen3.5-35B-A3B. Pure RL improves Qwen3.5-35B-A3B on SWE-bench Verified from 47.8% to 61.6% after only 1.8K training examples, a gain of 13.8 percentage points.

⚡ Installation

The following is an example installation on a CUDA 13.0 machine:

cd <this-repo>
uv sync
bash scripts/setup_verl.sh 0.8.0 cu130

See the Installation Guide for details.

⚡ Architecture

Agent Lightning v1.0 keeps the training architecture simple with three lightweight components:

  • Trainer: Runs verl and vLLM, builds training samples, and updates the policy.
  • API Gateway: Proxies model requests and captures training data.
  • Rollout Controller: Runs agents locally or as Kubernetes Jobs.

The Trainer creates rollouts, the Controller launches agents, and the Gateway turns interactions into training data, while agents continue to run with their real harnesses.

⚡ Results

We evaluate Agent Lightning v1.0 across several practical training domains, including Search R1, LLM-in-Sandbox, and Coding Agent. Pure RL delivers substantial improvements across all three domains, as shown below.

⚡ Documentation

SectionContent
InstallationBase environment and verl GPU stack
Quick StartLocal first run and end-to-end flow
BasicsComponents, rollouts, events, and trajectories
Trainer Configurationverl integration and trace aggregation
API Gateway ConfigurationGateway and model proxy settings
Controller ConfigurationLocal and Kubernetes runners
Asynchronous TrainingCollocated async collection and pause/drain

⚡ Examples

ExampleDescription
Calc-XPOC math reasoning example with AutoGen and MCP calculator tools, requiring only one GPU.
GSM8KPOC grade-school math reasoning example.
ScienceWorldInteractive science tasks in a text-based environment.
Search-R1Multi-turn retrieval and reasoning agent.
LLM-in-SandboxGeneral agent with computer and code execution tools.
Coding AgentCoding agent trained with repository tests.
Coding Agent: MoETrain Qwen3.5-35B-A3B with Megatron and R3.

⚡ Articles

⚡ Community Projects

// HOW IT'S BUILT

KEY FILES

.agents/skills/release/SKILL.mdREADME.md

// REPO STATS

18.6k stars