⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
skillopt-sleep
Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill consolidation, or says things like 'make my agent better the more I use it', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule background self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay through a selected backend -> consolidate validated CLAUDE.md/SKILL.md behind a held-out gate.
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Train agent skills like you train neural networks — with epochs, (mini-)batchsize, learning rates, and validation gates — but without touching model weights.
📖 For installation, data preparation, training/eval commands, configuration, and framework internals, start with the versioned SkillOpt documentation. A concise rendered overview is available in the Documentation & Reproduction Guide, and longer-form engineering analysis appears on the Technical Blog. We also maintain a Changelog for released and unreleased changes.
News 🔥🔥🔥
- [2026-07-24] 📰 SkillOpt in the news. Read the official Microsoft Research feature, along with recent coverage from VentureBeat, Synced (机器之心), Flowtivity, and The Decoder.
- [2026-07-02] 🚀 SkillOpt v0.2.0 is out on PyPI! Headline feature: SkillOpt-Sleep, a nightly offline self-evolution engine (harvest → mine → replay → consolidate behind a held-out validation gate), now shipped as the
skillopt-sleepCLI. It also includes experimental multi-objective, replay, and dream-rollout controls; the main CLI keeps conservative defaults and does not expose every experiment-harness control as a flag. The release source adds integration shells for Claude Code, Codex, Copilot, and Devin, plus an OpenClaw reference adaptation; these plugin/MCP files live in the repository rather than the PyPI wheel. It also adds SearchQA split materialization, Windows robustness, and hardened JSON parsing. See the release notes for full release details and contributor acknowledgements. - [2026-06-15] 😴 SkillOpt-Sleep (preview) — a nightly offline self-evolution companion for local coding agents (Claude Code / Codex / Copilot): review past sessions, replay recurring tasks, and consolidate validated skills behind a held-out gate. See
docs/sleep/README.mdfor what it is, how to use it, and results. - [2026-06-03] 🎉 gbrain, gbrain-evals, and darwin-skill have all integrated SkillOpt.
- [2026-06-02] 🎉 SkillOpt v0.1.0 is now available on PyPI! Install with
pip install skillopt. This initial release includes the full training loop (rollout → reflect → aggregate → select → update → evaluate), multi-backend support (OpenAI / Azure / Claude / Qwen / MiniMax), six built-in benchmarks, and WebUI dashboard.
Overview
Modern agent skills are usually hand-crafted, generated one-shot by a strong LLM, or evolved through loosely controlled self-revision — none of which behaves like a deep-learning optimizer for the skill itself, and none of which reliably improves over its starting point under feedback.
SkillOpt treats the skill document as the trainable state of a frozen agent, and trains it with the discipline that makes weight-space optimization reproducible. A separate optimizer model turns scored rollouts into bounded add / delete / replace edits on a single skill document; in the default paper-style path, a candidate edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, a rejected-edit buffer, and an epoch-wise slow / meta update make skill training stable while adding zero inference-time model calls at deployment.
The deployed artifact is a compact best_skill.md (typically 300–2,000
tokens) that runs against the unchanged target model. Across six
benchmarks, seven target models, and three execution harnesses (direct
chat, Codex CLI, Claude Code CLI), SkillOpt is best or tied-best on all
52 evaluated (model, benchmark, harness) cells and on GPT-5.5 lifts the
average no-skill accuracy by +23.5 points in direct chat, +24.8 inside
the Codex agentic loop, and +19.1 inside Claude Code. Optimized skill
artifacts transfer across model scales, between Codex and Claude Code
harnesses, and to nearby benchmarks without further optimization.
For the full method, ablations, and per-cell results see the paper; for a visual walkthrough of the loop see the project page; for deeper API / backend / benchmark docs see docs/.
🎬 Demo Video
https://github.com/user-attachments/assets/eb12d3bc-371c-467f-904d-91b61f339ed7
Extensibility & WebUI
Adding a new backend
A backend = a chat / exec target (e.g. openai_chat, claude_chat,
qwen_chat, minimax_chat, copilot_chat, openai_compatible, codex_exec,
claude_code_exec, cursor_exec, copilot_exec). If a provider implements the OpenAI Chat Completions
protocol, try the built-in openai_compatible backend before adding code. See
docs/guide/new-backend.md for the full
contract. Chat backends add a skillopt/model/<name>_backend.py module;
target-only exec backends use the shared harness in codex_harness.py.
Both register through common.py, backend_config.py, and
skillopt/model/__init__.py.
Adding a new benchmark
A benchmark = a skillopt/envs/<name>/ package with an adapter, a data loader,
a scored rollout helper, a YAML config, and optionally an initial seed skill.
See
docs/guide/new-benchmark.md for the full
contract; the simplest reference is skillopt/envs/searchqa/.
WebUI
Launch the monitoring dashboard (optional):
pip install -e ".[webui]"
python -m skillopt_webui.app
| Flag | Default | Description |
|---|---|---|
| `--p |
// HOW IT'S BUILT
KEY FILES