⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

eval-skills

@dzhng⭐ 975 stars

Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus example input→expected-output pairs. Pairs with [write-skills](../write-skills/SKILL.md) (the authoring principles every fix obeys).

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐ 975 on GitHubGitHub ↗

Popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

Skills

skills.sh

AI skills for building software factories

AI skills for building software factories. My personal library of domain-agnostic agent skills, reused across every project. Small, composable, and hackable — works with any harness that supports skills: Claude Code, Codex, opencode, Cursor, duet, and 70+ others.

npx skills add dzhng/skills

Add --list to pick individual skills, or copy any skills/<category>/<name>/ folder into your harness's skills directory (e.g. .claude/skills/).

From a clone, npm run install-skills does the same without the registry:

npm run install-skills              # into ~/.agents/skills, linked from ~/.claude/skills
npm run install-skills -- ../my-app # into a repo instead of the home directory
npm run install-skills -- --only write-spec,review ../my-app
npm run list-skills                 # names and categories

.agents/skills/<name>/ holds the real files (flat, category-free, with cross-category links rewritten to match); .claude/skills/<name> is a relative symlink into it, so both harnesses read one copy. Re-running overwrites the installed copies — a .claude/skills/<name> you keep as a real directory is left alone, and a .claude/skills that is already a symlink is left as is. Add --dry-run to see the plan first.

Why

Software is moving from tasks to factories: agents that pursue a goal autonomously until the output can be trusted. The hard part isn't breaking the goal into tasks — it's breaking it into independently verifiable pieces, and knowing where the pieces even are.

These skills run that loop. Treat the unknown as fog of war: map the terrain, carve it into territories that build and verify in isolation, and recursively re-slice whatever hides more map. And re-planning doesn't stop when planning ends — the spec is a living document, updated and re-sliced mid-implementation whenever the work teaches the agent that the plan is stale. Every piece must prove itself — architecture review, code review, and visual review against a baseline — before the loop moves on. Each iteration gets less wrong, until the goal is done.

A single autonomous run — 3 days, 18 minutes pursuing one goal

Proof: one unattended Codex run pursuing a single goal for 3d 18m on top of these skills, slicing and iterating until done.

How to use

Use a chained pipeline to build a feature, a research loop to discover what works, or individual skills as needed. Every skill stands alone.

The full loop — a big feature, start to finish

The full loop — explore, spec, build unattended, review the choices

  1. Map the fog. /explore-unknowns on the idea. It interviews you quadrant by quadrant and hands you rendered options, mocks, and decision tables to react to instead of asking you to imagine. By the end you know what the feature does.

  2. Codify. /write-spec on that map. Most decisions were already made upstream, so this pass is transcription — I don't read the spec. Anything genuinely new it hits, it asks about instead of deciding.

  3. Build. Kick off the loop:

    /goal /implement-spec specs/<feature>
    

    /goal is what puts the harness in loop mode — same move in Claude Code or Codex — and the spec drives it from there. A couple of hours for a small feature, two or three days for a large one. Add whatever framing fits: on the xyz branch, or using /codex as the implementer while you stay the parent orchestrator and reviewer.

  4. Review the choices, not the diff. The run ends by consolidating specs/<feature>/choices.md — every decision the agent made where the spec was silent, ranked least-confident first. That's the review surface. Send changes back and the next pass re-audits: every time the AI writes code, you audit what it chose.

    The rest fires on its own: a /review pass at the end of every slice, /screenshot-critique and /compare-screenshots on anything visual, /close-spec when the last slice lands, and a re-slice of the plan whenever implementation proves it stale.

Budget: 30 minutes to a few hours on steps 1–2, 30 minutes to a few hours on step 4. A run that goes two days is more like 2–3 hours on each end. Your time is in the bookends; the middle is unattended.

Research — learn through fast experiments

Use Auto Research when the next decision needs experimental evidence. It starts with one fast, revealing task, tests a short batch of hypotheses, checks combinations, and expands coverage as the approach improves. New failures become the focus; earlier tasks become regression checks.

/auto-research Reduce cost per task by at least 15% relative to the saved
baseline, without reducing task success. Start with one fast development task.

If the evaluator, metric, baseline, or required improvement is unclear, the skill asks before experimenting. Passing an evaluation and meeting an improvement target are separate requirements. The output includes the best verified artifact and a parameter-effect map: what was tested, where it helps or hurts, and how changes interact. Use that evidence to inform a spec when the research is ready for implementation.

À la carte — the spontaneous path

  • A brainstorm turns out to be a feature. /explore-unknowns works at the end of a discussion as well as at the start — run it to sweep for the angles neither of you thought of, then pick the loop up at step 2.

  • Any code change that didn't come from a spec. An ad hoc fix that touched more than expected: /review first (refactor-clean → code-review → write-docs), then /audit-choices. When the diff is too big to read, the choices ledger is how you still understand what is now in your codebase.

Skills

Engineering — slice, build, verify, repeat

SkillWhat it does
explore-unknownsWalk the user through mapping a task's unknowns quadrant by quadrant — known knowns first, then interviews, reactable artifacts, and blindspot passes — ending with a complete four-quadrant map.
auto-researchOptimize through fast, progressive experiments, producing a verified candidate and a map of parameter effects and tradeoffs.
write-specBreak a large feature into independently verifiable, human-reviewable slices with API seams and playable checkpoints.
implement-specBuild an existing spec to completion, one reviewable pass at a time, delegating independent slices in parallel.
implement-spec-with-codexRun implement-spec with Codex writing the code — you orchestrate, integrate, and review every pass.
close-specArchive a shipped spec and rewrite it from a build plan into a durable rationale record that points back at the code.
refactor-cleanRefactor by moving ownership to one clean concept instead of layering compatibility sediment beside the problem.
write-testsWrite tests one tracer bullet at a time that pin real behavior — not implementation details, config values, or lucky samples.
audit-testsMap contracts to independent test proof, consolidate redundant coverage, and remove test-only machinery without losing regression protection.
audit-performanceFind hot paths

// HOW IT'S BUILT

KEY FILES

skills/authoring/eval-skills/SKILL.mdREADME.md

// REPO STATS

975 stars