⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

install

@pierry⭐ 4 stars

Install the harness-kit pipeline into the current project. Copies agents, slash commands, skills, and hooks into the project's .claude/, plus AGENTS.md and CLAUDE.md at the repo root. Run after adding the harness-kit plugin from the marketplace.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐ 4 on GitHubGitHub ↗

New / niche

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

harness-kit

You describe a feature. Claude Code writes the spec, the plan, the code, and the tests, and opens the PR. Nothing moves forward until it passes a check.

harness-kit demo

Start

1 min read

harness-kit is a set of Claude Code agents that carry one idea through three phases: Spec, Build, and Ship. Every step writes a markdown file, a script checks its structure, a judge scores it, and only a passing step unlocks the next. You decide twice: the direction, and the PR.

Without itWith it
The problem lives in a chat threadA written problem with a numeric success metric
"Done" is whatever the agent decidedAcceptance criteria you can check one by one
Review is reading a diff and hopingEvery document scored and retried until it passes

Use it for features worth doing right. Skip it for one-line fixes and throwaway spikes. To see it move first, pull an idea through the pipeline in your browser.

/plugin marketplace add Pierry/harness-kit
/plugin install harness-kit@harness-kit

Restart Claude Code, open your repo, run /harness-kit:install, and restart again. It asks which eval judge to use and whether to turn on graph engineering; the defaults are fine. You need python3, git, and the gh CLI. To update: /plugin update harness-kit, restart, /harness-kit:update.

Phase 1: Spec

1 min read

You type a four-line brief, or build one in the brief builder.

/golden-path

Squad: checkout
Problem: Returning guests abandon checkout when a card is declined once.
Hypothesis: If we add one-tap retry, completion rises 5 points.
Success metric: checkout completion, from 71% to 76% within 30 days

It writes the PRD (problem, customers, metrics, rollout, risks) and you approve the direction: a wrong problem caught here costs one rewrite, not a feature. Then it writes the PRP, the engineering spec, with the files to change and acceptance criteria.

With Jev, Jev scores both documents instead of Claude. With graph engineering, each criterion gets a stable id like REQ-001 and trace/{feature}.yml is created.

Phase 2: Build

1 min read

It writes the plan (files, order, risks, tests), implements it in small commits that follow your conventions and linters, and runs your tests. Each step loops until it passes:

flowchart LR
    write[agent writes] --> check{structure ok and score 8.0+?}
    check -->|no, up to 3x| write
    check -->|yes| next[next step]

The score comes from small yes/no checks, such as "every metric has a baseline", so a failure names exactly what to fix. With graph engineering, the plan records each file it expects to touch, and dev fails if it touches any other file without a written reason.

Phase 3: Ship

1 min read

It writes the PR (summary, test plan, links) and you approve it; it opens as a draft. A monitor clears the pipeline when it merges.

With graph engineering, the PR carries a table linking each requirement to its code and the test that proves it, and the next feature marks those links VALIDATED or STALE. Every approved step logs its score to the quality dashboard, so you can see quality move over time. The cockpit shows it all live in a terminal: npm i -g @pieerry/harness-kit, then hk-tui.

Options

1 min read

SetupBest forAddsCost
Plain (default)most featuresthe gated pipelinenothing extra
Jev judgea judge that is not Claude grading ClaudeJev scores, Claude takes over when unsureabout $0.04 per million tokens
Graph engineeringshared or audited coderequirement ids, trace file, scope gatefree
Graph and Jevbothbothabout $0.04 per million tokens

Switch any time: /hk:eval jev or local, and /hk:graph manifest, full, or off. full adds a local graph of your decisions and call graph (FalkorDB with Graphiti on free NVIDIA models, and Joern), with no Docker. Install with HK_STATUSLINE=1 if you want the pipeline in your status line. When installed, semble, repowise, and context7 sharpen code search; otherwise it uses grep.

You wantRun
Idea to merged PR, approving each step/golden-path
Stop only at the two decisions/pipeline:run "<idea>"
Spec only/product-manager:run
Build from a spec you have/sse:run (--local to skip the PR)
Design a hard system first/system-design:run
Resume/pipeline:continue

Learn more

1 min read

The wiki, also in Portuguese and Spanish, has Getting Started with every step in detail, how scoring works in Evals and Jev and System One, traceability in Graph Engineering and Graph Theory, and the sources in References. The method is harness engineering from Birgitta Böckeler. The three agents (product-manager, staff-software-engineer, system-architect) are mapped in AGENTS.md.

One honest limit: neither judge is checked against human ratings yet, so treat 8.0 as a signal, not a measurement. Issues and PRs welcome. MIT license.

// HOW IT'S BUILT

KEY FILES

skills/install/SKILL.mdREADME.md

// REPO STATS

4 stars