⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

runtime-provisioner

@nealbridges⭐ 671 stars

VulnHunter sandbox-depth decision procedure. For each confirmed finding, classify its proof requirements, choose the cheapest sufficient sandbox (Docker container by default), record the runtime on the finding, and ensure Medium+ findings execute against the real dependency — no mocks at the boundary without an explicit, severity-capping EXECUTED-MOCK record. Use during the vulnhunt scan's Phase 2b after a finding is confirmed.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐ 671 on GitHubGitHub ↗

Popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

VulnHunter

[!NOTE] A maintained fork of Capital One's VulnHunter (Apache-2.0) — built to run on any agent harness, not just Claude Code. This fork's focus: harness portability, sandboxed (containerized) exploit validation, and measured-impact PoCs. See Why the changes · What this fork changes · The numbers.

From pattern-matching to provability.

VulnHunter is an open-source, agentic AI security tool that applies proactive, attacker-first analysis directly to source code.

Unlike traditional, passive SAST scanners that flag suspicious patterns and often cause false positives, VulnHunter reasons like an adversary. It identifies which defects are actually exploitable, maps prospective attack paths, and proposes targeted, evidence-backed fixes.

Modern software supply chains are deeply interconnected. A single vulnerability in a widely-used open-source component can ripple across thousands of enterprises simultaneously.

VulnHunter was developed internally at Capital One and open-sourced for the community. This fork carries that work forward — same methodology, reworked to run on any agent harness, with sandboxed (containerized) exploit validation and measured-impact PoCs as the roadmap. See What this fork changes.


Dual-use caution VulnHunter performs dual-use cybersecurity work (vulnerability discovery and exploitation). Expect guardrails: most commercially available models apply dual-use cyber safeguards, and aggressive exploitation behavior can trip rate limits or usage flags. VulnHunter's development and testing ran on open-weight, community-provided models — de-risked, abliterated, and uncensored — which are the models likely to matter for organizational use going forward. Audit only code you own or are otherwise authorized to audit.


[!IMPORTANT] Prerequisites & Model Requirements VulnHunter's methodology is built to run on open-weight, community-provided models — the de-risked, abliterated, uncensored ones organizations can actually deploy. A capable reasoning model is required; the strongest model your harness offers gives the best results, but the methodology does not depend on a specific vendor's frontier model. You supply your own model access.


What this fork changes

CapabilityUpstream (Capital One)This forkStatus
Harness portabilitySkills invoke Claude Code specifically; installer targets ~/.claude/skills; model gates hardcode Opus; harness pins claude-opus-4-8Skills are harness-portable prompt files (any harness with a skills directory + subagents); VULNHUNT_SKILLS_DIR / VULNHUNT_AGENTS_DIR / VULNHUNT_BIN_DIR / VULNHUNT_HOST_CMD / VULNHUNT_MODEL environment contract; model gates rephrased to "your harness's most capable reasoning model"Shipped
No-guess installerinstall.sh assumes ~/.claude/skillsExplicit directories, honors GROK_HOME semantics, writes the vh launcher to VULNHUNT_BIN_DIR/~/.local/bin, installs the vulnhunter-run skill + agent definition; Windows .cmd equivalents updatedShipped
vulnhunter-run operator skill— (absent)Unattended operator: clone → hunt → find-results → write/validate the scan manifest, with explicit stop rules and no improvisationShipped
Benchmark/judge hardeningFixed model + basic retryModel via environment, retry/backoff configuration, analyze_misses pipeline loss-point tracing, per-finding history trackingShipped
Harness-neutral report languageClaude-specific prose throughout the skillsHarness-neutral tool language (Agent → subagent, Claude CLI → harness session)Shipped
Sandbox-first exploit validationExploit tests may be static traces; runtime choice ad hocDocker-first runtime provisioning; the runtime recorded per finding; Medium+ severity must executeIn progress
Measured-impact PoCsPoCs are documents; impact assertedExecutable PoC + impact number in the finding (rows exposed, requests amplified, key-hours stranded)In progress

Why the changes

VulnHunter's methodology is host-agnostic by nature: it is prompt procedure, not tool binding. The upstream project grew up inside Claude Code — a coherent choice, and the right first home. But the agent-harness landscape has broadened, and a security methodology that installs into only one of them stops being an audit capability and starts being a vendor feature. This fork makes four changes, each with a reason.

1. Harness portability — your team's harness is not our harness

Every skill here is a portable prompt file with an explicit environment contract (VULNHUNT_SKILLS_DIR, VULNHUNT_AGENTS_DIR, VULNHUNT_MODEL, VULNHUNT_HOST_CMD), and the model gates now ask for your harness's most capable reasoning model instead of a specific product. Better means: the same methodology installs into whatever harness your team already runs — and becomes comparable across harnesses in benchmark runs, which is how this fork is developed.

2. A no-guess installer — "where do skills go" is a per-harness answer

The upstream installer copied skills into ~/.claude/skills unconditionally. On a machine running two harnesses — or a harness with a relocated home — that guess installs into the wrong place, silently. The fork's installer asks, or takes environment variables, and fails loudly with the exact instruction when the answer is missing. Better means: safe on multi-harness machines, correct under relocated homes, loud instead of silent when misconfigured.

3. Execution depth as a recorded decision — "provability" should not depend on model instinct

The original design already demands falsification and exploit tests. What it left open was how hard to work to actually execute them: static trace, mocked test, or a real containerized server. In one six-run benchmark against a single commit, that discretion produced anywhere from 3 to 42 findings — and opposite verdicts on the same sink, one proven against a mock, one closed by a test against a real server. This fork adds a runtime-provisioning procedure (Docker-first, recorded per finding) and a PoC discipline where impact is measured — rows leaked, ×-amplification, key-hours stranded — not narrated. Better means: a finding's validity no longer depends on which model had the instinct to stand up a container. (In progress — the build plan is on the public roadmap; ask in issues or watch the repo's Discussions.)

4. Operator ergonomics — a remediation loop compounds when it runs nightly

New in this fork: vulnhunter-run, an unattended operator that clones, hunts, locates results, and writes and validates the scan manifest with explicit stop rules. The benchmark tooling gains model configuration via environment, retry/backoff knobs, and loss-point analysis for missed findings. Better means: the difference between a tool you demo and a tool you schedule.

The numbers behind the fork

We benchmark VulnHunter against itself: six full scans of one real production Go service — same commit, five harness/model stacks. The numbers below are from those runs, and they're why this fork exists.

14×spread in confirmed findings across runs of the same commit. The process measuring itself was the first vulnerability — closing that gap is this fork's build plan.
42/42confirmed findings carried executable exploit tests — every one PASS, every one with its own PoC. No finding ships on a hunch.
55%of candidate findings eliminated or downgraded by the adversarial verification pass before reaching you. Others scan. VulnHunter litigates.
**315

// HOW IT'S BUILT

KEY FILES

skills/runtime-provisioner/SKILL.mdREADME.md

// REPO STATS

671 stars