⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
proof-before-done
>-
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
proof-before-done
A reusable skill that gates "done" claims with an eight-question verification protocol. The operational implementation of incremental engineering at the agent layer.
When an AI engineering agent (Claude Code, Codex, or any agent that consumes skill files) is about to claim work is "done," this skill fires and forces the agent to answer a specific set of questions in writing before the claim propagates downstream. Each question targets a recurring failure mode of agent-generated work.
The gate fires on every done-claim — there is no opt-out. What scales is depth: routine work passes with a short written statement, while risky surfaces (auth, migrations, deploy config, cross-service contracts, money paths) get the full eight questions. An agent that has mis-assessed its work will also mis-assess whether it needs a gate, so the decision to run it is not the agent's to make.
The eight questions
- Does my count match my list? Recount from ground truth, never from memory or from what the previous step said.
- Did I test what I asserted, or did I assert it? An unverified claim is fine when labeled unverified. Silent assertion is not.
- Are my magic numbers justified at the call site? Why 24h and not 12h, one line away from the constant.
- Are my string-encoded keys schemaed? One encoder, one decoder — not eighteen callers composing the same identifier inline.
- What does my work NOT solve? Written before the "what landed" section, or you don't know the boundary of what you fixed.
- Does my verification floor match my claim? A type-check proves syntax. If the claim is about runtime behavior, downgrade the claim or raise the floor.
- Is my framing honest about what I did? Re-read the opening sentence. Does it claim more than the floor proves?
- Did the verification's inputs come from reality, or from me? ← the one most teams are missing
Question 8 is the one that catches what the other seven structurally cannot. The first seven test whether the work was executed correctly; none of them tests whether the premise was right. If you wrote the artifact and also wrote its test fixtures, a green suite proves the two agree with each other — both can encode the same wrong assumption, so the suite cannot fail. For anything whose job is checking other work (a hook, a linter, a validator, a parser, an audit script), an invented fixture is an automatic block: capture at least one real input first.
The upstream companion
proof-before-done is the downstream half of a two-gate discipline:
| Stage | Gate | Fires at | What it gates |
|---|---|---|---|
| Decide | initiate-change | let's change X | Whether the proposal is well-formed enough to be built against |
| Do | (the work itself) | — | — |
| Declare done | proof-before-done (this repo) | this is done | Whether the claim of done is well-formed enough to propagate |
Together they bracket every deliberate change: one gate at the moment you decide to change something, one at the moment you claim it's finished. Each stands alone — installing this skill does not require the other.
What's in this repo
Two files, both useful on their own, more useful together:
-
SKILL.md— the portable skill. Copy it into your project's.claude/skills/directory (or your agent's equivalent), fill in four placeholders, and the gate is installed. -
INCREMENTAL-ENGINEERING.md— the longer methodology document that explains why the skill produces the results it does. Read this if you want to understand the discipline; skip it if you just want the tool.
Why this exists
Agents reach the bar the human holds, not the bar they state. Without an explicit gate at every step of the work, agents propagate plausible-looking "done" claims that don't survive contact with production. The cost compounds: each unverified step becomes a foundation that subsequent steps build on, and the bugs you find later are exponentially more expensive than the bugs you would have caught at the gate.
This skill is the gate. It is small, deliberate, and load-bearing. It was developed inside a multi-repository production codebase (ComOS — an AI-native commerce platform) and validated when a single afternoon's worth of work, run through the discipline, compressed roughly three weeks of architectural design into four hours across four repositories.
The QC series — what the gate produces at scale
The emergent behaviors this gate creates when run without exception — a corpus that converges instead of drifting, hunts its own flaws before symptoms appear, and learns its own shape — are documented in an eight-part series written from the production system the skill governs. Start at Part 1:
- The Skill That Grew a Mind
- Why "Done" Is the Most Dangerous Word in Software
- The Cure Is Cheap Because the Cause Is Close
- Self-Healing Is Search
- Question Everything — Especially Your Own Thoughts
- The Gate Binds Its Own Builders
- The Corpus Is a Neural Network
- The Quality of Your Aim — The Honesty Ratio Is the Loss Curve
Parts 1–5 make the case for why the gate works. Parts 6–8 show what it does at scale, including the constraints that bind the gate's own authors (Part 6) and the persisted loss curve the whole discipline optimizes (Part 8). If the README's neural-network framing below reads as metaphor, Part 7 is the article that argues it is structure.
What the skill is, structurally
The skill is not a checklist. It is a small neural network for high-quality coding, operating at the agent layer:
-
Activation function — the eight questions act as a per-step gate that decides whether "done" signals propagate forward to downstream work. Below the threshold (any question fails or is silent), work is reworked. At the threshold (all pass at the depth the work warrants), downstream work can safely depend on this step.
-
Backpropagation — when the gate catches a failure, the correction signal flows backward through the layers of work that produced it. The specific assertion, the reasoning chain, the missing convention, the discipline that didn't fire — each gets a credit-assigned update. Your team's documentation is the weight matrix; failures caught by the gate update the weights.
-
Bounded iteration — question 5 (what does my work NOT solve?) forces explicit scope declarations. This is what makes the methodology converge rather than thrash. Each iteration's contribution is scoped; future iterations land against a clear seam.
That mapping isn't decorative. The skill behaves like a network: forward pass produces the work, the activation function gates the "done" signal, the loss function (your judgment) fires when work falls short, backprop credit-assigns the correction, the weights (your docs) update permanently. Over many iterations, the network learns to catch its own failures before the human has to.
What you actually get when you install
The skill is the architecture of a quality-coding network. The weights — your team's accumulated documentation, conventions, decision logs, contributor docs — are what make the network produce your team's quality, not generic quality. Installing the template gives you the architecture. The firs
// HOW IT'S BUILT
KEY FILES