⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
claude-md-doctor
Give this repo's CLAUDE.md / AGENTS.md a checkup — size vitals vs official guidance, dead references, dead commands, stale claims — then backtest every rule against the repo's own Claude Code session history to see which rules were actually followed, ignored, or never used, and produce a doctor-style HTML report with evidence-cited prescriptions. Use when asked to check, diagnose, audit, review, improve, optimize, lint, grade, fix, clean up, shorten, or "doctor" CLAUDE.md, AGENTS.md, or agent instruction/memory files, or to find out whether CLAUDE.md rules actually work. Also use when a repo has NO memory file and the user wants one — "write/generate/suggest a CLAUDE.md (or hooks) from my sessions" — the skill mines the repo's real session history and drafts a proposed file with receipts.
Use with your AI agent
Open your project in any AI assistant that can read your files. Works with ChatGPT, Claude, Claude Code, Codex, Cursor, Hermes Agent, OpenClaw, Grok Bot, and more.
Download SKILL.mdYour agent needs access to this page’s linked instructions and your project files. Copying does not install or execute anything.
// RATINGS
// README
Linters check the file. Analytics grade your sessions. The doctor cross-examines one against the other — and cites receipts.
Quickstart
As a Claude Code plugin (recommended):
/plugin marketplace add agent-clinic/claude-md-doctor
then install claude-md-doctor from the /plugin menu. Or via the
skills.sh CLI:
npx skills add agent-clinic/claude-md-doctor
Or bare: copy skills/claude-md-doctor/ into ~/.claude/skills/.
Then, in any repo, just ask — "give my CLAUDE.md a checkup" — or invoke
directly: /claude-md-doctor:claude-md-doctor (bare install:
/claude-md-doctor). The report lands in .claude-md-doctor/report.html
plus machine-readable report.json.
Requires Python 3.9+ (standard library only). Everything runs locally; nothing leaves your machine.
What the exam covers
- Vitals — effective size vs the official guidance ("target under 200
lines per CLAUDE.md file" — Claude Code memory docs),
estimated token cost per session, structure, and pathology markers: stock
/initboilerplate never pruned, emphasis saturation, changelog accretion. - Records check — does everything the file points at exist? Dead file
paths, paths from a teammate's machine,
@importsthat don't resolve,pnpm/makecommands with no matching script,.claude/rules/scopes that match zero files. - Checkable claims — countable assertions ("2,100 tests across 180 files", "9 UI components; no dialog") verified against the repo. Inlined numbers rot; the doctor catches them.
- The report — a single self-contained HTML page: chart grade, chief complaint, per-finding evidence, the session-adherence History table, and concrete prescriptions, each footnoted with the official doc or study behind it (the evidence base lives in docs/RESEARCH.md).
It understands the real memory surface: CLAUDE.md, .claude/CLAUDE.md,
CLAUDE.local.md, nested files, .claude/rules/*.md (with paths: scopes),
@imports (depth 4, backtick-aware), claudeMdExcludes, ancestor
directories — and it treats the pointer-to-AGENTS.md pattern as healthy,
examining the target, while flagging the broken variant (pointer text without
@, which Claude Code never actually loads). A repo with an AGENTS.md but no
CLAUDE.md at all gets the doctor's simplest prescription: the official
one-line pointer, so Claude Code stops loading nothing.
The backtest — check if CLAUDE.md actually works in your sessions
Your own Claude Code session transcripts (~/.claude/projects/…) already
record whether past sessions actually followed each rule in your CLAUDE.md.
The doctor decomposes the file into rules and replays them against that
history — per rule, a verdict with receipts:
| Rule | Opportunities | Compliance | Verdict |
|---|---|---|---|
| Never import the legacy API types | 12 | 100% | healthy |
Run verify before you finish | 2 | 0% | ignored |
| Never hardcode a colour | 0 | — | inert |
Behind every number: matched excerpts, and for finish-ordering rules a session-timeline strip showing exactly what ran after the last edit. Two ideas drive the verdicts (full taxonomy). Every rule gets an enforcement class — the cheapest reliable detector:
| Class | Detector | Binds |
|---|---|---|
hook | gate over tool calls (commands, edits, orderings) — prevents | the agent |
linter/test | static analysis over the code itself | every agent and every human |
judge | LLM audit, post-hoc, with a stated reliability ceiling | audit only |
~70% of real-world directives land in the first two — laws waiting to be passed. And every violation is triaged by cause, because the cause picks the medicine:
| Cause | What happened | Medicine |
|---|---|---|
| defiance-proven | the agent echoed the rule, then broke it | block-mode gate — the reminder already lost |
| defiance | violated in fresh context | warn-hook, then block |
| dilution | drowned late in a heavy session | slim the file, move the rule to point-of-use |
| absence | non-root rule lost to compaction | re-inject; never block |
The arming ladder (reminder → warn → block) is set per rule from its own
violation forensics, and review-then-arm hook proposals are written to
the exam folder — nothing is ever installed automatically. Every checkup
also emits a share-safe card (grade, hearts, doctor's note — aggregates
only, never a string from your repo) and a claude-md-health.svg badge
for your README. Matcher fires are
sample-verified before they count, because matchers have bugs; unverified
results are banner-labeled provisional. Research shows agents silently skip
mandated steps while outputs still pass checks; only behavioral evidence
catches that — and it's free, sitting in your transcript history.
Run it on your own repo: the rules you'd bet on being followed are rarely the ones that are.
No CLAUDE.md? The doctor writes your chart
Most repos have no memory file at all (18 of the 20 on our own machine). But their session transcripts already contain the unwritten rulebook, and the same engine that backtests rules can run in reverse — mine the history, then validate the checkable candidates against it:
| Signal | Example | Becomes |
|---|---|---|
| repeated corrections | "no, use pnpm not npm" typed in 3 sessions | a rule |
| failed → fixed pairs | npm test fails, pnpm test works, again | a rule (often a hook) |
| re-discovery | agent reads package.json at every session start | a fact, stated once |
| permission denials | you rejected git push twice | a "never" rule |
| repeated preambles | the same context paragraph pasted each session | a fact |
Grouped signals survive only with recurrence (≥2 sessions or ≥3
occurrences; re-discovery needs 3 distinct sessions) and carry recency
flags — a preference the repo moved past is marked stale for the judge
pass to decline. Corrections reach the judge ungated (wording varies too
much to group), deduped and capped, and are judged hardest. Each accepted
mechanically-checkable rule is then replayed through the backtest for
precise counts. The result is PROPOSED-CLAUDE.md: a lean
draft where every line carries its receipt as an HTML comment (stripped at
load, so it costs the adopter nothing), hook-class rules arrive as
review-then-arm guard proposals ("born mechanized"), and the report shows
the re-discovery tax your sessions have been paying. The draft is held to
the same 200-line vitals this tool grades everyone else on — the generator
refuses to prescribe the disease it diagnoses. Nothing is installed and no
CLAUDE.md is written for you — the draft lands in the exam folder
(.claude-md-doctor/), and adoption is your move. Repos that do have a CLAUDE.md get the
same mining as a gap analysis: rules you keep dictating by hand that
the file never says.
FAQ
Why does Claude ignore my CLAUDE.md? Usually
// HOW IT'S BUILT
KEY FILES