⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

v1.2.0

recursive-decomposition

@massimodeluisa⭐ 51 stars

Decompose dense codebase-wide, multi-document, PDF, and aggregation work even when the input fits the context window, following Recursive Language Models (Zhang, Kraska, Khattab, 2025). Use when the user asks to analyse all files, a whole repo, all docs, large PDFs, or to aggregate or multi-hop across scattered sources. Skip one file, one function, a single needle, or a one-page PDF conversion. Triggers: long context, context rot, large codebase, many files, all files, big document, multi-document, PDF, aggregate, summarize everything, codebase-wide, multi-hop, recursive, sub-agents, map-reduce.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐ 51 on GitHubGitHub ↗

Growing

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README


The problem

Large codebases, dozens of documents, long reports: as the context grows, models miss details, link distant parts by guesswork and lose accuracy. The Recursive Language Models paper calls it context rot.

What it does

When a task is codebase-wide, multi-document, a pile of PDFs, or a dense aggregate, even if it fits the window, the skill makes the agent treat the input as an environment to query instead of text to swallow:

  1. Size the input before reading anything.
  2. Filter the search space with searches, not reads.
  3. Chunk what remains into batches of 5 to 10 files or natural units.
  4. Recurse at depth 1 with one sub-agent per batch, each with a self-contained brief. Sub-agents do not spawn sub-agents.
  5. Verify the merged answer on a small window against the sources.
  6. Synthesise programmatically, with file and line references.

PDFs go through anydoc first (npx -y @firecrawl/anydoc FILE -o .firecrawl/out.md), then grep the markdown. Cloud firecrawl parse is for OCR, -Q, or -S. A one-page convert job is not this skill.

Install

With the skills CLI:

npx skills add massimodeluisa/recursive-decomposition-skill

Add -g for a user-level install, -a claude-code (or another agent) to target one agent.

As a Claude Code plugin:

claude plugin marketplace add massimodeluisa/recursive-decomposition-skill
claude plugin install recursive-decomposition@recursive-decomposition-skill

Manual: copy skills/recursive-decomposition into ~/.claude/skills/ (or your agent's skills directory) and restart the agent.

Usage

  • /recursive-decomposition applies the protocol to the current task.
  • /recursive-decomposition src/ sizes that input first, then runs the protocol.

The skill also activates on its own for prompts like:

Analyze error handling patterns across this entire codebase
Find all TODO comments in the project and categorize by priority
What API endpoints are defined across all route files?
Summarize the key decisions from all meeting notes in docs/
Find security issues across all Python files

How it works

SituationApproach
One file, one function, or a single needleRead directly
Linear aggregate or list-everything, and completeness mattersDecompose
Pairwise, quadratic, or multi-hop across scattered sourcesDecompose, even under 30k tokens
10+ files or 50k+ tokensDecompose
Under 30k tokens and a localised answerRead directly

Results reported in the paper:

TaskDirect modelWith RLM
Multi-hop QA (6 to 11M tokens)70%91%
Linear aggregationbaseline+28 to 33%
Quadratic reasoningunder 0.1%58%
Context scaling2^14 tokens2^18 tokens

RLM runs were about 3x cheaper than summarisation baselines.

Eval

Fixture is a git submodule, not files copied into this repo: tccao/mortgage-doc-rag (MIT), 131 public-domain mortgage PDFs, about 63 MB. Thanks to that project for the corpus.

git submodule update --init --depth 1 skills/recursive-decomposition/evals/files/mortgage-doc-rag
bash .github/scripts/eval-skill.sh check

Pre/post on that tree, same prompt (count, 10 largest, titles from at most three files):

PathWhat ranTimeTitles from the 3 largest
With skillsize first, anydoc on 2 digital PDFs, firecrawl parse on 1 scan (anydoc exit 3)size 0.02 s; anydoc 1.64 s; OCR parse 15.45 sAPPRAISAL OF REAL PROPERTY; TILA RESPA Integrated Disclosure; Uniform Residential Appraisal Report
Without skillpdftotext on all 131 PDFs1.65 stwo titles from digital PDFs; empty on the largest file (scan, 4 bytes)

The naive path is faster and blind on scans. The skill is slower because it OCRs one file, and that is the file pdftotext cannot read. Details: evals/README.md.

Repository structure

recursive-decomposition-skill/
├── .claude-plugin/          plugin.json, marketplace.json (the repo is the plugin)
├── .github/                 bash validator and CI workflow
├── skills/recursive-decomposition/
│   ├── SKILL.md             protocol, rules, patterns
│   ├── references/          rlm-strategies, cost-analysis, codebase-analysis, document-aggregation
│   └── evals/               trigger queries, mortgage PDF submodule, bash scorer
├── assets/                  social preview, logo (light and dark)
├── AGENTS.md · CONVENTIONS.md · CONTRIBUTING.md · CHANGELOG.md
└── LICENSE

Acknowledgments

This skill is based on the Recursive Language Models paper. Thanks to the authors:

Recursive Language Models, Alex L. Zhang, Tim Kraska, Omar Khattab, arXiv:2512.24601, December 2025. Abstract · PDF

Eval corpus: tccao/mortgage-doc-rag (MIT), public-domain mortgage PDFs. Thanks to that project.

PDF conversion: Firecrawl anydoc and firecrawl-parse. Thanks to Firecrawl.

This skill is an independent project and is not affiliated with the paper authors, MIT, tccao, or Firecrawl.

Author

Massimo De Luisa: [massimo.deluisa.bio](https://massi

// HOW IT'S BUILT

KEY FILES

skills/recursive-decomposition/SKILL.mdREADME.md

// REPO STATS

51 stars