⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
rubik-claude
Solve a complex task automatically with plan-mode-style rigor but no approval prompts. Runs a structured cycle of scramble (understand), plan, rotate (execute in verified slices), inspect (adversarial review by a fresh reviewer) and adjust, so a mid-tier model can reach the quality of a larger one. Invoke explicitly with /rubik-claude followed by the task, optionally with depth=1 to 4. Use only when the user explicitly invokes it by name (rubik-claude); never start it on your own for an ordinary request.
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
rubik-claude
Plan-mode rigor without the approval prompts. One command makes your agent plan, build in verified steps, get attacked by a fresh reviewer, and fix what holds up. It only interrupts you for a real decision.
/rubik-claude add rate limiting to the API and update the tests
Why
A single fast pass is where most agent mistakes come from:
| Without it | With rubik-claude |
|---|---|
| Starts editing before the goal is clear | Restates the goal, constraints and unknowns first |
| "Done" means "I think it works" | Success criteria are written down before work starts and checked after |
| Bugs from step 2 surface at step 6 | Each slice is verified before the next one builds on it |
| The author reviews their own work | A fresh reviewer who didn't write it tries to break it |
| Plan mode asks you to approve every step | The plan is written where you can read it, but never blocks |
The idea: spend extra passes on planning and independent review so a mid-tier model gets closer to the quality of a larger one.
Does it work?
Tested on eight tasks, once with the skill and once as a plain prompt (same model, independently re-graded): 49/50 assertions with the skill, 47/50 without.
The whole gap comes from one task. Where existing code quietly breaks an invariant (an inventory race that a plain run noticed but left unfixed), the review passes caught and fixed it. On the other seven tasks a plain run was already as correct, though the skill's reviewers still found real defects in the agent's own drafts, and the run took about 2-5x longer.
One run per cell, one model, small tasks: a signal, not a benchmark. Per-eval table, cost, the places agents didn't follow the skill, and all limits: evals/RESULTS.md. A walkthrough of one run: docs/example-run.md.
Install
One line (Claude Code, personal skills):
git clone https://github.com/gmflaubert/rubik-claude ~/.claude/skills/rubik-claude
Per project: clone into <repo>/.claude/skills/rubik-claude/ instead. For other agents that read SKILL.md, put the folder wherever that agent loads skills from.
Then run /rubik-claude <task>.
How it works
- Scramble: restate the goal, constraints and unknowns.
- Plan: written down with checkable success criteria; never waits for approval.
- Rotate: execute one slice at a time and verify each before moving on.
- Inspect: a fresh reviewer (subagent, or a separate in-context pass) tries to break the result.
- Adjust: verify each finding, fix the real ones, re-inspect what changed.
- Solved or ask: stop when criteria pass and review is clean; otherwise ask you one targeted question.
Usage
/rubik-claude <task>
/rubik-claude <task> depth=3
Without depth the skill scores the task and picks a tier:
| Tier | For | Max rotations |
|---|---|---|
| 1 Light | small, clear | 2 |
| 2 Standard | multi-step | 3 |
| 3 Deep | ambiguous or risky | 4 |
| 4 Max | high stakes | 5 |
Higher tiers add a pre-mortem, then a comparison of two candidate approaches, then a final audit. The rotation cap keeps cost bounded.
It only runs when you invoke it. The description says so, which is portable across agents. In Claude Code you can make that a hard guarantee by adding disable-model-invocation: true to the frontmatter (the portable spec doesn't allow that key, so it isn't in the repo copy).
Portability
The skill uses no tool names specific to one product. Where subagents exist, the reviewer runs in a fresh context. Where they don't, the review runs as a separate in-context pass and the report says so.
Layout
SKILL.md: the workflowreferences/tiers.md: tier scoring and budgetsreferences/review-prompts.md: reviewer, pre-mortem and approach-comparison templatesevals/evals.json: test prompts and assertions;evals/RESULTS.md: results and limits
Contributing
Issues and PRs welcome, especially failing cases where the review missed a real bug. Add the case to evals/ so it stays fixed.
License
MIT
// HOW IT'S BUILT
KEY FILES