⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

arena

@jakeschincariol⭐ 234 stars

>-

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐ 234 on GitHubGitHub ↗

Popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

The Arena skill

When Claude keeps giving you bad answers, make 100 versions of it fight to the death. One Claude Code skill. Free, MIT, no signup, no API key, nothing to connect.

You use it when you are not happy with what Claude gave you. Instead of asking again and again, /arena spins up 100 sub-agents and gives every one of them the exact same task, word for word. Each one gets a different strategy card: a way of reasoning, a workflow and a strategy, so each one attacks the task differently. Then they pair off. Each attacks the other's solution, each defends and revises its own, and a judge scores the match on a written rubric. The loser is out. 100 becomes 50, then 25, 13, 7, 4, 2, 1.

What comes back is the one solution that survived, the attacks it beat on the way, and how many rounds it took.

It never touches your files. Every sub-agent writes inside .arena/. The winning answer comes back to you, and using it is your call.

Install

Paste this into Claude:

https://github.com/Jakeschincariol/arena-skill

Install this skill, then confirm /arena works.

Or copy the folder yourself, in Claude Code:

git clone https://github.com/Jakeschincariol/arena-skill.git
cp -r arena-skill/skills/arena ~/.claude/skills/

Or as a plugin:

/plugin marketplace add Jakeschincariol/arena-skill
/plugin install arena-skill@arena-skill

Claude Code namespaces plugin skills, so installed as a plugin it shows up as /arena-skill:arena. Copy the folder instead if you want plain /arena. Project-local instead of global: copy the same folder into your repo's .claude/skills/.

It needs Claude Code, because it spawns sub-agents with the Agent tool, and Python 3.8 or newer. Nothing to pip install.

Use it

/arena
/arena --quick write the headline for our pricing page
/arena --agents 32 fix the flaky test in tests/test_api.py
/arena --seed 7 plan my launch week, I have 6 hours a day

On its own, /arena takes your last request as the task and the answer you did not like as the one to beat. Claude can also reach for it without the slash, when you say something like "that's a bad answer, make them compete". When it starts itself like that, it asks before it spends anything.

flagwhat it does
--agents NN competitors. Default 100.
--quick16 competitors. The everyday setting.
--seed SSame seed, same cards and same bracket. Default random, and recorded.
--wave WSub-agents per wave. Default 10. Only raise it if you raised Claude Code's limit.

How it works

  1. Spawn. N sub-agents, one Agent call each. Every one gets the same task text, byte for byte (a test checks it), plus one strategy card. There are 15 reasoning modes (first principles, inversion, analogy, adversarial, constraint first, worked example, Socratic, contrarian, systems thinking and more), 12 workflows (draft, critique, rewrite; outline first; test first; research, then synthesise; three drafts, pick one...) and 12 strategies (simplest thing that works, maximal rigour, user empathy first, edge cases first, speed...). That is 2,160 different cards. The dealer gives each agent its own, with no repeats.
  2. Attack. Solutions are paired, and the pairing avoids putting two agents with the same reasoning mode against each other. Each side attacks the other's solution through its own card: what is wrong, what requirement it missed, the exact input that breaks it. Up to 7 attacks, labelled FATAL, MAJOR or MINOR.
  3. Defend. Each side answers every attack it took, conceding or rebutting with evidence, then rewrites its solution to fix everything it conceded.
  4. Judge. A separate judge sub-agent reads both revised solutions, checks every attack itself, and scores both on the rubric: correctness 30, completeness against the task 25, robustness to the attacks raised 20, specificity 15, clarity 10. The judge never sees the cards. bracket.py does the arithmetic: the higher total goes through, and a solution with a verified fatal flaw cannot beat one without. The loser is out.
  5. Repeat. The survivors carry their revised solutions into the next round. An odd number means one bye, never to the same agent twice while anyone else is waiting for one.
  6. Result. One solution left. You get it, the attacks it survived, its card and the round count. If you started from an answer you rejected, a last judge compares the winner with it blind and the skill tells you the score, even when the old answer wins.

The 100-agent bracket, from python3 skills/arena/bracket.py plan:

  round  alive  matches  bye  sub-agent calls  waves
  spawn    100        -    -              100     10
      1    100       50    -              250     25
      2     50       25    -              125     13
      3     25       12  yes               60      8
      4     13        6  yes               30      5
      5      7        3  yes               15      3
      6      4        2    -               10      3
      7      2        1    -                5      3
  total                                 595     70

  alive per round: 100 -> 50 -> 25 -> 13 -> 7 -> 4 -> 2 -> 1

The whole tournament lives in one JSON file, managed by bracket.py. The main Claude session only runs the loop and never reads the hundreds of solution files: every sub-agent writes its work to disk and replies with one line, so the main context holds receipts and bookkeeping, not the work. If the conversation gets compacted halfway through, bracket.py next picks it up from the file.

What it costs

The skill is free. The tokens are yours, and 100 agents is a lot of them.

agentsroundssub-agent callswaves of 10
100, the default759570
64637949
32518728
16, --quick49116
834310

Add one call when there is a rejected answer to beat. Every call reads the task and one or two solutions, so the bill grows with the size of the task. Use --quick or --agents 16 for everyday things and save the full 100 for the answer that matters. The skill prints these numbers before it starts, and python3 skills/arena/bracket.py plan --agents N prints them any time.

The tool

bracket.py is standard-library Python. It is the reason the orchestrator never loses track.

python3 bracket.py plan --agents 100        # rounds, calls, waves. Writes nothing
python3 bracket.py init --agents 100 --seed 7 --task-file task.md [--baseline-file old.md]
python3 bracket.py next                     # what to do now, with the exact command
python3 bracket.py prompts attack           # write every sub-agent brief, list the jobs in waves
python3 bracket.py pairings                 # this round's matches, including the bye
python3 bracket.py collect                  # record the judges' verdicts
python3 bracket.py record r3-m07 a042 --reason "..."   # or record one by hand
python3 bracket.py advance                  # eliminate the losers, pair the survivors
python3 bracket.py status                   # alive and eliminated, per round
python3 bracket.py winner                   # the survivor, the attacks it survived, the rounds

The tests run a full 100-agent tournament through that command line with random winners and check it ends with exactly one survivor, plus 16, 7 and 1 agents, the dealer's guarantees, and every phase from spawn to the final check with stand-in sub-agents:

python3 -m unittest discover -s tests -v

The fine print

"100 versions of Claude" means 100 sub-agents of the model you are running. Not 100 different models. What makes them different is the card. The dealer guarantees that no two agents get the same card, that every reasoning mode, workflow and strategy is dealt as evenly as possible, and that up to 144 agents, no two agents even s

// HOW IT'S BUILT

KEY FILES

skills/arena/SKILL.mdREADME.md

// REPO STATS

234 stars