⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

basic-skill

@darkrishabh⭐ 798 stars

Summarize small CSV files and identify the highest revenue month.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐ 798 on GitHubGitHub ↗

Popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

agent-skills-eval

A test runner for Agent Skills.

Write a SKILL.md, drop in some evals, and find out — empirically — whether your skill actually makes the model better at the task.

npm version CI license: MIT node docs TypeScript

Documentation · Quickstart · SDK · agentskills.io


Why this exists

Agent Skills — the open standard from Anthropic for giving agents domain knowledge — make it easy to ship a SKILL.md and assume your agent is now better at the task. The hard part is proving it.

agent-skills-eval is the missing piece. It runs your skill against the same prompts twice — once with_skill loaded into context, once without_skill (baseline) — has a judge model grade both outputs, and gives you a side-by-side report. If the skill doesn't make a measurable difference, you'll see it. If it does, you have receipts.

It's the test framework for the Agent Skills ecosystem, separated from any specific agent runtime so it works wherever your skills do.

Quickstart

npx agent-skills-eval ./skills \
  --target gpt-4o-mini \
  --judge gpt-4o-mini \
  --baseline \
  --strict

That's it. Point it at a folder of skills, give it a target model and a judge model, and it produces a workspace with full artifacts and a static HTML report.

agent-skills-workspace/
└── iteration-1/
    ├── meta.json            # run metadata
    ├── benchmark.json       # rolled-up pass/fail per skill
    ├── eval-basic/
    │   ├── with_skill/      # output, timing, judge grading
    │   └── without_skill/   # ↑ same, with the skill stripped
    └── report/
        └── index.html       # the visual report

Open iteration-1/report/index.html and you have a real, evidence-backed answer to "is my skill working?"

What you get

with_skill vs without_skillEvery eval runs both ways so you can see the actual lift from the skill — or its absence.
Judge-graded outputsUse any chat model as a judge. Pass/fail with cited assertions, not vibes.
TypeScript SDK + CLIOne-liner CLI for CI, full SDK for custom pipelines, custom providers, and dashboards.
OpenAI-compatible by defaultWorks out of the box with OpenAI, Together, Groq, Anthropic via OpenAI-compat layers, local Llama servers — anything that speaks the OpenAI chat API.
Tool-call assertionsDeterministic checks for agents that call tools, not just generate text.
Portable artifactsJSON + JSONL all the way down. Run today, diff tomorrow. Plug into your own dashboard.
Static HTML reportsA drop-in report site you can publish anywhere — no infrastructure.
Fully spec-compliantImplements the full agentskills.io specification: SKILL.md validation, evals/evals.json, official iteration-N artifact layout, frontmatter rules.

Install

npm install agent-skills-eval

Or run directly without installing:

npx agent-skills-eval --help

How it works

The mental model is straightforward. For every eval defined in your skill:

                ┌─────────────────────────────┐
                │       same prompt           │
                └───────────────┬─────────────┘
                                │
                ┌───────────────┴─────────────┐
                ▼                             ▼
        ┌──────────────┐              ┌──────────────┐
        │ with_skill   │              │without_skill │
        │ SKILL.md in  │              │ baseline,    │
        │ context      │              │ no skill     │
        └──────┬───────┘              └──────┬───────┘
               │                             │
               ▼                             ▼
          target model                  target model
               │                             │
               ▼                             ▼
            output                        output
               │                             │
               └──────────┬──────────────────┘
                          ▼
                   ┌─────────────┐
                   │  judge      │  scores both against
                   │  model      │  the same assertions
                   └──────┬──────┘
                          ▼
                  pass / fail per side

The judge sees the eval's expected_output and assertions and grades each side independently. The --baseline flag is what enables the comparison; without it you only get the with_skill run.

YAML config

For anything beyond a quick command, drop a config file at the root of your project:

# agent-skills-eval.yaml
root: ./skills
workspace: ./agent-skills-workspace
baseline: true
target: gpt-4o-mini
judge: gpt-4o-mini
baseUrl: https://api.openai.com/v1
apiKeyEnv: OPENAI_API_KEY
include:
  - "skills/**"
exclude:
  - "**/draft-*"
evalIds:
  - "basic"
concurrency: 4
layout: iteration
strict: true
report:
  enabled: true
  title: Agent Skills Report
logging:
  format: pretty   # pretty | jsonl | silent
  verbose: false
  color: auto
targetParams:
  temperature: 0
judgeParams:
  temperature: 0
OPENAI_API_KEY=... npx agent-skills-eval --config agent-skills-eval.yaml

CLI flags always override config values.

Options reference

These are configuration-file options and CLI defaults. See --help for the available CLI flags; targetParams, judgeParams, and logging.snippetLength are config-only. Logging flags use --log-format, --log-file, --verbose, and --no-color; report flags use --report, --no-report, --report-title, and --report-output. Supplied CLI flags override matching config values.

OptionTypeDefaultWhat it does
rootstring.Directory scanned recursively for SKILL.md files. Positional CLI arg.
workspacestring./agent-skills-workspaceOutput directory for run artifacts (per-eval outputs, grading.json, benchmark.json, meta.json).
baselinebooleanfalseWhen true, runs each eval both with_skill and without_skill so the report shows the lift the skill provides. When false, only with_skill runs.
targetstringgpt-4o-miniModel under evaluation — the one the skill is meant to help.
judgestringvalue of targetModel that grades the rubric assertions. Set it to a stronger model than target for more reliable grading.
baseUrlstringOPENAI_BASE_URL env, else requiredBase URL of the OpenAI-compatible API for both target and judge.
apiKeyEnvstringOPENAI_API_KEYName of the environment variable holding the API key (the key itself is never written to config).
includestring[]all discovered skillsGlob(s) matched against each skill's pa

// HOW IT'S BUILT

KEY FILES

examples/basic-skill/SKILL.mdREADME.md

// REPO STATS

798 stars