⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
eval-genius
>-
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
// README
Everyone says "you need evals." Almost nobody says when.
Evals are suddenly everywhere. Every AI talk, every launch post, every hiring thread says you need them. Then you sit down to actually do it, and the questions start.
Does a prompt tweak really need an eval, or is that overkill? At which point in the build does the first one go in? What is an "eval harness", concretely, beyond a folder named evals/? Which metric, how many examples, does a judge model even count? And when a number finally comes out, is 34 out of 40 good? Is a 3-point gain real, or noise? Or did the run just crash quietly and report a pass?
Answer these by feel and you get exactly what most teams have: a benchmark nobody trusts, a gate that has been green for a month because it grades nothing, and a number in the README that would not survive one sharp question.
Meet Eval Genius
Eval Genius is that missing judgment, packaged as a skill your AI agent runs with you. It thinks like a measurement engineer: decide what "better" means before you look, push every check you can down to plain code, treat a crash as unmeasured rather than a pass, and trust the number last.
It is not a course you have to read first. You describe where you are, in plain words, and it takes the next step, whether that step is "you don't need one yet" or "here is the gate, and here is why this run cannot be trusted."
There are two doors in, and neither asks for eval vocabulary. If you can already say what "good" looks like, it works top-down from that promise. If all you have is "the outputs are sometimes wrong and I don't know what to measure," it works bottom-up instead: it reads your real bad outputs with you, names the error categories, and turns each one into something measurable. Same discipline, entered from wherever you actually stand.
It is tool-agnostic and dependency-free: a SKILL.md plus a few standard-library Python scripts. It runs in Claude Code, Codex, and Droid, or any agent that loads skills, or from your terminal on its own.
See it in 20 seconds
A bill-splitting agent sounds sure, and is quietly wrong. Eval Genius makes it earn the word "ready."
What you can ask it
Real questions, answered from wherever you actually are:
- "Do I need evals for my chatbot, or is that overkill right now?"
- "Where does an eval even go in my build?"
- "My summaries are sometimes wrong and I don't even know what to measure."
- "Is my skill's description actually firing on the right prompts?"
- "I got 34 out of 40, is that good?"
- "Is this 3-point gain real, or noise?"
- "We swapped the agent's scaffold. Is it better now?"
- "Can my agent be tricked into leaking a secret or calling a tool it shouldn't?"
- "Can we put a number in the launch post?"
- "Calibrate my LLM judge against some human labels."
No setup ritual, no vocabulary you have to learn first. Describe the situation, get the next move.
What it does for you
It walks the whole path, and meets you at any point on it, including the start:
- Decides whether you need an eval at all, and which kind belongs at your stage, from first prototype to production.
- Starts from your failures when there is no promise yet. Real bad outputs get read by hand, clustered into named error categories, and each category becomes the thing to measure.
- Picks the eval: what to measure, which grader (code first, a judge only where no assertion works), which metric, how many examples, adopt a public benchmark or build your own.
- Builds and gates it: fixture, runner, scorer, reporter, a bar written before the run, and a CI gate that ends in PASS, FAIL, or CANNOT-MEASURE and only compares two runs when they truly measured the same thing, the same way.
- Tests whether it holds up under attack: adversarial cases, prompt injection, unsafe tool calls, data exfiltration, graded by code on what actually happened and reported per attack type, never hidden inside one "safety score."
- Reads the result with you: against the bar you wrote, with noise bounds, per-item diffs, and a harness-bug check before any surprising number is believed.
- Writes it up honestly, with caveats, tiers, and the comparison rule stated out loud.
- Refuses the shortcuts that produce pretty lies: bars moved after the fact, blended scores, run-until-green, and judges nobody calibrated.
Working evals with Jev
Some of what an eval checks is a plain typed decision: real defect or not, which failure class, positive or negative sentiment, does this match the brand voice. A full reasoning model is overkill for those, and grading them by hand at volume is the real waste. Eval Genius can route exactly that residue to *Jev
// HOW IT'S BUILT
KEY FILES