Best AI Agent Skills for LLM Evals and Testing (2026)
A. Frans
Published August 15, 2026
Table of Contents
Most people testing an AI agent are testing the wrong layer. They check whether the final answer looked right, then move on. But when a Claude Code setup misbehaves, the answer is rarely where it broke. The skill never triggered. The wrong tool got called. A subagent silently returned nothing and the main loop wrote around the hole.
That gap is why eval tooling for agents looks different from eval tooling for a chatbot. You are not scoring one string. You are scoring a path.
The good news is that a Claude Code run leaves an unusually readable trail. Skills are files with frontmatter. Subagent dispatches land in JSONL. Hooks are scripts with stdout you can capture. MCP calls are JSON-RPC payloads. Almost every surface an assertion would want to read is already structured, which is more than you get from most agent frameworks.
The bad news is that the skills directory does not have a clean "eval" category, and about half of what surfaces when you search for one is a Python library or an MCP server rather than a skill. This page sorts that out.
The five stages of an agent eval loop
Before picking anything, it helps to know which stage you are missing. Teams usually have one or two of these and think they have all five.
| Stage | What it answers | Who usually owns it |
|---|---|---|
| 1. Capture | What did the agent do, step by step? | Tracing / observability |
| 2. Failure discovery | Which of those runs went wrong, and how? | Log review, annotation |
| 3. Golden set | Which cases must never regress? | Curated fixtures |
| 4. Scoring | Did this run pass, and by whose ruler? | Assertions plus LLM-as-judge |
| 5. Gate | Does this change ship? | Pre-commit, CI, or a manual stop |
Skills that gate the work (stage 5)
Verification Before Completion
Part of Jesse Vincent's superpowers collection. It forces the agent to check its work against the stated requirements before declaring a task done, which sounds obvious until you count how many agent runs end with a confident summary of something that never ran.
This is the cheapest thing on this page and the one I would install first. It is not a scoring framework. It is a stop sign, and most agent failures are failures to stop.
License: MIT. Trust: community-reviewed. Deep dive: /skills/sp-verification
Double Shot Latte
A Claude Code plugin from obra that evaluates whether Claude should keep going or hand back. Narrower than verification, and useful in a different spot: long autonomous runs where the risk is not a wrong answer but an agent that grinds through fifteen more steps after the useful work finished.
At 112 stars and unreviewed security status, treat this as a read-the-source install. It is small enough that reading it takes ten minutes.
License: MIT. Trust: community, unreviewed. Deep dive: /skills/double-shot-latte
Harness
Sets up guardrails as a meta-skill, establishing constraints an agent development workflow has to operate inside. Last pushed May 2026, license listed as unknown, unreviewed. The idea is sound and the maintenance signal is weak, so read it as a pattern to copy rather than a dependency to adopt.
Deep dive: /skills/harness
Skills that write the tests (stages 3 and 4)
Skill Creator
Anthropic's own, and the piece most people miss: skill-creator does not just scaffold a SKILL.md, it helps you write evals for the skill. Those evals check whether a given prompt causes Claude to read the skill, follow the sequence, and produce the artifacts you expected.
That is exactly the "did the right thing trigger" assertion the opening paragraph was about, and it ships from the vendor with an audited security status. If you are writing your own skills and not writing evals for them, start here.
Trust: official, audited, MIT. Deep dive: /skills/skill-creator · Our full review: Skill Creator review
Fast Agent
evalstate/fast-agent is a framework for building and evaluating agents, with strong support for MCP, ACP and skills. Apache-2.0, actively pushed, 3,891 stars on a repo where the repo is the project rather than a monorepo slice, so that number means something.
Use it when you have outgrown ad-hoc scripts and want a harness that can define an agent, run it against a fixture set, and score the result in one place.
Deep dive: /skills/fast-agent
Capture and observability (stage 1)
Phoenix
Arize's open-source observability and evaluation stack. This is a real product with 11,053 stars, an active push history, and a specific job: trace what your LLM application did, then run evaluations over those traces.
One caveat worth stating plainly. Phoenix is a Python application you run, not a folder you drop in ~/.claude/skills/. Our directory lists it because agents can drive it, but installing it means standing up a service. Budget an afternoon, not a command.
License: listed as Other, check the repo before commercial use. Deep dive: /skills/phoenix
Web Eval Agent
An MCP server from refreshdotdev that autonomously exercises a web application and reports what it found. Apache-2.0, 1,240 stars. This one covers the case the file-based evals cannot reach: your agent shipped a UI change and you want something to click through it.
Last push was February 2026, which is the oldest date in this list. Check that it still runs against your stack before wiring it into CI.
Install: claude mcp add web-eval-agent -- npx -y refreshdotdev/web-eval-agent · Deep dive: /skills/web-eval-agent
Adversarial and safety testing
ISC Bench
wuyoscar/ISC-Bench probes what its author calls internal safety collapse: turning an agent into a sensitive data leak through the conversation itself rather than through a code vulnerability. Community tier, unreviewed, license listed as Other.
If you run agents that touch customer data, this class of test belongs in your plan even if you build your own version. Prompt injection through tool output is the failure mode most teams have not written a single test for.
Deep dive: /skills/isc-bench · Related: agent skills for penetration testing
Archestra
An enterprise platform with guardrails, an MCP registry and a gateway. AGPL-3.0, which matters if you were planning to embed it in a closed product. Relevant here because a gateway is where you can enforce policy on every tool call rather than trusting each agent to behave.
Deep dive: /skills/archestra
How to read the trust column
Three things about this directory that affect how you should read every row on this page.
Star counts on monorepo-hosted skills belong to the parent repo. A skill living inside anthropics/skills inherits the star count of the whole repository, and so does every other skill in there. When a skill is a folder inside a large collection, judge it by the maintainer and the folder's own history, not by the number next to it. Only treat stars as a signal when the repo is the project, which on this page means Phoenix, Fast Agent, Web Eval Agent and Bifrost.
Install strings need checking. A large share of skill entries across the wider directory space carry a shorthand install command that does not exist as a shell command. For skills distributed through Anthropic's collection, the working path is the plugin marketplace: add anthropics/skills as a marketplace inside Claude Code, then install the bundle you want. For standalone repos, cloning into ~/.claude/skills/<name>/ is what puts the SKILL.md where Claude reads it. For anything labelled an MCP server, claude mcp add is the real command. If a command fails, that is the tell.
Unreviewed means unreviewed. Four of the entries above have unreviewed security status. A skill is instructions your agent will follow with your credentials loaded. Read the SKILL.md before installing it, the same way you would read a shell script someone handed you.
A working starting point
If you want one path rather than nine options, this is the order that gets a real gate in place fastest.
- Install Verification Before Completion. It costs nothing and stops the most common failure.
- Use Skill Creator to write evals for the two or three skills you rely on most.
- Add a golden set of ten to twenty prompts you know the correct behaviour for, stored as plain files in the repo.
- Wire those into whatever already blocks your merges. If your CI does not run them, they are documentation.
- Only then reach for Phoenix or Fast Agent, when the volume of runs makes eyeballing impossible.
The teams I have seen do this well started with five test cases and a script. The teams that stalled started by choosing a platform.
FAQ
Do I need a separate eval tool, or is skill-creator enough?
For a handful of custom skills, skill-creator plus a small golden set is enough, and it has the advantage of being official and audited. You need more once you are running agents at volume and cannot read every trace by hand. That is the point where Phoenix or Fast Agent pays for itself.
What is the difference between an eval and a unit test here?
A unit test asserts on a deterministic output. An eval scores a non-deterministic one, usually with a mix of hard assertions and a model-graded rubric. Both belong in the same CI step. Our guide to agent skills for writing unit tests covers the deterministic half.
Why do so many "eval skills" turn out to be MCP servers or Python libraries?
Because the category is young and the labels are loose. A skill is a markdown file with instructions Claude reads. An MCP server is a process exposing tools over a protocol. A library is code you import. All three show up under the same search, and they install in completely different ways. Check the install command before you plan the afternoon.
Should I use LLM-as-judge scoring?
For subjective dimensions, yes, with a caveat: validate the judge itself against human labels before you trust it. A judge prompt that agrees with you 70% of the time will happily approve a regression. Build the golden set first, then check whether your judge reproduces it.
What about testing whether the right skill triggered at all?
That is the assertion most people are missing and the one Claude Code makes easy, since skill invocation shows up in the run's structured events. Assert on the sequence, not only on the final text. A run that produced the right answer through the wrong path will break the next time the input shifts.
Share this article
⚙Related Tools
📄Related Articles
Best AI Observability and Evaluation Tools in 2026: Ship Reliable LLM Apps
12 min read
Best AI Agent Evaluation and Observability Tools in 2026: Ship Reliable AI
11 min read
Skill Creator Review: Is Anthropic's Meta-Skill Worth Installing in 2026?
8 min read
Best AI Agent Skills for QA & Test Automation in 2026
9 min read
Best AI Agent Skills for Unit Tests in 2026
9 min read
Get More AI Tool Guides
New comparisons and guides every week. Join thousands of professionals staying ahead of the AI curve.