Skip to main content
Guide9 min read·Updated August 15, 2026
🧩

Best AI Agent Skills for LLM Evals and Testing (2026)

B

A. Frans

Published August 15, 2026

Agent SkillsEvalsTestingObservabilityClaude Code

Most people testing an AI agent are testing the wrong layer. They check whether the final answer looked right, then move on. But when a Claude Code setup misbehaves, the answer is rarely where it broke. The skill never triggered. The wrong tool got called. A subagent silently returned nothing and the main loop wrote around the hole.

That gap is why eval tooling for agents looks different from eval tooling for a chatbot. You are not scoring one string. You are scoring a path.

The good news is that a Claude Code run leaves an unusually readable trail. Skills are files with frontmatter. Subagent dispatches land in JSONL. Hooks are scripts with stdout you can capture. MCP calls are JSON-RPC payloads. Almost every surface an assertion would want to read is already structured, which is more than you get from most agent frameworks.

The bad news is that the skills directory does not have a clean "eval" category, and about half of what surfaces when you search for one is a Python library or an MCP server rather than a skill. This page sorts that out.

The five stages of an agent eval loop

Before picking anything, it helps to know which stage you are missing. Teams usually have one or two of these and think they have all five.

StageWhat it answersWho usually owns it
1. CaptureWhat did the agent do, step by step?Tracing / observability
2. Failure discoveryWhich of those runs went wrong, and how?Log review, annotation
3. Golden setWhich cases must never regress?Curated fixtures
4. ScoringDid this run pass, and by whose ruler?Assertions plus LLM-as-judge
5. GateDoes this change ship?Pre-commit, CI, or a manual stop
Stage 5 is the one people skip, and it is the only one that changes behaviour. An eval you look at after shipping is a report. An eval that blocks the merge is a test.

Skills that gate the work (stage 5)

Verification Before Completion

Part of Jesse Vincent's superpowers collection. It forces the agent to check its work against the stated requirements before declaring a task done, which sounds obvious until you count how many agent runs end with a confident summary of something that never ran.

This is the cheapest thing on this page and the one I would install first. It is not a scoring framework. It is a stop sign, and most agent failures are failures to stop.

License: MIT. Trust: community-reviewed. Deep dive: /skills/sp-verification

Double Shot Latte

A Claude Code plugin from obra that evaluates whether Claude should keep going or hand back. Narrower than verification, and useful in a different spot: long autonomous runs where the risk is not a wrong answer but an agent that grinds through fifteen more steps after the useful work finished.

At 112 stars and unreviewed security status, treat this as a read-the-source install. It is small enough that reading it takes ten minutes.

License: MIT. Trust: community, unreviewed. Deep dive: /skills/double-shot-latte

Harness

Sets up guardrails as a meta-skill, establishing constraints an agent development workflow has to operate inside. Last pushed May 2026, license listed as unknown, unreviewed. The idea is sound and the maintenance signal is weak, so read it as a pattern to copy rather than a dependency to adopt.

Deep dive: /skills/harness

Skills that write the tests (stages 3 and 4)

Skill Creator

Anthropic's own, and the piece most people miss: skill-creator does not just scaffold a SKILL.md, it helps you write evals for the skill. Those evals check whether a given prompt causes Claude to read the skill, follow the sequence, and produce the artifacts you expected.

That is exactly the "did the right thing trigger" assertion the opening paragraph was about, and it ships from the vendor with an audited security status. If you are writing your own skills and not writing evals for them, start here.

Trust: official, audited, MIT. Deep dive: /skills/skill-creator · Our full review: Skill Creator review

Fast Agent

evalstate/fast-agent is a framework for building and evaluating agents, with strong support for MCP, ACP and skills. Apache-2.0, actively pushed, 3,891 stars on a repo where the repo is the project rather than a monorepo slice, so that number means something.

Use it when you have outgrown ad-hoc scripts and want a harness that can define an agent, run it against a fixture set, and score the result in one place.

Deep dive: /skills/fast-agent

Capture and observability (stage 1)

Phoenix

Arize's open-source observability and evaluation stack. This is a real product with 11,053 stars, an active push history, and a specific job: trace what your LLM application did, then run evaluations over those traces.

One caveat worth stating plainly. Phoenix is a Python application you run, not a folder you drop in ~/.claude/skills/. Our directory lists it because agents can drive it, but installing it means standing up a service. Budget an afternoon, not a command.

License: listed as Other, check the repo before commercial use. Deep dive: /skills/phoenix

Web Eval Agent

An MCP server from refreshdotdev that autonomously exercises a web application and reports what it found. Apache-2.0, 1,240 stars. This one covers the case the file-based evals cannot reach: your agent shipped a UI change and you want something to click through it.

Last push was February 2026, which is the oldest date in this list. Check that it still runs against your stack before wiring it into CI.

Install: claude mcp add web-eval-agent -- npx -y refreshdotdev/web-eval-agent · Deep dive: /skills/web-eval-agent

Adversarial and safety testing

ISC Bench

wuyoscar/ISC-Bench probes what its author calls internal safety collapse: turning an agent into a sensitive data leak through the conversation itself rather than through a code vulnerability. Community tier, unreviewed, license listed as Other.

If you run agents that touch customer data, this class of test belongs in your plan even if you build your own version. Prompt injection through tool output is the failure mode most teams have not written a single test for.

Deep dive: /skills/isc-bench · Related: agent skills for penetration testing

Archestra

An enterprise platform with guardrails, an MCP registry and a gateway. AGPL-3.0, which matters if you were planning to embed it in a closed product. Relevant here because a gateway is where you can enforce policy on every tool call rather than trusting each agent to behave.

Deep dive: /skills/archestra

How to read the trust column

Three things about this directory that affect how you should read every row on this page.

Star counts on monorepo-hosted skills belong to the parent repo. A skill living inside anthropics/skills inherits the star count of the whole repository, and so does every other skill in there. When a skill is a folder inside a large collection, judge it by the maintainer and the folder's own history, not by the number next to it. Only treat stars as a signal when the repo is the project, which on this page means Phoenix, Fast Agent, Web Eval Agent and Bifrost.

Install strings need checking. A large share of skill entries across the wider directory space carry a shorthand install command that does not exist as a shell command. For skills distributed through Anthropic's collection, the working path is the plugin marketplace: add anthropics/skills as a marketplace inside Claude Code, then install the bundle you want. For standalone repos, cloning into ~/.claude/skills/<name>/ is what puts the SKILL.md where Claude reads it. For anything labelled an MCP server, claude mcp add is the real command. If a command fails, that is the tell.

Unreviewed means unreviewed. Four of the entries above have unreviewed security status. A skill is instructions your agent will follow with your credentials loaded. Read the SKILL.md before installing it, the same way you would read a shell script someone handed you.

A working starting point

If you want one path rather than nine options, this is the order that gets a real gate in place fastest.

  • Install Verification Before Completion. It costs nothing and stops the most common failure.
  • Use Skill Creator to write evals for the two or three skills you rely on most.
  • Add a golden set of ten to twenty prompts you know the correct behaviour for, stored as plain files in the repo.
  • Wire those into whatever already blocks your merges. If your CI does not run them, they are documentation.
  • Only then reach for Phoenix or Fast Agent, when the volume of runs makes eyeballing impossible.

The teams I have seen do this well started with five test cases and a script. The teams that stalled started by choosing a platform.

FAQ

Do I need a separate eval tool, or is skill-creator enough?

For a handful of custom skills, skill-creator plus a small golden set is enough, and it has the advantage of being official and audited. You need more once you are running agents at volume and cannot read every trace by hand. That is the point where Phoenix or Fast Agent pays for itself.

What is the difference between an eval and a unit test here?

A unit test asserts on a deterministic output. An eval scores a non-deterministic one, usually with a mix of hard assertions and a model-graded rubric. Both belong in the same CI step. Our guide to agent skills for writing unit tests covers the deterministic half.

Why do so many "eval skills" turn out to be MCP servers or Python libraries?

Because the category is young and the labels are loose. A skill is a markdown file with instructions Claude reads. An MCP server is a process exposing tools over a protocol. A library is code you import. All three show up under the same search, and they install in completely different ways. Check the install command before you plan the afternoon.

Should I use LLM-as-judge scoring?

For subjective dimensions, yes, with a caveat: validate the judge itself against human labels before you trust it. A judge prompt that agrees with you 70% of the time will happily approve a regression. Build the golden set first, then check whether your judge reproduces it.

What about testing whether the right skill triggered at all?

That is the assertion most people are missing and the one Claude Code makes easy, since skill invocation shows up in the run's structured events. Assert on the sequence, not only on the final text. A run that produced the right answer through the wrong path will break the next time the input shifts.

Share this article

📬

Get More AI Tool Guides

New comparisons and guides every week. Join thousands of professionals staying ahead of the AI curve.