Blog
Evaluating AI coding agents and Claude Code skills on your own tasks.
-
How to test Claude Code skills: from 5 manual prompts to a 200-row eval suite
Go from eyeballing five manual prompts to a sandboxed, 200-row skill eval suite with precision/recall gates you can run in CI — one step at a time.
-
Introducing Coder Eval: evaluate coding agents on your tasks, not a leaderboard
How do you evaluate AI coding agents on your own tasks? Coder Eval runs a real agent in a sandbox against YAML tasks and scores what it actually produced.