<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Coder Eval blog</title><description>Evaluating AI coding agents and Claude Code skills on your own tasks.</description><link>https://coder-eval.com</link><language>en</language><atom:link href="https://coder-eval.com/rss.xml" rel="self" type="application/rss+xml"/><item><title>How to test Claude Code skills: from 5 manual prompts to a 200-row eval suite</title><link>https://coder-eval.com/blog/how-to-test-claude-code-skills</link><guid isPermaLink="true">https://coder-eval.com/blog/how-to-test-claude-code-skills</guid><description>Go from eyeballing five manual prompts to a sandboxed, 200-row skill eval suite with precision/recall gates you can run in CI — one step at a time.</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate><category>claude-code</category><category>skills</category><category>evals</category><category>tutorial</category></item><item><title>Introducing Coder Eval: evaluate coding agents on your tasks, not a leaderboard</title><link>https://coder-eval.com/blog/introducing-coder-eval</link><guid isPermaLink="true">https://coder-eval.com/blog/introducing-coder-eval</guid><description>How do you evaluate AI coding agents on your own tasks? Coder Eval runs a real agent in a sandbox against YAML tasks and scores what it actually produced.</description><pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate><category>announcement</category><category>agent-evals</category><category>claude-code</category><category>skills</category><category>benchmarking</category></item></channel></rss>