# Coder Eval > Open-source framework to evaluate and benchmark AI coding agents (Claude Code, > Codex, Google Antigravity/Gemini) and Claude Code skills: sandboxed runs of real > agents against declarative YAML tasks, weighted 0.0-1.0 scoring, a > skill_triggered activation check, A/B experiments, and CI gates. Install: `uv tool install coder-eval` (Python 3.13+, Apache-2.0). - Docs: https://coder-eval.com/docs - Full docs as one file: https://coder-eval.com/llms-full.txt - Docs llms.txt (per-page index): https://coder-eval.com/docs/llms.txt - How it compares (vs SWE-bench, SkillsBench): https://coder-eval.com/docs/comparison - Blog: https://coder-eval.com/blog - GitHub: https://github.com/UiPath/coder_eval - PyPI: https://pypi.org/project/coder-eval/