Everyone has opinions on AI coding tools. Not enough teams publish the data.
This post covers a structured 90-day study of AI coding assistants across a 70-engineer enterprise team — frontend, backend, platform, and QA. The metrics tracked: PR velocity, review cycles, defect escape rate, time-to-first-review, and self-reported time savings. Tools used: Claude Code (primary), GitHub Copilot (one squad for comparison), Cursor (three frontend engineers).
The results were good enough to justify the investment. They were also more nuanced than the vendor benchmarks suggest — and two findings changed how to think about rollout strategy entirely.
Study Setup
A 6-month baseline was established before any tooling changes. PR metrics from GitHub, defect data from Jira, monthly velocity surveys. Then a structured rollout:







