Do AI Coding Tools Make Teams Faster? The 2026 Evidence
By the CTO Coach TeamReviewed against primary sourcesPublished
Controlled experiments say AI coding tools make developers anywhere from 55% faster (GitHub, 2022) to 19% slower (METR, 2025), depending on the study. Large surveys say adoption is near-universal and trust is low. This page lays the primary studies side by side, with each one's design, sample and setting, so you can see why the headline numbers disagree and what a CTO should measure instead.
We report each study's design. We do not pool effect sizes, because the studies measure different tasks, developers and outcomes. Every figure below was read on the linked primary page in October 2026.
Do AI coding tools make developers faster?
The honest answer is that it depends on the task, the developer and the codebase. A 2022 GitHub experiment found a 55% speed-up on a well-defined task, METR's 2025 trial found experienced developers took 19% longer on real issues in familiar repositories, and METR's 2026 follow-up found smaller effects with serious selection caveats. Surveys report perceived gains alongside low trust.
| Study | Year | Sample | Design | Setting | Headline finding |
|---|---|---|---|---|---|
| GitHub Copilot experiment | 2022 | 95 professional developers | Randomised controlled experiment | One task: write an HTTP server in JavaScript | Copilot group 55% faster (95% interval 21% to 89%) |
| METR early-2025 study | 2025 | 16 experienced open-source developers, 246 issues | Randomised controlled trial | Their own mature repositories (over 1 million lines of code) | Tasks took 19% longer with AI; developers had predicted a 24% speed-up |
| METR design update | 2026 | 57 developers, 143 repositories, 800+ tasks | Randomised at task level | Own open-source projects, including smaller and greenfield ones | -18% task time for original developers (n=10), -4% for new developers (n=47); wide intervals |
| DORA Accelerate State of DevOps 2024 | 2024 | Survey respondents | Survey (association) | Software delivery teams | Higher AI adoption associated with an estimated 1.5% lower throughput and 7.2% lower stability |
| DORA State of AI-assisted Software Development 2025 | 2025 | Nearly 5,000 technology professionals | Survey (association) | Software delivery teams | 90% adoption; AI now linked to higher throughput |
| Stack Overflow Developer Survey 2025 | 2025 | 33,244 responses on trust | Survey | Developers worldwide | 84% use or plan to use AI tools; 46% distrust vs 33% trust accuracy |
Why do the headline numbers disagree?
They measure different things. The Copilot experiment timed one self-contained task. METR timed real issues in large repositories the developers already knew well. Surveys capture perception and association across whole teams. A lab task, a real codebase and a self-report will not give the same number, so none of them is wrong. Read each as evidence about its own setting.
Three differences explain most of the gap.
- Task type. A well-specified greenfield task favours AI assistance. Work in a mature, familiar codebase leaves less to gain.
- Developer experience. METR's participants had contributed to their repositories for years, and the tools used then were mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet.
- What is measured. Time to finish a task is not delivery performance, quality or maintainability.
What do the controlled experiments show?
Two bodies of controlled evidence exist. GitHub's 2022 experiment randomly split 95 professional developers: the Copilot group averaged 1 hour 11 minutes against 2 hours 41 minutes for the control group, and 78% finished versus 70%. METR's 2025 trial found experienced developers took 19% longer with AI, while believing afterwards it had sped them up by 20%.
The perception gap in METR's result is worth a CTO's attention. After the study, developers still believed AI had sped them up by 20%, although the measured effect was a slowdown. Self-reported productivity can diverge from measured productivity.
METR marks its 2025 results as out of date. The July 2025 post says "These results are out of date" and points to a February 2026 update. That update (METR, February 2026), titled "We are Changing our Developer Productivity Experiment Design", reports a follow-up that began in August 2025 with 57 developers across 143 repositories and more than 800 tasks. It estimates an 18% reduction in task time for the ten returning developers (interval -38% to +9%) and a 4% reduction for 47 newly recruited developers (interval -15% to +9%). Both intervals include zero.
The authors list serious caveats in the same update. Developers increasingly declined to participate because they would not work without AI, which likely biases the estimated speed-up downward. Between 30% and 50% of developers said they withheld some tasks they did not want to do without AI. Time tracking was unreliable for developers running several agents at once. The authors believe developers are likely more sped up in early 2026 than in early 2025, but call their data "very weak evidence" for the size of that change. METR also says it is changing its experiment design. Cite the 2026 update, not only the 2025 headline.
What do the large surveys show?
The surveys show high adoption, positive perceived productivity and low trust. DORA's 2025 report found 90% adoption, a median of two hours a day with AI and over 80% saying productivity improved, but only 24% trusting AI "a lot" or "a great deal". Stack Overflow's 2025 survey found 84% using or planning to use AI tools and more developers distrusting (46%) than trusting (33%) its accuracy.
| Measure | DORA 2024 | DORA 2025 | Stack Overflow 2025 |
|---|---|---|---|
| Adoption | More than 75% rely on AI for at least one daily responsibility | 90% (up 14%); median two hours a day; 65% rely heavily | 84% use or plan to use; 51% of professional developers daily |
| Perceived productivity | More than one third report moderate to extreme gains | Over 80% say productivity improved | 52% agree AI tools or agents positively affected productivity |
| Code quality | 25% more adoption associated with 3.4% higher code quality | 59% report a positive influence | 66% name "almost right, but not quite" as the top frustration |
| Trust | 39% little or no trust in AI-generated code | 24% a lot or a great deal; 30% a little or not at all | 46% distrust vs 33% trust accuracy; 3% highly trust |
These are surveys of perception and association. They cannot show that AI caused any improvement.
What is the delivery and stability trade-off?
DORA's 2024 report associated higher AI adoption with an estimated 1.5% decrease in delivery throughput and a 7.2% reduction in delivery stability, while a 25% increase in adoption was associated with 7.5% higher documentation quality, 3.4% higher code quality and 3.1% faster code review. DORA's 2025 report reversed the throughput finding, linking AI adoption to higher throughput, and gave no stability figure in its announcement.
Two cautions apply. DORA's announcement does not state the adoption increase behind the throughput and stability estimates, and association is not causation. The practical lesson is that gains in individual speed do not automatically become gains in delivery. Stack Overflow's 2025 survey adds that 45.2% of respondents say debugging AI-generated code is more time-consuming.
How much do developers trust AI-generated output?
Trust is low and, in Stack Overflow's data, the main frustration is code that is almost right. In 2025, 46% of respondents distrusted AI accuracy and 33% trusted it, with only 3% highly trusting it. DORA found 24% trusting AI a lot or a great deal. A CTO should plan for review effort, not assume AI output is ready to ship.
What should a CTO measure before and after an AI rollout?
Measure delivery performance and quality before you roll out AI tools, then again afterwards, using the same metrics. DORA's current model has five: change lead time, deployment frequency and failed deployment recovery time (throughput), plus change fail rate and deployment rework rate (instability). Add developer satisfaction and review time, and compare against a baseline.
A before-and-after checklist
- Record a baseline for the five DORA metrics for at least a full quarter before the rollout (DORA).
- Pick two or three task types where you expect gains, and track them separately from mature-codebase work.
- Track review time and the share of changes needing rework, not only time to first commit.
- Survey developers on trust and on where AI output needed correction.
- Re-measure after one and two quarters, and compare teams that adopted early against teams that did not, if you can.
- Report results to the board in business terms; see board-ready CTO communication.
For the decision of whether to build or buy AI capability, see AI build vs buy; for strategy and governance, see CTO AI strategy; for team structure, see managing AI teams; and for what changes for the leader, see engineering leadership in the AI era.
Limitations of this review
We read the primary pages for each study but not the full papers, so we report what those pages state. The studies differ in tasks, tools, populations and years, and the AI tools themselves have changed since 2022. We did not include every study. Vendor-run research such as GitHub's should be read with that in mind. The METR results are the only ones here from a research organisation independent of an AI tool vendor, and METR itself says its 2025 results are out of date. If you spot an error, tell us; our methodology page logs corrections.
Frequently asked questions
Do AI coding tools make developers faster?
It depends on the study. GitHub's 2022 experiment found the Copilot group finished one task 55% faster. METR's 2025 trial found experienced developers took 19% longer on real issues in their own repositories, and its 2026 follow-up found smaller effects with wide intervals and acknowledged selection bias. Measure the effect in your own organisation.
What did the METR study find?
METR's July 2025 randomised trial, with 16 experienced open-source developers and 246 real tasks, found that tasks took 19% longer with AI tools allowed. Developers had predicted a 24% speed-up and afterwards believed AI had sped them up by 20%. METR now labels these results out of date and points to a February 2026 update.
What do the DORA reports say about AI?
DORA's 2024 report associated higher AI adoption with an estimated 1.5% lower delivery throughput and 7.2% lower delivery stability. Its 2025 report found 90% adoption, a median of two hours a day, over 80% reporting productivity gains and a positive link to throughput, but only 24% trusting AI a lot or a great deal. Both are surveys, so they show association, not causation.
Do developers trust AI-generated code?
Mostly not yet. In Stack Overflow's 2025 survey, 46% of developers distrusted the accuracy of AI tools and 33% trusted it, and only 3% highly trusted it. DORA's 2024 report found 39% had little or no trust in AI-generated code. The top frustration in the Stack Overflow survey, at 66%, was AI solutions that are almost right but not quite.
How should a CTO measure the impact of AI coding tools?
Record a baseline for DORA's five metrics (change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate) before rollout, then compare after one and two quarters. Add review time, rework and developer trust. Self-reported speed-ups can diverge from measured results, as METR's 2025 trial showed.
Sources
- Research: quantifying GitHub Copilot's impact on developer productivity and happiness GitHub, 2022
- Measuring the impact of early-2025 AI on experienced open-source developer productivity METR, 2025
- We are changing our developer productivity experiment design METR, 2026
- Announcing the 2024 DORA report Google Cloud, 2024
- DORA report 2025: State of AI-assisted Software Development Google, 2025
- DORA metrics: the four keys (current five-metric model) DORA, 2026
- Developer Survey 2025: AI Stack Overflow, 2025
How we source and check figures: Methodology.
Cite this research
Free to cite with attribution. Copy the citation or a link-back snippet.
Get the citation