Productivity benchmarks

Verified 4 September 2026 against the cited pages.

Contents

Read each supported claim and caveat alongside the finding. Expand the method details to check the sample and measurement. A 55% task speed-up on one synthetic task and a 19% slowdown on real issues are both real results; neither is a QA budget number.

Controlled and independent studies

METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, 10 Jul 2025 (arXiv 2507.09089)

Finding: AI-allowed issues took 19% longer (CI +2% to +39%); developers expected 24% faster and afterwards believed 20% faster

Supports: Task efficiency (negative)

Caveats: Elite developers on repositories they know well; small sample; authors note possible learning effects after several hundred hours

Sample and method details
  • Sample and method: 16 experienced maintainers, 246 real issues, randomized AI-allowed vs not; Cursor Pro with Claude 3.5/3.7 Sonnet
  • Metric and denominator: Time per issue
  • Measured or self-reported; sponsor: Measured (screen recording); independent non-profit

METR, “We are changing our developer productivity experiment design”, 24 Feb 2026

Finding: Completion time -18% (CI -38% to +9%) original cohort; -4% (CI -15% to +9%) new cohort; negative values mean faster completion

Supports: Inconclusive; speedup point estimates with substantial selection bias

Caveats: 30-50% of developers reported withholding some tasks; recruitment and task selection make the effect estimate unreliable; both confidence intervals include no effect; study design being revised

Sample and method details
  • Sample and method: 57 developers, 800+ tasks, 143 repositories, late-2025 tools
  • Metric and denominator: Time per task
  • Measured or self-reported; sponsor: Measured; independent

METR, self-reported impact survey, 11 May 2026

Finding: Median self-reported uplift about 3x

Supports: Perception only

Caveats: Authors: their own RCT participants over-estimated by 40 points

Sample and method details
  • Sample and method: 349 technical workers, survey
  • Metric and denominator: Self-estimated speed-up
  • Measured or self-reported; sponsor: Self-reported; independent

Peng, Kalliamvakou, Cihon, Demirer, “The Impact of AI on Developer Productivity: Evidence from GitHub Copilot”, 13 Feb 2023

Finding: 55.8% faster (71.2 vs 160.9 minutes; CI 21% to 89%)

Supports: Task efficiency

Caveats: Single synthetic task; code quality not assessed

Sample and method details
  • Sample and method: 95 freelancers, one JavaScript HTTP-server task, randomized
  • Metric and denominator: Task completion time
  • Measured or self-reported; sponsor: Measured; vendor-affiliated (Microsoft Research, GitHub, MIT)

Cui et al., “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments”, SSRN Sep 2024; Management Science Feb 2026

Finding: +26.1% completed tasks (standard error 10.3); larger gains for less experienced developers

Supports: Output volume

Caveats: Wide interval; no quality, throughput or cost measure

Sample and method details
  • Sample and method: 4,867 developers at three firms, randomized Copilot access
  • Metric and denominator: Completed tasks (pull requests) per developer
  • Measured or self-reported; sponsor: Measured telemetry; Microsoft-affiliated

GitHub and Accenture enterprise study, 13 May 2024

Finding: +8.7% PRs, +15% merge rate, +84% successful builds

Supports: Output volume

Caveats: No confidence intervals or n published

Sample and method details
  • Sample and method: Randomized, sample size not disclosed
  • Metric and denominator: PRs, merge rate, build success
  • Measured or self-reported; sponsor: Telemetry plus survey; vendor

GitHub code-quality RCT, 18 Nov 2024

Finding: 53% more likely to pass all unit tests; readability, reliability, maintainability +2 to +4%

Supports: Quality (small effect)

Caveats: Artificial task

Sample and method details
  • Sample and method: 202 developers, blinded review by 25 developers
  • Metric and denominator: Unit-test pass rate, quality scores
  • Measured or self-reported; sponsor: Measured; vendor

McKinsey, “Unleashing developer productivity with generative AI”, 27 Jun 2023

Finding: Documentation in about half the time; code generation about half; refactoring about two thirds; under 10% saving on complex tasks; junior developers 7-10% slower

Supports: Task efficiency

Caveats: Tiny in-house sample; lab conditions

Sample and method details
  • Sample and method: About 40 in-house developers, lab tasks
  • Metric and denominator: Task time
  • Measured or self-reported; sponsor: Measured (lab); consultancy

Surveys and telemetry datasets

DORA / Google Cloud, Accelerate State of DevOps 2024, 22 Oct 2024 (announcement)

Finding: Throughput -1.5%, stability -7.2%, documentation quality +7.5%, code quality +3.4%, review speed +3.1%; 39% little or no trust in AI code

Supports: Individual perception up; delivery down

Caveats: Cross-sectional association, not causal

Sample and method details
  • Sample and method: About 3,000 survey respondents; regression on AI adoption
  • Metric and denominator: Change per 25% increase in AI adoption
  • Measured or self-reported; sponsor: Self-reported; Google Cloud with sponsors

DORA, State of AI-assisted Software Development 2025, 23 Sep 2025 (announcement)

Finding: 90% use AI; over 80% perceive productivity gain; throughput association now positive; stability association still negative; 30% little or no trust

Supports: Amplifier framing

Caveats: Survey; no cost data; seven-capability model published Dec 2025

Sample and method details
  • Sample and method: About 5,000 respondents plus 100 hours of interviews
  • Metric and denominator: Associations; adoption shares
  • Measured or self-reported; sponsor: Self-reported; Google Cloud; partners include GitHub and GitLab

DORA, ROI of AI-assisted Software Development, May 2026 (coverage)

Finding: Illustrative 39% first-year ROI, about 8-month payback, J-curve with an initial dip; task gains 35-40% greenfield vs 10% or less on complex legacy code

Supports: Cost-benefit model only

Caveats: Authors call it a high-uncertainty estimate

Sample and method details
  • Sample and method: Framework with an illustrative 500-engineer organization
  • Metric and denominator: Modelled ROI
  • Measured or self-reported; sponsor: Modelled; Google Cloud; full PDF gated

Faros AI, “The AI Productivity Paradox”, 23 Jul 2025

Finding: +21% tasks, +98% PRs, +91% review time, +9% bugs per developer, +154% PR size; no company-level correlation with improvement

Supports: Individual up, organization flat

Caveats: Customer base only

Sample and method details
  • Sample and method: 10,000+ developers, 1,255 teams, telemetry
  • Metric and denominator: Per-developer deltas; company-level correlation
  • Measured or self-reported; sponsor: Measured; engineering-analytics vendor

Faros AI, “AI Engineering Report 2026: The Acceleration Whiplash”, 12 Apr 2026

Finding: +33.7% throughput, +16.2% merge rate, +54% bugs per developer, +242.7% incidents-to-PR ratio, +441.5% review time, +31.3% PRs merged without review

Supports: Throughput up, quality down

Caveats: Not controlled; may not generalize

Sample and method details
  • Sample and method: 22,000 developers, 4,000+ teams; lowest- vs highest-adoption periods
  • Metric and denominator: Pre/post deltas
  • Measured or self-reported; sponsor: Measured; vendor

Uplevel, Copilot impact study (2024; page updated 2026)

Finding: No significant productivity gain; 41% higher bug-introduction rate

Supports: Contradictory evidence

Caveats: Single customer base

Sample and method details
  • Sample and method: About 800 developers, telemetry
  • Metric and denominator: Bugs, cycle time
  • Measured or self-reported; sponsor: Measured; analytics vendor

Atlassian, State of Developer Experience 2025, 9 Jul 2025

Finding: 68% save 10+ hours per week with AI; 50% lose 10+ hours to friction; 16% of time spent coding

Supports: Perceived time saved

Caveats: Perception

Sample and method details
  • Sample and method: 3,500 developers and managers, six countries
  • Metric and denominator: Hours per week
  • Measured or self-reported; sponsor: Self-reported; vendor

Stack Overflow Developer Survey 2025, Jul 2025

Finding: 84% use or plan to use AI; 46% distrust accuracy vs 33% trust; 66% cite “almost right” code; 45% say debugging AI code takes longer; 17.9% use AI mostly for testing

Supports: Adoption and trust

Caveats: 2026 survey not yet public at verification date

Sample and method details
  • Sample and method: 49,000+ respondents, 177 countries
  • Metric and denominator: Share of developers
  • Measured or self-reported; sponsor: Self-reported

BCG, AI at Work 2026, 3 Jun 2026

Finding: 42% save a workday or more weekly; 66% get no guidance on what to do with saved time

Supports: Time saved not redirected

Caveats: Cross-industry

Sample and method details
  • Sample and method: 11,749 workers, 14 markets
  • Metric and denominator: Share of workers
  • Measured or self-reported; sponsor: Self-reported; consultancy

MIT NANDA, “The GenAI Divide: State of AI in Business 2025”, Jul 2025

Finding: “95% of organizations are getting zero return”; 5% of integrated pilots extract value

Supports: Enterprise pilot to P&L gap

Caveats: Authors: “directionally accurate”; not software-specific

Sample and method details
  • Sample and method: 52 interviews, 153 survey responses, 300 deployments
  • Metric and denominator: Share of pilots with P&L impact
  • Measured or self-reported; sponsor: Self-reported; academic preliminary

Quality-engineering surveys

World Quality Report 2024-25 (Capgemini, Sogeti, OpenText), 22 Oct 2024

Finding: 68% using or with roadmap for generative AI in QE; 72% report faster automation; 57% lack a comprehensive automation strategy; 34% implementing efficient QE practices

Supports: Adoption

Caveats: No cost-of-quality or budget-share figure publicly available

Sample and method details
  • Sample and method: 1,750+ senior executives, 33 countries
  • Metric and denominator: Share of organizations
  • Measured or self-reported; sponsor: Self-reported; vendor-sponsored

World Quality Report 2025-26, 13 Nov 2025

Finding: 89% piloting or deploying (37% production, 52% pilot); 15% enterprise-wide; average perceived productivity boost 19%; one third report minimal gains; barriers: data privacy 67%, integration 64%, hallucination 60%

Supports: Adoption; perceived productivity

Caveats: Full report registration-gated

Sample and method details
  • Sample and method: 2,000+ senior executives, 22 countries
  • Metric and denominator: Share of organizations; average perceived gain
  • Measured or self-reported; sponsor: Self-reported; vendor-sponsored

Consultancy and analyst positions

Bain, Technology Report 2024, “Beyond Code Generation”, 25 Sep 2024

Finding: “efficiency improvements of about 10% to 15% on average”; a reported 30% gain on eligible activities “represents a net efficiency improvement of 15% across developers’ total time”; “companies fail to monetize even these gains because they’re unable to reposition the saved time”

Supports: Capacity, explicitly not savings

Caveats: Opaque method

Sample and method details
  • Sample and method: Client experience and leader survey (n undisclosed)
  • Metric: “Efficiency” share of developer time
  • Measured or self-reported; sponsor: Self-reported; consultancy

Bain, Technology Report 2025, “From Pilots to Payoff”, 23 Sep 2025

Finding: “10% to 15% productivity boosts” but “the time saved isn’t redirected”; 25-30% with end-to-end process change; code generation is “25% to 35% of the time from initial idea to product launch”

Supports: Capacity, not savings

Caveats: Opaque method

Sample and method details
  • Sample and method: SaaS developer-team surveys (2,000-20,000 FTE teams)
  • Metric: Productivity share
  • Measured or self-reported; sponsor: Self-reported; consultancy

McKinsey, “The AI revolution in software development”, 1 Apr 2026

Finding: Top quintile 16-30% productivity and 31-45% quality; “simply giving developers AI tools does not meaningfully move the needle”; “success cases are still few and far between”

Supports: Top-quintile outcomes

Caveats: Method not transparent

Sample and method details
  • Sample and method: About 300 public companies plus client cases
  • Metric: Productivity and quality improvement
  • Measured or self-reported; sponsor: Analyst-derived; consultancy

Deloitte, “AI can help banks unleash a new era of software engineering productivity”, 24 Apr 2025

Finding: “save between 20% and 40% in software investments for the banking industry by 2028”; one bank observed 20% among a group of engineers using AI in testing

Supports: Cost-avoidance forecast

Caveats: Explicit prediction disclaimer

Sample and method details
  • Sample and method: Projection model
  • Metric: Share of banking software spend
  • Measured or self-reported; sponsor: Modelled; consultancy

Deloitte, “How can organizations engineer quality software in the age of generative AI?”, 28 Oct 2024

Finding: 45% use generative AI for coding, 38% lack confidence in results; LLM code success 90% on simple prompts to 42% on complex

Supports: Adoption and risk

Caveats: No original productivity measurement

Sample and method details
  • Sample and method: 2,770-organization survey plus 40 interviews; cites third parties
  • Metric: Adoption and confidence
  • Measured or self-reported; sponsor: Self-reported; consultancy

Gartner press releases: Apr 2024, Oct 2024, May 2025, Jul 2025, Jul 2026

Finding: 90% of enterprise engineers to use AI assistants by 2028; AI tools “will generate modest productivity increases by augmenting existing developer work patterns”; 71% of leaders find augmenting workflows a pain point; “tiny teams are not a cost optimization tactic”

Supports: Adoption only

Caveats: No public Gartner statistic on AI testing productivity; research notes and the AI-augmented testing Magic Quadrant are paywalled

Sample and method details
  • Sample and method: Surveys of 300-598 leaders; predictions
  • Metric: Adoption forecasts
  • Measured or self-reported; sponsor: Self-reported and forecast; analyst firm

Not verified (do not cite)

  • Any World Quality Report cost-of-quality or QA-budget-share figure (full reports are registration-gated).
  • Any Gartner statistic on AI-augmented testing productivity.
  • The full DORA 2026 ROI report contents beyond press coverage.
  • A Bain or Deloitte 2026 quantified testing gain.
  • Stack Overflow Developer Survey 2026 (URL returned 404 on 4 Sep 2026).
  • The Deloitte banking per-institution dollar figure (units ambiguous in extraction).