Read each supported claim and caveat alongside the finding. Expand the method details to check the sample and measurement. A 55% task speed-up on one synthetic task and a 19% slowdown on real issues are both real results; neither is a QA budget number.
Finding: Completion time -18% (CI -38% to +9%) original cohort; -4% (CI -15% to +9%) new cohort; negative values mean faster completion
Supports: Inconclusive; speedup point estimates with substantial selection bias
Caveats: 30-50% of developers reported withholding some tasks; recruitment and task selection make the effect estimate unreliable; both confidence intervals include no effect; study design being revised
Finding: Documentation in about half the time; code generation about half; refactoring about two thirds; under 10% saving on complex tasks; junior developers 7-10% slower
Supports: Task efficiency
Caveats: Tiny in-house sample; lab conditions
Sample and method details
Sample and method: About 40 in-house developers, lab tasks
Metric and denominator: Task time
Measured or self-reported; sponsor: Measured (lab); consultancy
Finding: 90% use AI; over 80% perceive productivity gain; throughput association now positive; stability association still negative; 30% little or no trust
Supports: Amplifier framing
Caveats: Survey; no cost data; seven-capability model published Dec 2025
Sample and method details
Sample and method: About 5,000 respondents plus 100 hours of interviews
Metric and denominator: Associations; adoption shares
Measured or self-reported; sponsor: Self-reported; Google Cloud; partners include GitHub and GitLab
Finding: Illustrative 39% first-year ROI, about 8-month payback, J-curve with an initial dip; task gains 35-40% greenfield vs 10% or less on complex legacy code
Supports: Cost-benefit model only
Caveats: Authors call it a high-uncertainty estimate
Sample and method details
Sample and method: Framework with an illustrative 500-engineer organization
Metric and denominator: Modelled ROI
Measured or self-reported; sponsor: Modelled; Google Cloud; full PDF gated
Finding: 84% use or plan to use AI; 46% distrust accuracy vs 33% trust; 66% cite “almost right” code; 45% say debugging AI code takes longer; 17.9% use AI mostly for testing
Supports: Adoption and trust
Caveats: 2026 survey not yet public at verification date
Sample and method details
Sample and method: 49,000+ respondents, 177 countries
Finding: 68% using or with roadmap for generative AI in QE; 72% report faster automation; 57% lack a comprehensive automation strategy; 34% implementing efficient QE practices
Supports: Adoption
Caveats: No cost-of-quality or budget-share figure publicly available
Sample and method details
Sample and method: 1,750+ senior executives, 33 countries
Metric and denominator: Share of organizations
Measured or self-reported; sponsor: Self-reported; vendor-sponsored
Finding: “efficiency improvements of about 10% to 15% on average”; a reported 30% gain on eligible activities “represents a net efficiency improvement of 15% across developers’ total time”; “companies fail to monetize even these gains because they’re unable to reposition the saved time”
Supports: Capacity, explicitly not savings
Caveats: Opaque method
Sample and method details
Sample and method: Client experience and leader survey (n undisclosed)
Metric: “Efficiency” share of developer time
Measured or self-reported; sponsor: Self-reported; consultancy
Finding: “10% to 15% productivity boosts” but “the time saved isn’t redirected”; 25-30% with end-to-end process change; code generation is “25% to 35% of the time from initial idea to product launch”
Supports: Capacity, not savings
Caveats: Opaque method
Sample and method details
Sample and method: SaaS developer-team surveys (2,000-20,000 FTE teams)
Metric: Productivity share
Measured or self-reported; sponsor: Self-reported; consultancy
Finding: Top quintile 16-30% productivity and 31-45% quality; “simply giving developers AI tools does not meaningfully move the needle”; “success cases are still few and far between”
Supports: Top-quintile outcomes
Caveats: Method not transparent
Sample and method details
Sample and method: About 300 public companies plus client cases
Metric: Productivity and quality improvement
Measured or self-reported; sponsor: Analyst-derived; consultancy
Finding: “save between 20% and 40% in software investments for the banking industry by 2028”; one bank observed 20% among a group of engineers using AI in testing
Supports: Cost-avoidance forecast
Caveats: Explicit prediction disclaimer
Sample and method details
Sample and method: Projection model
Metric: Share of banking software spend
Measured or self-reported; sponsor: Modelled; consultancy
Finding: 90% of enterprise engineers to use AI assistants by 2028; AI tools “will generate modest productivity increases by augmenting existing developer work patterns”; 71% of leaders find augmenting workflows a pain point; “tiny teams are not a cost optimization tactic”
Supports: Adoption only
Caveats: No public Gartner statistic on AI testing productivity; research notes and the AI-augmented testing Magic Quadrant are paywalled
Sample and method details
Sample and method: Surveys of 300-598 leaders; predictions
Metric: Adoption forecasts
Measured or self-reported; sponsor: Self-reported and forecast; analyst firm
Not verified (do not cite)
Any World Quality Report cost-of-quality or QA-budget-share figure (full reports are registration-gated).
Any Gartner statistic on AI-augmented testing productivity.
The full DORA 2026 ROI report contents beyond press coverage.
A Bain or Deloitte 2026 quantified testing gain.
Stack Overflow Developer Survey 2026 (URL returned 404 on 4 Sep 2026).
The Deloitte banking per-institution dollar figure (units ambiguous in extraction).