Research library / Reviewed 5 September 2026

Documents behind the research

Search 30 curated sources. Each record identifies what was accessible, what it supports and what it cannot establish. Publisher links open the originals.

Download register (CSV) · JSON · QA companion brief v1.7.0 (13-page PDF)

30 sources

G01 / Forecast / Industry

Software engineering trends for 2025 and beyond

Gartner · 2025-07-01 · Reviewed 2026-09-05

Forecasts 90% of enterprise software engineers using AI code assistants by 2028.

Method, limits and access

A forecast about tool use, not a measured QE outcome or a savings estimate.

Access: Public text

G02 / Analyst perspective / Industry

How to Select an AI-Augmented Software Testing Platform

Gartner · 2026-02-03 · Reviewed 2026-09-05

The public abstract frames platform selection around quality and business value beyond test automation.

Method, limits and access

Only the abstract was reviewed; detailed criteria, findings and vendor assessments require licensed access.

Access: Public abstract; full report gated

G03 / Analyst perspective / Industry

Journey Guide to Improving Software Testing in the AI Age

Gartner · 2026-04-02 · Reviewed 2026-09-05

The public summary positions continuous quality as a response to more complex AI-era delivery.

Method, limits and access

The full guide was not available for review. No proprietary maturity model is reproduced.

Access: Public abstract; full report gated

G04 / Forecast / Industry

Smaller software engineering teams by 2029

Gartner · 2026-07-07 · Reviewed 2026-09-05

Forecasts 60% of organizations adopting smaller engineering teams at scale by 2029; emphasizes team redesign and platform enablement.

Method, limits and access

Not evidence for reducing QE headcount; the release distinguishes restructuring from cost optimization.

Access: Public text

M01 / Survey / Industry

Unlocking the value of AI in software development

McKinsey · 2025-11-03 · Reviewed 2026-09-05

Surveyed nearly 300 leaders; 100 assessed impact. The top outcome quintile reported stronger productivity, quality and speed.

Method, limits and access

Self-reported outcomes and selected top performers; no causal estimate and no representative bank savings rate.

Access: Public text

M02 / Case study / Industry

The AI revolution in software development

McKinsey · 2026-04-01 · Reviewed 2026-09-05

Describes agent-based delivery and an unnamed global bank implementation as a strategic direction.

Method, limits and access

Consultancy-reported case; its speed and cost claims lack a public independent counterfactual. Do not transfer them into a business case.

Access: Public text

W01 / Survey / Industry

World Quality Report 2025–26: public findings

Capgemini / Sogeti / OpenText · 2025-11-13 · Reviewed 2026-09-05

Reports widespread experimentation, limited enterprise scale and material privacy, integration, reliability and skills barriers.

Method, limits and access

Vendor-sponsored survey of over 2,000 executives in 22 countries and 10 sectors. Question-specific sample sizes are not in the public release.

Access: Public release; full report registration

D01 / Survey / Industry

State of AI-assisted Software Development 2025

DORA / Google Cloud · 2025-09-23 · Reviewed 2026-09-05

AI adoption interacts with the surrounding engineering system; organizational foundations matter for delivery outcomes.

Method, limits and access

Observational research and reported associations do not isolate the effect of a tool. Partner-sponsored research.

Access: Public text

D02 / Qualitative study / Industry

Balancing AI tensions: moving from AI adoption to effective software development

DORA · 2026-03-10 · Reviewed 2026-09-05

Analysis of 1,110 open-ended Google engineer responses highlights verification work, context limitations and review friction.

Method, limits and access

One-company qualitative sample from Q3 2025; prompts may have focused respondents on code generation.

Access: Public text

E01 / Randomized study / Delivery evidence

Early-2025 AI on experienced open-source developer productivity

METR · 2025-07-10 · Reviewed 2026-09-05

16 developers and 246 tasks: access to early-2025 AI increased completion time by 19% (95% CI: +2% to +39%).

Method, limits and access

Experienced developers on familiar repositories; narrow task and tool setting, not a claim about all developers or 2026 agents.

Access: Public text

E02 / Study update / Delivery evidence

Updated developer productivity study

METR · 2026-02-24 · Reviewed 2026-09-05

Point estimates suggest faster work: returning developers −18% time, new recruits −4%; both confidence intervals include no effect.

Method, limits and access

Selection effects and concurrent-agent time measurement weaken the estimates. This is not a clean temporal trend against the earlier trial.

Access: Public text

E03 / Randomized study / Delivery evidence

The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers

Cui and colleagues / Management Science · 2026-02-27 · Reviewed 2026-09-05

Across 4,867 developers, the pooled estimate was 26.08% more completed tasks, with standard error 10.3 percentage points.

Method, limits and access

Coding-assistant field experiments, not QE-specific labor savings. Public author PDF is the February 2025 working paper; publication followed in 2026.

Access: Publisher abstract + open author paper

E04 / Case study / AI for QE

Automated Unit Test Improvement using Large Language Models at Meta

Alshahwan and colleagues / FSE 2024 · 2024-02-14 · Reviewed 2026-09-05

TestGen-LLM improves existing tests using build, reliability and coverage filters before developer review.

Method, limits and access

Single-company evaluation. Generated tests and test-a-thon recommendations use different denominators; acceptance does not establish defect reduction.

Access: Public text

E05 / Case study / AI for QE

LLM-powered bug catchers: Meta ACH

Meta Engineering · 2025-02-05 · Reviewed 2026-09-05

Uses mutation-guided generation: candidate tests must detect a modeled fault while passing the original code.

Method, limits and access

Tests target selected fault concerns; effectiveness depends on mutation relevance and human review. Company-authored implementation report.

Access: Public text

E06 / Case study / AI for QE

LLM-Based Automated Diagnosis Of Integration Test Failures At Google

Ziftci and colleagues / Google · 2026-04-13 · Reviewed 2026-09-05

Reports 90.14% diagnosis accuracy on 71 manually evaluated failures and deployment on 52,635 distinct failing tests.

Method, limits and access

The small reviewed accuracy sample is different from the deployment population. Feedback usefulness is not diagnostic accuracy.

Access: Public text

E07 / Case study / AI for QE

FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale

Li and colleagues / UT Austin and Uber · 2025-11-18 · Reviewed 2026-09-05

Uses dynamic call-graph context and analysis to repair flaky tests in a Go monorepo.

Method, limits and access

Six-month enterprise case. Reported repair and acceptance percentages use different populations; weak fixes can change test semantics.

Access: Public text

A01 / Framework / QE for AI

Artificial Intelligence Risk Management Framework: Generative AI Profile

NIST · 2024-07-26 · Reviewed 2026-09-05

Provides a cross-sector structure for identifying and managing generative-AI risks throughout the lifecycle.

Method, limits and access

Voluntary guidance; translate it into system-specific controls and testable requirements.

Access: Public text

A02 / Framework / QE for AI

AI RMF Core: Govern, Map, Measure, Manage

NIST · 2023-01-26 · Reviewed 2026-09-05

Connects lifecycle risk management with representative evaluation, documented uncertainty and ongoing monitoring.

Method, limits and access

RMF 1.0 remains the referenced framework; NIST says it is being revised. It is not a product certification.

Access: Public text

A03 / Framework / QE for AI

Top 10 for Agentic Applications 2026

OWASP GenAI Security Project · 2025-12-09 · Reviewed 2026-09-05

Threat guidance extends evaluation to agent goals, tool use, identity, supply chains and memory.

Method, limits and access

Community risk taxonomy; use it to design scenarios, not to infer measured incident rates for a bank.

Access: Public text

A04 / Framework / QE for AI

OWASP GenAI LLM Top 10 2026

OWASP GenAI Security Project · 2026-08-03 · Reviewed 2026-09-05

The 2026 guide updates LLM application risks and explicitly pairs model-component risks with agentic-system risks.

Method, limits and access

Resource page dated August 3; release announcement September 1 has a September 2 dateline. Downloaded PDF retains date placeholders. Reviewed as the linked 2026 edition.

Access: Public text

A05 / Framework / QE for AI

Agent Control Standard (ACS)

OWASP GenAI Security Project · 2026-09-01 · Reviewed 2026-09-05

Introduces common hooks for runtime policy enforcement and observability across agent frameworks.

Method, limits and access

Newly donated open standard; evaluate implementation coverage and integration maturity before relying on portability.

Access: Public text

R01 / Supervisory guidance / Financial services

Generative and agentic AI: technology, cyber security and operational resilience

OSFI · 2026-07 · Reviewed 2026-09-05

Discusses scoped agent identities, tool restrictions, testing, traceability, approval and fallback practices for financial institutions.

Method, limits and access

Technology Risk Bulletin offers sound practices that complement existing guidelines; it is not a new standalone binding AI rule.

Access: Public text

R02 / Supervisory guidance / Financial services

Guideline E-23: Model Risk Management (2027)

OSFI · 2025-09-11 · Reviewed 2026-09-05

Risk-based model governance and lifecycle expectations provide a planning reference for applicable AI systems.

Method, limits and access

Effective May 1, 2027. Assess applicability under institutional model policy; do not label every coding assistant a regulated model by default.

Access: Public text

T01 / Product documentation / Technology landscape

Tosca Agentic AI documentation

Tricentis · Living documentation · Reviewed 2026-09-05

Documents natural-language support for finding test assets, explaining results and generating test cases.

Method, limits and access

Capability description for Tosca 2026.1; not an independent assessment of accuracy, enterprise fit or return.

Access: Public text

T02 / Product documentation / Technology landscape

Deterministic Visual AI tests

Applitools · Living documentation · Reviewed 2026-09-05

Describes visual comparison against approved baselines within existing testing frameworks.

Method, limits and access

Vendor product description. Validate false alerts, missed changes, dynamic content and baseline governance on your applications.

Access: Public text

T03 / Product documentation / Technology landscape

Evaluation concepts and workflows

LangSmith / LangChain · Living documentation · Reviewed 2026-09-05

Documents offline datasets, evaluators and online feedback loops for AI applications.

Method, limits and access

Tool documentation does not establish evaluator validity. Calibrate judges and human labels to the target task.

Access: Public text

T04 / Product documentation / Technology landscape

About GitHub Copilot code review

GitHub · Living documentation · Reviewed 2026-09-05

Documents AI review suggestions within pull requests and editors.

Method, limits and access

Review suggestions supplement reviewer judgment; a product capability is not evidence of prevented defects.

Access: Public text

T05 / Product documentation / Technology landscape

Red teaming for AI agents

Promptfoo · Living documentation · Reviewed 2026-09-05

Documents adversarial testing of agent tool use, access boundaries and context.

Method, limits and access

A red-team suite samples a threat model. Passing it does not prove that an agent is secure.

Access: Public text

T06 / Product documentation / Technology landscape

View and compare evaluation results

Microsoft Foundry · Living documentation · Reviewed 2026-09-05

Documents run-level and row-level evaluation comparison, including quality and operational metrics.

Method, limits and access

Platform evaluation requires representative datasets, appropriate metrics and deployment-specific thresholds.

Access: Public text

S01 / Survey / Industry

Developer Survey 2025: AI

Stack Overflow · 2025 · Reviewed 2026-09-05

Reported trust in AI accuracy remains mixed; distrust exceeds trust in the survey responses.

Method, limits and access

Self-selected developer survey. Overall sample size and AI-question response counts are different denominators.

Access: Public text

The local archive contains the successfully retrieved public PDFs and a checksum manifest. Gated reports were not downloaded. This site links publisher originals and publishes its own synthesis; it does not rehost third-party reports.

Research coverage, evidence grades and exclusions · Current audience PDF exports