AI × QE Briefings Executive · Strategic vision

Quality engineering in the AI era

Executive · Strategic vision / September 2026

Quality engineering
for the AI era

AI × quality engineering · Industry research

Conceptual connected quality studio with human review and a release gate

02 / Strategic target state

A shared quality capability

Pan horizontally to explore the full diagram. A text description is available below.

Strategic target state: two quality missions use a shared assurance platform to create trustworthy delivery outcomes AI for QE Test design Failure diagnosis Test maintenance QE for AI Behavior and grounding Agent actions Drift and resilience SHARED QUALITY CAPABILITY Testable intent Requirements → task contracts → expectations Independent assurance Tests + evaluations + human review Operational learning Versioned evidence + observed failures Curated cases improve the next evaluation Delivery outcomes Faster validated change Less avoidable rework Trustworthy AI Evidence of behavior Bounded authority Reusable capability Shared evaluation assets Portable controls Strategic vision · outcomes are ambitions to validate, not reported bank results
Strategic target state · analyst outlook informs the ambition; the capability design is our synthesis. [G01][M01][D01][A01]
Research behind this design & diagram description
Finding

Gartner forecasts broader AI assistant use. McKinsey’s survey connects reported impact with changes across delivery and the operating model. DORA describes AI as an amplifier of the surrounding system.

Design implication

Invest in shared evaluation assets, platform controls and operational learning that serve both AI-assisted engineering and AI products. These sources do not establish a bank-specific return.

Diagram in words

AI for QE and QE for AI share testable intent, independent assurance and operational learning. That common capability supports validated delivery, trustworthy AI behavior and reuse across teams.

Inspect methods and source documents →

03 / Industry adoption

The constraints are organizational

What constrains adoptionShare of respondents reporting each barrier · World Quality Report 2025–26
Data privacy
67%
Integration complexity
64%
Hallucination / reliability
60%
AI / ML skill gaps
50%

Overlapping responses, not shares of a whole. Survey: >2,000 executives; question-specific n unavailable in the public release. Source: Capgemini, Sogeti and OpenText, November 2025. [W01]

04 / Evidence in context

Task gains depend on the setting

Pan horizontally to explore the full diagram. A text description is available below.

Two randomized-study findings measured different outcomes: a 26.08 percent increase in completed tasks in field experiments and a 19 percent increase in completion time in METR early-2025 study CUI ET AL. / THREE FIELD EXPERIMENTS METR / EARLY-2025 TRIAL +26.08% completed tasks +19% completion time Developer sample 4,867 Precision SE 10.3 percentage points Setting Coding assistants in three firms Developer / task sample 16 / 246 Precision 95% CI: +2% to +39% Setting Experienced developers; familiar repos The unit, task and user population determine what transfers. Separate outcomes; do not pool these percentages or treat them as QE savings rates.
Two randomized-study findings · different outcomes and settings; neither percentage is a QE savings benchmark. [E01][E02][E03]
Research behind this design & diagram description
Finding

Cui et al. report a 26.08% increase in completed tasks across three field experiments. METR’s early-2025 trial reports a 19% increase in completion time for experienced developers on familiar repositories.

Design implication

Segment pilots by task and user population, retain verification effort, and measure quality alongside throughput. METR’s later follow-up has selection concerns; it does not supply a clean causal time trend.

Diagram in words

Cui: 4,867 developers, completed tasks +26.08%, standard error 10.3 percentage points. METR: 16 developers, 246 tasks, completion time +19%, 95% confidence interval +2% to +39%. The studies measure distinct outcomes and should not be pooled.

Inspect methods and source documents →

05 / Explore the value logic

Capacity and realized value

$45.0killustrative net economic impact3.3% capacity released · 0.45% of addressable spend

Annual value bridge $ thousands · per $10M addressable QA spend
Base assumptions, illustrative annual values
Capacity equivalent$330000
Uncaptured capacity−$165000
AI and pilot cost−$100000
Quality allowance−$20000
Extra human effort$0
Net economic impact$45000
Net benefit vs. adoption All other current assumptions held fixed · $ thousands

Illustrative assumptions, not a forecast. Capacity is effort equivalent; cash savings need Finance-approved capture.

Activity share 55% · eligibility 60% · AI/pilot cost 1.0% · quality allowance 0.2%.

06 / Analyst and industry signals

The industry direction and its limits

Source / evidence typeWhat it saysStrategic implication
Gartner · forecast90% of enterprise software engineers using AI code assistants by 2028.Plan for widespread assisted work; tool use alone does not establish quality or savings.
McKinsey · surveyNearly 300 leaders surveyed; 100 assessed impact. Top performers reported broader delivery improvements.Invest in workflow and organizational change. These self-reported outcomes do not isolate causality.
DORA · observational researchAI interacts with the surrounding engineering system.Improve feedback, testing and platform foundations alongside adoption.

Public sources, reviewed September 2026. Gartner’s gated testing reports contribute public abstracts only; no proprietary vendor ranking is reproduced. [G01][G02][G03][M01][D01]

07 / Research beyond code generation

Industrial cases reveal reusable patterns

Industrial exampleObserved approachCapability to develop
Meta · TestGen-LLM / ACHGenerated tests face build, reliability and coverage checks; ACH adds fault-oriented validation.An independent test-quality gate, with meaningful assertions.
Google · Auto-DiagnoseFailure summaries link likely causes to relevant log lines in the review workflow.Diagnosis with inspectable evidence and reviewer feedback.
Uber / UT Austin · FlakyGuardTargeted execution context supports flaky-test repair in an enterprise Go repository.Repair validation that preserves the original test intent.

Company-specific cases demonstrate mechanisms. They do not establish a transferable bank savings rate or a product comparison. [E04][E05][E06][E07]

08 / Portfolio choices

The first workflows need a clear quality check

Pan horizontally to explore the full diagram. A text description is available below.

Proposed portfolio screen by action authority and ease of verifying the result Easy to verify Needs expert judgment BOUNDED STARTING POINTS Test drafts with an executable contract Failure summaries linked to logs CONTROLLED EXECUTION Sandbox test execution Draft pull requests through scoped tools BUILD THE EVALUATION CAPABILITY Risk analysis and exploratory test ideas Expert review defines useful output SEPARATE AUTHORITY DECISION Production changes and customer actions Require action-level evidence and recovery Advice / draft artifacts Greater action authority
Proposed portfolio screen · qualitative categories, not a scored market ranking. [E04][E06][A01][A03]
Research behind this design & diagram description
Finding

Industrial cases demonstrate bounded test generation and diagnosis. Agent guidance adds the consequences of actions to the evaluation problem.

Design implication

Select an initial workflow with an independent check and limited authority; assess judgment-heavy and consequential workflows separately.

Diagram in words

A two-by-two screen separates easy-to-verify results from expert judgment, and advice from greater action authority. Test drafts and log-linked diagnosis are candidate starting points.

Inspect methods and source documents →

09 / Research to strategy

What the findings change

Pan horizontally to explore the full diagram. A text description is available below.

Research-to-design map connecting DORA, Meta, NIST and OSFI findings with architecture decisions RESEARCH OBSERVATION DESIGN RESPONSE EVIDENCE TO RETAIN DORA / verification effort Generation can shift work into review. Join effort with run telemetry Measure net workflow effort Prep · review · correction D02 Meta / filtered test generation Passing is not proof of added value. Independent test-quality gate Protect approved assertions Stable runs · fault detection E04 · E05 NIST / lifecycle assurance Assurance continues in production. Version cases; feed failures back Re-evaluate configuration changes Case history · drift · release A01 · A02 OWASP + OSFI / agent actions Tools expand the impact of failure. Enforce permissions outside the model Test denied actions and recovery Identity · decision · approval A03 · R01
Research → proposed design → evidence to retain. The architecture choices are our synthesis, not publisher endorsements. [D02][E04][E05][A01][A02][A03][R01]
Research behind this design & diagram description
Finding

The sources address different failure modes: verification effort, weak generated tests, lifecycle change and the consequences of agent actions.

Design implication

Translate each finding into an observable design requirement. Review burden needs effort telemetry; test generation needs a quality gate; changing AI needs repeatable evaluation; tool use needs enforceable permissions.

Diagram in words

DORA informs joined effort measurement. Meta informs independent test-quality gates. NIST informs versioned evaluation and feedback. OWASP and OSFI inform external permission enforcement. Each design response has an associated evidence record.

Inspect methods and source documents →

10 / Continuous assurance

Quality improves through feedback

Pan horizontally to explore the full diagram. A text description is available below.

Proposed AI evaluation architecture with versioned application configurations, held-out cases, independent evaluation, release decisions, monitoring and feedback Versioned application Model · prompt · retrieval Tools · permissions · memory Evaluation corpus Typical + adversarial cases Held-out set · human labels Baseline and candidate runs Same tasks and test conditions Record outputs AND tool actions Independent scoring Deterministic contracts Calibrated judges + expert review Quality by scenario; cost and latency versioned traces Release gate Agreed floors hold? Hold + investigate Failed case → diagnosis Revise config or control no Controlled deployment Canary / bounded exposure Human escalation + rollback yes Production monitoring Outcomes · drift · denials Trace samples · incidents · rework Case curation Review oracle + privacy Add failures to regression new cases Proposed application-level assurance · evaluate again when context, behavior or authority changes
Proposed evaluation and release architecture · the unit of assurance is the configured application. [A01][A02][T03][T06]
Research behind this design & diagram description
Finding

NIST treats evaluation and risk management as lifecycle activities. Evaluation product documentation supports versioned datasets and comparisons; judge calibration still requires local evidence.

Design implication

Compare baseline and candidate configurations against held-out cases, gate releases on agreed floors, then turn reviewed production failures into regression cases.

Diagram in words

Versioned application configurations and a held-out evaluation corpus feed comparable runs and independent scoring. A failed release gate holds the change for investigation. A passing gate leads to controlled deployment and production monitoring; curated failures refresh the corpus.

Inspect methods and source documents →

11 / Operating model

Ownership across the quality lifecycle

Pan horizontally to explore the full diagram. A text description is available below.

Proposed operating responsibilities for product teams, the platform and independent challenge Product and QE owner Owns behavior and release outcomes Shared platform team Context, identity and runtime Service + evidence export Delivery team Test intent, assertions and review Accept through existing CI Delivery and risk owners Exceptions + evidence review Scope and escalation advice accountability Shared services Independent challenge Operations and value owners Incidents, adoption and capacity use
Proposed ownership model · adapt to existing institutional accountability. [D01][M01][A01]
Research behind this design & diagram description
Finding

DORA and McKinsey place AI adoption in the surrounding delivery organization. NIST frames governance across the lifecycle.

Design implication

Assign a product owner for outcomes, a platform owner for services and a separate challenge function; include operations and value capture.

Diagram in words

Product and QE accountability flows to the delivery team. Shared platform services enable delivery; delivery and risk owners challenge its evidence. Operations and value owners use the resulting outcomes.

Inspect methods and source documents →

12 / People and capability

The quality role gains new disciplines

CapabilityWhat practitioners need to doObservable learning evidence
Test intent and oraclesDefine expected behavior and challenge plausible but weak assertions.A test catches a relevant fault and preserves the contract.
Evaluation and curationBuild representative cases; calibrate expert labels and model judges.Evaluation results survive a held-out scenario review.
Agent assuranceInspect tool authority, context boundaries and recovery behavior.A denied-action exercise leaves no unauthorized side effect.
Workflow ownershipMeasure review and correction work; coach teams through real tasks.Task-level outcomes include the cost of verification.

Proposed capability agenda informed by research on verification work and lifecycle assurance; it is not a staffing reduction plan. [D02][G04][A01][E05]

13 / Technology strategy

The platform outlasts a model

Pan horizontally to explore the full diagram. A text description is available below.

Enterprise assurance platform overview; inspect components or follow the guided paths Team workflows IDE · workbench Delivery systems Repository · CI AI applications RAG · agents Context Permitted snapshots QE runtime Bounded plan + model Action gateway Identity + policy Checks Sandbox + validity Evidence store Run + artifact + decision Regression corpus Reviewed failure cases Release authority Reviewer + change gate Evaluation Held-out + adversarial
Inspect a component: Select a component in the diagram
Proposed logical architecture · select a component to inspect its role and research basis. [A01][A03][R01][E04]
Research behind this design & diagram description
Finding

NIST frames assurance across the lifecycle. OWASP and OSFI identify agent identity, tool authority and traceability as control concerns; Meta separates generation from validation.

Design implication

Separate context, generation, policy enforcement, independent evaluation and release authority. Preserve a shared evidence record across those boundaries.

Diagram in words

Engineers and delivery systems supply tasks and approved context. The runtime proposes actions through a policy gateway. Independent sandbox checks and application evaluations supply evidence to the release authority. Traces and decisions feed the evidence store; curated failures become regression cases.

Inspect methods and source documents →

14 / Investment architecture

Shared foundations and domain expertise

Pan horizontally to explore the full diagram. A text description is available below.

Proposed investment split between product-specific quality assets and reusable assurance services PRODUCT-SPECIFIC ASSETS Payments Transaction invariants Failure and recovery cases Customer service AI Grounded response criteria Escalation and tool scenarios Engineering workflows Repository test contracts Review and rework baselines REUSABLE ASSURANCE SERVICES Dataset registry · Evaluation runners · Identity + gateway · Evidence store · Recovery tooling REPLACEABLE MODEL AND TOOL ADAPTERS Procure capabilities against local tests; preserve access to configurations, artifacts and evidence.
Proposed investment architecture · layer size does not represent budget allocation. [D01][M01][T03][T06]
Research behind this design & diagram description
Finding

Organizational research emphasizes delivery foundations. Evaluation tools expose datasets and run comparisons, but task fit still requires local validation.

Design implication

Fund reusable services centrally while product teams own domain-specific quality assets. Use adapters and evidence export requirements to support substitution.

Diagram in words

Three example domains contribute different quality assets. They use a common assurance layer, which connects through replaceable model and tool adapters.

Inspect methods and source documents →

15 / Capability roadmap

Evidence gates the next horizon

Pan horizontally to explore the full diagram. A text description is available below.

Proposed capability roadmap with three horizons and evidence gates for stronger AI participation 01 Trusted assistance Draft tests + diagnose failures Independent checks; human approval 02 Connected assurance Shared evaluation and context services Cross-team evidence and feedback 03 Controlled agents Multi-step work with bounded authority Action-level controls and recovery Gate: local quality + net effort Gate: repeatable outcomes across teams Gate: safe actions + tested resilience Proposed horizons, not a dated forecast or measured maturity score
Proposed capability horizons · progression depends on demonstrated outcomes and authority, rather than calendar dates. [M01][D01][A01][R01]
Research behind this design & diagram description
Finding

Industry research places AI adoption within delivery-system and operating-model change. Assurance guidance makes the consequences of actions and ongoing monitoring relevant to expansion.

Design implication

Develop bounded assistance first, connect shared assurance services next and expand agent authority only after local quality, repeatability and recovery evidence supports it.

Diagram in words

Horizon one is trusted assistance, gated by local quality and net effort. Horizon two is connected assurance, gated by repeatability across teams. Horizon three is controlled agents, gated by safe actions and tested resilience.

Inspect methods and source documents →

16 / Strategic measures

The leadership evidence contract

Pan horizontally to explore the full diagram. A text description is available below.

Proposed leadership evidence contract linking strategic goals, measures and decisions STRATEGIC OUTCOME EVIDENCE TO EXAMINE LEADERSHIP DECISION Delivery capacity Net effort with review and rework Redeploy verified capacity Trustworthy outcomes Fault detection and escaped defects Hold the quality floor Reusable capability Results by team and task Fund the shared service Controlled autonomy Denials and recovery tests Set the next authority limit Every readout includes task mix, sample, period, costs and uncertainty. No organizational results are claimed here.
Proposed leadership scorecard · measures and decisions, with no invented performance values. [D02][E01][E03][A01]
Research behind this design & diagram description
Finding

Studies use different tasks, populations and outcome measures; verification effort can offset generation gains.

Design implication

Tie each strategic objective to observable evidence and an explicit decision. Keep quality and control floors alongside capacity measures.

Diagram in words

Four rows connect delivery capacity, trustworthy outcomes, reusable capability and controlled autonomy to evidence and leadership decisions.

Inspect methods and source documents →

17 / Risk and accountability

Governance follows the system lifecycle

Pan horizontally to explore the full diagram. A text description is available below.

Proposed governance loop connecting system inventory, authority, evidence and ongoing review System inventory Purpose + dependencies Named accountable owner Authority and appetite Allowed data and actions Escalation + stop rules Evidence review Evaluation + control results Exceptions + open risks Operational review Incidents + service changes Revisit scope and authority Approved scope New evidence OSFI context: July 2026 bulletin = sound practices; revised E-23 takes effect 1 May 2027.
Proposed governance loop · assess applicability through existing risk functions. [A01][R01][R02]
Research behind this design & diagram description
Finding

NIST links governance to lifecycle risk management. OSFI discusses AI accountability and resilience; revised E-23 has a future effective date.

Design implication

Maintain an inventory and an explicit authority boundary, then revisit them after incidents or material changes. This diagram is not a compliance certification.

Diagram in words

Inventory informs authority limits. Evidence review supports an approved scope. Operations feed incidents and service changes back into the inventory and authority decision.

Inspect methods and source documents →

18 / Closing decision

The next decision: capability, foundations and scope

Ambition

Quality as a shared capability

Choose the delivery constraints and AI risks the organization needs to address.

Investment

Foundations and people

Develop evaluation skills, platform integration and accountable ownership.

Expansion

Value with evidence

Grow the capabilities that improve outcomes within their control boundary.

Vision: AI contributes more work while the organization strengthens its ability to judge and govern that work.

19 / Strategic choices in practice

Choose the shared assurance model

Pan horizontally to explore the full diagram. A text description is available below.

Three strategic choices: local assistants, shared assurance services, bounded agents 01 Local assistants Fast individual adoption Duplicated context and checks Local learning stays fragmented 02 Shared assurance Portable tests and evidence Domain-owned expectations Common evaluation and controls 03 Bounded agents Multi-step delegated work Stronger action authority Higher recovery obligations Use to learn tasks Recommended strategic center Earn through evidence Sequencing choice: standardize the proof before expanding delegated authority. Trade-off: shared services require ownership; agent breadth requires independently tested containment.
Authored strategic choices · no bank maturity or savings claim. [G01][D01][A01][A03]
Research behind this design & diagram description
Finding

Industry outlook and lifecycle guidance support broader AI participation and independent assurance.

Design implication

Use shared assurance as the strategic center; delegate actions only as boundaries and evidence mature.

Diagram in words

Local assistants accelerate individual tasks. Shared assurance makes expectations, evaluation and evidence reusable. Bounded agents add delegated actions and stronger recovery obligations. These are design choices, not measured maturity scores.

Inspect methods and source documents →

20 / Strategic choices in practice

What changes in a payment-API workflow

Pan horizontally to explore the full diagram. A text description is available below.

Illustrative payment API change: current working hypothesis and proposed target workflow CURRENT WORKING HYPOTHESIS · VALIDATE WITH THE BANK Interpret requirement Interpret intent Write tests Local prompts Review change Rebuild evidence Release Late evidence assembly PROPOSED TARGET · NEGATIVE PAYMENT AMOUNTS MUST BE REJECTED Domain contract Owner-approved rule Versioned test oracle Bounded generation Permitted context only Sandbox tests Independent proof Original passes Mutant fails Reviewer decides Evidence-led release Existing release owner Failure → new test case Shared: identity, evaluation runners, evidence schema. Domain: payment rules, test cases, acceptance and service outcomes.
Illustrative financial-services workflow · current-state row is a hypothesis, not an observed bank condition. [E04][E05][D02][A03]
Research behind this design & diagram description
Finding

Test-generation studies use quality gates; verification work remains part of the workflow.

Design implication

Join a domain-owned payment contract to shared evidence services and existing release authority.

Diagram in words

The current-state hypothesis has local interpretation, test writing, review and late evidence assembly. The target versions the approved payment rule, generates tests in a scoped sandbox, independently checks fault detection and retains the reviewer decision and release evidence.

Inspect methods and source documents →

21 / Strategic choices in practice

Sequence shared assets before wider agency

HorizonShared institutional assetDomain responsibility / trade-off
Prove the workflowTask identity, approved models, baseline measurement and evidence contract.Payment owner defines behavior; delivery team measures review effort. Keep scope narrow.
Reuse the assuranceEvaluation runners, policy enforcement and reusable failure cases.Own domain cases, acceptance and outcomes. Shared services need funded owners.
Delegate bounded actionsScoped tool adapters, recovery exercises and operational evidence.Service owner accepts consequence and recovery duty. Expand one authority boundary at a time.

22 / Client adoption assumptions

Client adoption requires funded foundations

Adoption decisionRequired evidence accepted · owner accountable · foundation work funded
Business & peopleExpected outcomes · reviewer capacity · offshore handoffs · measured baseline
Runnable engineeringInfrastructure · CI/CD · test data · dependency control · testable applications
Supported servicesFrameworks & devices · AI access · run evidence · platform ownership
Authored adoption assumptions. Validate each application and workflow; unknown requirements remain open. Assumptions and evidence →

Fund the missing capabilities and reviewer capacity before promising scale. A workflow can progress only when its required evidence is accepted.

23 / QE modernization

QE modernization and AI adoption

01Trusted work

Reviewed scenarios, expected outcomes and ownership

AI drafts with domain reviewA reviewer can identify a wrong answer
02Repeatable execution

Reliable suites, isolated data and controlled dependencies

AI drafts tests that run in CIAnother engineer reproduces the result
03Shared QE services

Reusable environments, contracts, artifacts and support

Assistance works across teamsA second team onboards and operates it
04Broader AI execution

Evaluated tools, bounded actions and recovery

Agents execute within approved scopeFailures stop, evidence persists, people can intervene
Proposed capability progression. Reviewed assistance can begin while foundations improve. Broader execution requires evidence for its specific scope. Modernization workstreams and sources

Develop reliable, reusable QE capability alongside reviewed AI assistance. Broader execution depends on the evidence available for the selected workflow.

24 / Fintech evidence and adoption

Financial-services outcomes

Libra Internet Bank

35% less time

Test-creation timeIndex: previous time = 100

UiPath customer case. Normalized from the reported reduction; review effort is unspecified. F1

Goldman Sachs

36% → 72%

Unit-test coverageShare of a selected module (%)

Diffblue customer case. Coverage measures exercised code; it does not establish financial correctness. F2

Fiserv

65% fewer incidents

Major incidentsIndex: previous year = 100

Tricentis modernization case. AI testing was a later pilot; this is not an AI-attributed reduction. F4

Separate outcomes and denominators. Each panel has its own meaning; no combined savings rate is calculated.

Reported customer outcomes use different measures and evidence types. Use them to select a trial; establish the client’s own baseline.

25 / Fintech evidence and adoption

A focused first engagement

Requirements ready

Onboarding design

Draft scenarios in the current test-management workflow.

Proof

Accepted coverage and measured review effort.

Execution repeatable

Payment API tests

Generate candidates in the existing framework and CI pipeline.

Proof

Correct behavior passes; a known fault fails.

Diagnostics available

Failure triage

Draft diagnoses alongside the existing disposition process.

Proof

Faster correct decisions with traceable evidence.

Proposed entry points. Choose one using the client's constraints; inspect prerequisites and prepare a trial brief.

Select the entry point that fits the client’s foundations. Build reusable assets and retain evidence for the next adoption decision.

Our Banking Client / Authored payment scenario

What AI contributes, from requirement to release

01 / Understand the change Clarify retry behaviorPAY-142 · accepted criteria 02 / Design test cases Draft scenarios and boundariesMATRIX-142 · reviewed cases 03 / Draft unit tests Test retry and dedupe logicUNIT-142 · developer review 04 / Build API / UI tests Turn cases into test codePR-142 · reviewed test commit 05 / Prepare data and stubs Draft isolated test inputsFIX-142 + STUB-142 · validated 06 / Select and run tests Recommend affected coverageRUN-142 · runner verdict 07 / Investigate failures Link evidence to a hypothesisDEF-142 · reviewed diagnosis 08 / Repair and retest Propose reviewed changesRERUN-143 · fresh proof 09 / Review release evidence Summarize gaps and retain casesPACK-142 · human decision
Nine steps, one payment scenario. AI drafts or recommends; engineers review; CI executes; release owners decide.

AI assists the testing team. For code, fixtures, failure evidence and release handoffs, open the banking engineering walkthrough →

Slide notes

Dictionary · terms & acronyms ↗ (opens in a new tab)