Libra Internet Bank
35% less time
Test-creation timeIndex: previous time = 100
UiPath customer case. Normalized from the reported reduction; review effort is unspecified. F1
AI × QE / Financial-services evidence
Banks and fintech firms are using AI to expand testing capability. Use their reported results to choose a focused trial, define the evidence you need and decide where to begin.
Seven cases connect test design, legacy unit tests, change intelligence and QE modernization. Named vendor cases, a bank report and a controlled coding study answer different questions. These are peer examples; Our Banking Client remains a fictional proposal.
01 / What has been reported
Libra Internet Bank
Test-creation timeIndex: previous time = 100
UiPath customer case. Normalized from the reported reduction; review effort is unspecified. F1
Goldman Sachs
Unit-test coverageShare of a selected module (%)
Diffblue customer case. Coverage measures exercised code; it does not establish financial correctness. F2
Fiserv
Major incidentsIndex: previous year = 100
Tricentis modernization case. AI testing was a later pilot; this is not an AI-attributed reduction. F4
Our conclusion: start where a team can independently judge the output and measure the work. Public cases support that opportunity, but they do not establish a transferable QA savings rate. No case reviewed here supplies an independently audited, comparable bank-wide net AI-QE saving.
02 / Inspect the example behind the headline
AI in testing covers direct test-generation examples. QE foundations covers broader modernization results. Adjacent AI covers coding and change-risk work that informs the approach without measuring test-team savings.
Named customer / vendor case
35% less test-creation time
Reported bank-level test design time
UiPath reports 35% less time creating test cases and AI generation for more than 60% of cases used. The pilot began with two onboarding teams.
Named customer / vendor case
36% → 72% coverage
Reported unit-test coverage for a selected module
Diffblue reports this coverage increase within 24 hours. A separate application received 3,211 generated tests; review took one day.
Bank-authored operational report
5% → 100% of changes screened
Reported change-request screening coverage
DBS reports an 81% decrease in monthly average change-related incidents with AI risk scoring. It separately describes JIRA Assist for user stories and test cases.
Named customer / vendor case
65% fewer major incidents
Reported year-over-year incidents
Tricentis reports the decline during testing standardization across a large application estate.
Named customer / vendor case
5+ weeks → 5 days
Reported regression cycle
The case reports shorter regression cycles within an 18-month transformation and releases moving from every three to four months to monthly.
Bank-authored controlled coding experiment
30.98 → 17.86 minutes
Mean self-reported time per Python challenge
The controlled crossover study involved over 100 engineers. Analysis retained 172 task observations after removing six duplicates and 22 unsolved tasks.
Anonymous supplier case
42% less QA cycle effort claimed
Supplier-reported QA cycle effort
TestingXperts reports AI use in test design, automation, impact analysis, reporting and investigation.
03 / Patterns to adopt
These are our recommendations drawn from the cases and measurement guidance, to be adapted to the client's environment.
Choose one onboarding or payment change with clear expected outcomes. Preserve the current toolchain where it supports the trial.
Proof to collect. A domain reviewer can accept or reject each proposed scenario.
Coverage and passing tests are useful signals. Challenge money movement, rounding and retries with known failures and reviewed expected outcomes.
Proof to collect. A duplicate transfer or wrong balance causes the candidate test to fail.
Assign owners to environments, fixture reset, dependency models and reusable frameworks. Track their improvements separately from AI.
Proof to collect. Another engineer reproduces the run and locates its original evidence.
Use comparable work, the same readiness level and explicit task allocation. Include failures, review, rework and setup; record missing observations.
Proof to collect. The team can explain selection, exclusions and uncertainty in the result.
Observe risk-based test selection or triage suggestions before changing required suites. Inspect missed defects and incorrect groupings.
Proof to collect. Mandatory payment checks still run, and wrong suggestions are visible.
Coach domain QA and automation engineers together. Name a support owner and test transfer to a second team before expanding across 50–100 staff.
Proof to collect. The second team uses and maintains the pattern without its original authors.
04 / Conditions for adoption
Assess the required capabilities for the selected application and workflow. Reviewed scenario drafting can begin when its own inputs and reviewers are ready; broader AI execution depends on repeatable tests, controlled environments and trustworthy evidence.
One adoption model, two views: six modernization workstreams connect to the twelve assumptions in the readiness worksheet. A marked cell means that some capabilities in that workstream are required.
| Modernization workstream | Test design | Automation | Failure diagnosis only | Triage & retest |
|---|---|---|---|---|
| Test strategy and testability | Required | Required | Required | Required |
| Containerization and environment lifecycle | Outside scope | Required | Outside scope | Required |
| Service virtualization and contract fidelity | Outside scope | Required | Outside scope | Required |
| Data, fixtures and domain checks | Outside scope | Required | Outside scope | Required |
| CI execution and diagnostic evidence | Required | Required | Required | Required |
| Shared platform service and team skills | Required | Required | Required | Required |
AI access, permitted context, evaluation, baseline measurement and funding also apply to every workflow. “Outside scope” does not mean a capability is unnecessary elsewhere.
Readiness evidence. A known defect is caught by an independent assertion, and another engineer reproduces the result.
Accountable role. QE lead + application developers
Readiness evidence. Two runs do not interfere, a clean runner recreates the environment, and failed runs leave no unowned resources.
Accountable role. Platform / DevOps owner
Readiness evidence. A deliberate contract mismatch fails a check, and timeout and duplicate scenarios can be replayed.
Accountable role. Integration lead + provider owner
Readiness evidence. The retry scenario rejects a duplicate transfer or a wrong journal entry even if the API reports success.
Accountable role. Domain QA + data / application engineer
Readiness evidence. A failed required check stops promotion; evidence remains accessible after environment cleanup.
Accountable role. Developers + QE automation + DevOps
Readiness evidence. A second squad runs the pattern, troubleshoots a failure and knows who owns the next action.
Accountable role. QE platform product owner + supplier delivery lead
Client assumption. These capabilities are unverified until the client supplies evidence. A roadmap or license purchase is insufficient. Capture each required gap, a named owner, remediation cost and a review date; unresolved execution dependencies hold that scope.
Containerization is an implementation option where suitable. Legacy, mobile and batch testing may need reproducible VMs, devices or reserved integration environments. Service substitutes require owned behavior and separate real-integration checks.
Assess client dependencies → · Modernization sources and technical guidance →
05 / How a generated test earns trust
For PAY-142, the accepted behavior is one transfer and the expected journal and balances. A test that passes a deliberately duplicated transfer is inadequate even if its code coverage is high.
Overview: a proposed candidate-test review and execution boundary.
The tour follows a rejected candidate; acceptance is a separate path.
A separate candidate that meets all agreed checks may enter the maintained suite. Rework alone does not establish acceptance.
06 / Make the trial persuasive
Shared conditions · Comparable tasks, same framework, environment and acceptance rules
Compare accepted outcomes · Include unsuccessful attempts, review, rework and operating costs
ANZ's experiment used self-reported time and excluded unsolved tasks; its quality results also differ by metric. METR's February 2026 update reports selection problems and uncertain effect estimates. Our proposal is to record every assigned attempt, keep unsuccessful work visible and use the same acceptance criteria in both groups. F6, F9
| Measure | What to record | Decision use |
|---|---|---|
| Accepted work | Agreed coverage, correct assertions and reviewer disposition | Compare usable outputs, including rejected attempts. |
| Active effort | Drafting, review, repair, reruns and support in non-overlapping stages | Calculate net effort per accepted task; record setup separately. |
| Quality | Known-defect detection, critical omissions, flaky failures and escapes | Pause expansion when required correctness checks fail. |
| Elapsed time | Queue, environment and provider waits alongside active time | Identify whether the critical path actually shortens. |
| Economic capture | Tool/runtime costs and a named redeployment or budget mechanism | Distinguish usable capacity from cash savings. |
Use client baselines to agree targets before starting. Keep inconclusive results as inconclusive; a small successful pilot supports a next increment, not an enterprise effect-size claim.
07 / A practical client invitation
Requirements ready
Draft scenarios in the current test-management workflow.
ProofAccepted coverage and measured review effort.
Execution repeatable
Generate candidates in the existing framework and CI pipeline.
ProofCorrect behavior passes; a known fault fails.
Diagnostics available
Draft diagnoses alongside the existing disposition process.
ProofFaster correct decisions with traceable evidence.
Manual case writing dominates; requirements and domain reviewers are available.
Assess this workflow’s dependencies → · Plan QE modernization →
A buildable service, stable API tests and repeatable test data already exist.
Assess this workflow’s dependencies → · Plan QE modernization →
Test runs already retain usable logs, traces, assertions and defect dispositions.
Assess this workflow’s dependencies → · Plan QE modernization →
“Let's choose one workflow your team already owns. We'll agree what good looks like, keep the existing tools where they fit, and compare reviewed outcomes and total effort. You'll leave with reusable assets and evidence for the next decision.”
Keep the framework when it fits. The trial can improve the work around it: drafting scenarios, producing reviewed pull requests or assembling evidence. Compare incremental benefit before buying a replacement platform.
Begin with a coached group of domain testers and automation engineers. Domain QA owns expected behavior and acceptance; engineers own executable checks and integration. Reuse the resulting templates and coaching with a second team before wider rollout.
Agree the required quality checks, total-effort measurement and a review date. Continue only when the evidence supports the next scope. If readiness or net benefit is insufficient, retain the baseline, reviewed assets and a specific remediation backlog.
The existing fictional case assumes 75 offshore QA staff, with eight in the pilot cohort. Use weeks 1–2 to establish readiness and the baseline, weeks 3–8 for the bounded payment pilot, and weeks 9–12 for a second-team reuse trial. This is an illustrative planning sequence; change scope or timing when prerequisites are missing.
Commercial structure. Treat QE modernization as a costed dependency workstream with owners, acceptance evidence and review dates. Separate assessment and foundation remediation from pilot delivery, licenses/runtime, coaching and ongoing support. The client approves scope and acceptance conditions before wider adoption; the case's setup allowance does not price a real engagement.
08 / Evidence you can inspect
Reviewed 2026-09-07. Unspecified publication dates remain unknown; retrieval dates do not establish when an outcome occurred. KPMG supplies industry context, not an additional independently measured client case. No licensed analyst report or publisher endorsement is implied. Source register CSV · JSON.
UiPath · Named customer / vendor case
Inspect: Bringing AI-powered software testing to banking teams; Outcomes.
Public full text · Reviewed 2026-09-07Diffblue · Named customer / vendor case
Inspect: Results; Time spent writing unit tests and manual-effort footnote.
Public full text · Reviewed 2026-09-07DBS · Bank-authored operational report
Inspect: Change management; JIRA Assist; AI-powered change risk scoring.
Public full text · 2024 reporting year · Reviewed 2026-09-07Tricentis · Named customer / vendor case
Inspect: Company overview; Embracing the future with AI; Result.
Public full text · Reviewed 2026-09-07Tricentis · Named customer / vendor case
Inspect: Results; Looking to the future.
Public full text · Reviewed 2026-09-07Chatterjee, Liu, Rowland and Hogarth · Bank-authored controlled coding experiment
Inspect: Sections 3, 4.6, 4.9 and 5; PDF pp. 4, 6–7, 11–13.
Public full text · 2024-04-17 · Reviewed 2026-09-07TestingXperts · Anonymous supplier case
Inspect: Summary; AI-Led Test Design Transformation; Impact-Driven Test Prioritization.
Public full text · Reviewed 2026-09-07KPMG · Consultancy synthesis
Inspect: Pages 10–13; used for context, not as an independent outcome study.
Public full text · 2025 · Reviewed 2026-09-07METR · Independent research / methodological update
Inspect: Selection effects; raw estimates and confidence intervals.
Public full text · 2026-02-24 · Reviewed 2026-09-07Diffblue · Product documentation
Inspect: How it works: regression tests capture current behavior.
Public full text · Reviewed 2026-09-07