AI × QE Briefings Executive · Fintech case

Our Banking Client: strategic vision

Our Banking Client / Illustrative banking scenario / Executive

Our Banking Client: shared QA platform

A payment-release story: 75 offshore QA staff, heavy manual testing and constrained test environments.

75
offshore QA staff

One payment release.
A platform the next team can reuse.

Executive briefing • Strategic vision • Client scenario

Our Banking Client / Illustrative banking scenario / Executive

The QA organization

75 offshore QA staff across five squads; 45 focus on manual and domain testing. Environments and vendor test slots limit parallel work.

Manual & domain testing
45 people
Automation engineering
15 people
Data & environments
8 people
QA leads & coordination
7 people

The pilot draws eight QA staff from this population. More engineers alone do not create more vendor test capacity.

Our Banking Client / Illustrative banking scenario / Executive

QE testing scenario: one timeout, two debits

PAY-142 simulates a CAD 100 transfer accepted before a timeout. The customer retries; an injected duplicate-posting defect tests whether QE catches the second debit.

  1. CustomerSend CAD 100
  2. Payment APIKey: pay-142
  3. ProviderAccept; response times out
  4. Customer retriesReuse pay-142
Injected test defect2 journals

Payer CAD 800
Recipient CAD 200

FAIL · duplicate posting
Expected test outcome1 journal

Payer CAD 900
Recipient CAD 100

PASS · expected state
QE testing scenario. The injected defect causes the duplicate posting; a timeout alone does not imply two debits. Initial balances: CAD 1,000 and CAD 0. One journal contains one debit and one credit. No fees, FX or other activity.

Expected: one transfer, one balanced journal and one customer-visible result. This is a QE test scenario, not a reported client incident.

Our Banking Client / Illustrative banking scenario / Executive

Where the release work goes

Modelled hands-on effort for one bounded release test pack. Eight stages add to 300 hours.

Requirements
20h
Test design
40h
Test data
30h
Environment
25h
Automation
55h
Execution
80h
Triage & retest
35h
Release evidence
15h
300 person-hours for this pack. Each activity belongs to one stage. Baseline already includes its ordinary review and rework.

The pack spans 10 business days. Waiting time and human effort use different units and must be measured separately.

Our Banking Client / Illustrative banking scenario / Executive

The platform investment

Common context, fixtures and evidence serve every squad. Application teams retain business ownership.

1. QA contextStories, tests, defectsVersioned business rules 2. AI workbenchDraft and explainProposed artifacts only 3. Human reviewApprove expected resultsReview test-code changes 4. Test runnersExisting CI + frameworksPinned build and fixtures 5. Run evidenceAssertions, logs, tracesAI drafts failure summaries 6. Release reviewQA + release ownerDefects and exceptions Shared foundation: synthetic data · dependency stubs · artifact IDs · ownership · model configuration
Read the top row left to right, then the lower row right to left. Blue is system context or execution, purple is AI assistance, amber is human review and green is evidence. Arrows show work handoffs.

Existing test frameworks and pipelines execute approved work. AI supports preparation and interpretation.

Our Banking Client / Illustrative banking scenario / Executive

Before / after: separate modernization from AI

01 / Manual QE Copy cases and run by handShared data + scarce vendor slotsReconcile screenshots and logs 02 / Modernized QE Reviewed automated testsIsolated data + virtual servicesCI retains repeatable run evidence 03 / AI-assisted QE Draft cases, code and fixturesSuggest scope, diagnosis and repairSummarize proof for human review
Establish the baselineActive work · vendor waits · coverage gaps
Prove repeatabilityEnvironment readiness · reruns · raw evidence
Isolate added AI valueCandidate acceptance · review · rework
Compare the same payment scope. Modernization removes execution bottlenecks; assess incremental AI value against that repeatable baseline.

Same payment scope. The two target states are proposed; no productivity or quality improvement has been measured.

Our Banking Client / Illustrative banking scenario / Executive

Delivery maturity changes the first move

The same tool creates different work in a shared manual environment and a repeatable delivery system.

01

Shared & manual

Begin with requirements, test design and evidence drafts. Establish stable fixtures and repeatable runs before expanding generation.

02

Mixed maturity

Pilot payment APIs and one web journey. Add synthetic fixtures, provider stubs and evidence links alongside AI assistance.

03

Repeatable delivery

Extend generation and triage across more application teams. Evaluate test selection in shadow mode before changing required coverage.

These profiles illustrate effort, not adoption approval. Review twelve dependency areas before committing to scope, dates or benefits.

Our Banking Client / Illustrative banking scenario / Executive

What 78 hours of gross capacity means

Mixed-maturity target: 300 hours becomes 222, including 40 hours of review.

Baseline
300h
Assisted target
182h work40h

Purple segment = review. Total assisted effort = 222h.

  1. 78hgross capacity
    300 − 222
  2. 66hafter operation
    78 − 12
  3. 33husable capacity
    66 × 50%
  4. 15 packsto recover 480h setup
    rounded up, steady assumptions

Conditional capacity arithmetic; adoption prerequisites are unverified. Tool, cloud and vendor charges are excluded. This is not cash ROI or measured AI impact.

Our Banking Client / Illustrative banking scenario / Executive

Ownership within the offshore model

OwnerAccountability
Shared enablementOwn context adapters, templates, fixtures, runner integration and coaching.
Application QA teamsOwn business scenarios, expected results, execution and defect disposition.
Product and release ownersResolve requirement questions and accept the release evidence.
Delivery managementProvide repository access, overlap hours and an explicit backlog for freed capacity.

Reskill domain testers through paired reviews. Agree supplier incentives and reuse ownership before scaling.

Our Banking Client / Illustrative banking scenario / Executive

Observed time framework for the API pilot

Record actual pilot time from baseline through repeated API runs. Example windows remain subject to readiness.

Observation status: not yet recorded

  1. Planning window · weeks 1–2

    Observe the baseline

    Record API test preparation, active effort, waits and rework for one comparable pack.

    Evidence gate: Reviewed API scope + usable baseline
  2. Planning window · weeks 3–8

    Run the API pilot

    Record setup separately; measure review, execution, triage and retest over repeated packs.

    Evidence gate: Required assertions pass + comparable run evidence
  3. Planning window · weeks 9–12

    Confirm repeatability

    Measure second-team setup, support effort and repeat runs before choosing the next scope.

    Evidence gate: Reusable pattern + evidence-based next step
Record actual dates and elapsed time; track active person-hours, waiting reasons, review and rework separately.

Planning windows are illustrative. Environment, data, virtualization and reviewer readiness determine when each phase can progress.

Our Banking Client / Illustrative banking scenario / Executive

The evidence needed to expand

MeasureExpansion evidence
QualityAll mandatory payment assertions pass; no unexplained skips or weakened checks.
EffortInclude preparation, review, correction, retesting and platform operation.
DeliveryReport elapsed time, waiting reasons and failed-run recovery separately.
ReuseA second squad runs the pattern and maintains it without the pilot authors.

A faster draft is insufficient. Compare tasks of similar scope and retain conventional work as a reference.

Our Banking Client / Illustrative banking scenario / Executive

Brainstorm questions

The next decision: assess one payment workflow with QA, development and platform leads.

  1. “Across the 75 QA staff, where do people spend time and where are they waiting?”
  2. “Which payment journey has stable requirements, observable outcomes and an owner who can join a pilot?”
  3. “Could one platform team remove duplicated work across the squads?”
  4. “If we return capacity, which backlog or release commitment will use it?”

Agree the bounded pilot and its evidence before promising wider rollout. See the discovery brief for owners, commercial scope and timing.

Our Banking Client / Illustrative banking scenario / Executive

Adoption depends on funded foundations

Scope follows proven capabilityInputs + people + evidence + approved AI access + a measured baseline

Drafting

Reviewed scenarios

Approved rules and observable outcomes.

Domain QA accepts the scenario matrix.Assess test design →

Diagnosis only

Evidence-linked triage

Retained failures, known dispositions and usable diagnostics.

A reviewer checks the diagnosis; no retest runs.Assess diagnosis →

Execution

Reviewed test changes

Add repeatable environments, CI, fixtures, provider control and reliable runners.

Independent assertions challenge the candidate.Assess automation →
Authored adoption scopes. Every scope requires named platform ownership and all of its readiness checks; these summaries do not replace the assessment. Reuse across teams needs demonstrated support and reusable evidence. Compare modernization dependencies.

Our Banking Client has not been assessed. The dependency backlog can change scope, cost and timing; the 480-hour setup allowance is not a client quote.

Our Banking Client / Illustrative banking scenario / Executive

Modernization enables wider AI adoption

01Trusted work

Reviewed scenarios, expected outcomes and ownership

AI drafts with domain reviewA reviewer can identify a wrong answer
02Repeatable execution

Reliable suites, isolated data and controlled dependencies

AI drafts tests that run in CIAnother engineer reproduces the result
03Shared QE services

Reusable environments, contracts, artifacts and support

Assistance works across teamsA second team onboards and operates it
04Broader AI execution

Evaluated tools, bounded actions and recovery

Agents execute within approved scopeFailures stop, evidence persists, people can intervene
Proposed capability progression. Reviewed assistance can begin while foundations improve. Broader execution requires evidence for its specific scope. Modernization workstreams and sources

For Our Banking Client, fund repeatable payment tests and shared QE services alongside reviewed AI assistance. Validate this capability roadmap through client discovery.

Our Banking Client / Illustrative banking scenario / Executive

Peer evidence for the payments pilot

Libra Internet Bank

35% less time

Test-creation timeIndex: previous time = 100

UiPath customer case. Normalized from the reported reduction; review effort is unspecified. F1

Goldman Sachs

36% → 72%

Unit-test coverageShare of a selected module (%)

Diffblue customer case. Coverage measures exercised code; it does not establish financial correctness. F2

Fiserv

65% fewer incidents

Major incidentsIndex: previous year = 100

Tricentis modernization case. AI testing was a later pilot; this is not an AI-attributed reduction. F4

Separate outcomes and denominators. Each panel has its own meaning; no combined savings rate is calculated.

Peer results support a trial. Outcomes for Our Banking Client remain unmeasured.

Our Banking Client / Illustrative banking scenario / Executive

Our Banking Client: where to begin

Requirements ready

Onboarding design

Draft scenarios in the current test-management workflow.

Proof

Accepted coverage and measured review effort.

Execution repeatable

Payment API tests

Generate candidates in the existing framework and CI pipeline.

Proof

Correct behavior passes; a known fault fails.

Diagnostics available

Failure triage

Draft diagnoses alongside the existing disposition process.

Proof

Faster correct decisions with traceable evidence.

Proposed entry points. Choose one using the client's constraints; inspect prerequisites and prepare a trial brief.

Choose one entry point, validate its prerequisites and agree how reviewers will judge the result. Wider rollout depends on local evidence.

Our Banking Client / Illustrative banking scenario / Executive

Four layers of AI-assisted QE

  1. 1Assist a person Individual copilotsCopilot drafts a retry test; QA reviews and runs it. Proof to expandUseful, correct drafts after review
  2. 2Assist a workflow Connected workflow assistantsA CI assistant explains failures with links to evidence. Proof to expandTraceable, useful recommendations
  3. 3Delegate a bounded task Bounded QE agentsAn agent drafts a patch and reruns sandbox tests. Proof to expandRepeatable actions within limits
  4. 4Scale proven capability Shared QE AI platformSquads reuse context, evaluations, fixtures and evidence. Proof to expandA second team can adopt and operate it
Shared foundations · start on day oneContext · data · reliable tests · CI · QE modernization · owners
Layer 4 supports layers 1–3. Scale useful assistance without having to delegate more actions.

Authored roadmap. Expand with validated prerequisites and client evidence.

Our Banking Client / Illustrative banking scenario / Executive

Before / after: the shared QE architecture

Before / scenario baseline

Manual QE waits for test capacity

Payments QAWeb / mobile QASettlement QA
↓Manual cases, reruns and evidence↓
Few shared test environments

Backend services are not virtualized

Vendor test slots limit parallel runs

After / pilot target

Squads reuse supported QE services

Payments QAWeb / mobile QASettlement QA
↓Owned interfaces and reusable assets↓
Shared QE platform

Context + fixtures + runners + evidence

AI drafts and explains · people review

QE modernization: isolated runs and controlled dependencies. AI assistance: preparation and interpretation.

Starting point: heavy manual QE and scarce vendor environments. Pilot target: controlled service tests, with real integration still required.

Our Banking Client / Illustrative banking scenario / Executive

Outcomes beyond hours

01

Catch payment faults

Catch every critical seeded fault

Reviewed fault scenarios and independent balance assertions

02

Cover critical journeys

Exercise every critical scenario

AI drafts edge cases; domain QA approves the scenario matrix

03

Explain each release decision

Complete every decision record

CI captures evidence; AI drafts a source-linked summary

Supporting proofTrust repeat runsValidate real integrationsEnable the next squad

Baseline: Not recorded · Observed after: Not recorded · Targets proposed for pilot agreement.

Business aims: fewer escaped payment defects and emergency fixes. These proposed pilot measures do not establish those outcomes.

Our Banking Client / Illustrative banking scenario / Executive

What AI contributes, from requirement to release

01 / Understand the change Clarify retry behaviorPAY-142 · accepted criteria 02 / Design test cases Draft scenarios and boundariesMATRIX-142 · reviewed cases 03 / Draft unit tests Test retry and dedupe logicUNIT-142 · developer review 04 / Build API / UI tests Turn cases into test codePR-142 · reviewed test commit 05 / Prepare data and stubs Draft isolated test inputsFIX-142 + STUB-142 · validated 06 / Select and run tests Recommend affected coverageRUN-142 · runner verdict 07 / Investigate failures Link evidence to a hypothesisDEF-142 · reviewed diagnosis 08 / Repair and retest Propose reviewed changesRERUN-143 · fresh proof 09 / Review release evidence Summarize gaps and retain casesPACK-142 · human decision
Nine steps, one payment scenario. AI drafts or recommends; engineers review; CI executes; release owners decide.

PAY-142 is an authored scenario. Each AI output needs review; test runners execute and people decide.

Our Banking Client / Illustrative banking scenario / Executive

Follow one payment test through the handoffs

01 / Approved intent PAY-142 → AC-142Same-key retry = one transfer 02 / AI candidate Status-only test rejectedPR-142 adds money assertions 03 / Challenge the test RUN-142 · injected defectExpected 1 journal; observed 2 04 / Investigate and repair DEF-142 links the evidenceHuman approves the proposed fix 05 / Fresh retest RERUN-143 · retained evidenceCorrect build passes; mutant fails 06 / Release review PACK-142 links both runsOwner checks remaining gates
An illustrative failure-to-fix trace. A deliberately faulty build must fail; a passing retest alone is not release approval.

Illustrative artifacts and outcomes. Preserve failed evidence; a passing retest alone does not approve release.

Slide notes

Dictionary · terms & acronyms ↗ (opens in a new tab)