AI × QE / A quality transformation briefing
From manual QA to AI-assisted delivery.
Start with the work your teams do every day. See where AI can help, what the platform needs, and how to prove the value.

For executives and sponsors: the opportunity, target state and leadership choices.
Open executive briefings → 02 / THE STORYMake it concrete75 offshore QA staff. A payment retry. Nine AI-assisted handoffs to inspect.
Explore the client scenario → 03 / ARCHITECTUREShow how it worksFor technical SDs and leads: shared services, toolchains and delivery maturity.
Open architecture briefings → 04 / NEXT STEPScope the opportunityPick a workflow, understand the baseline and agree what a pilot must prove.
Plan the discovery →The fintech story / Illustrative case
75 QA staff.
One payment release.
What happens when the payment provider accepts a transfer, but the customer sees a timeout? Use this familiar problem to explain the platform, the workflows and the human decisions.
Follow AI from requirement to release →- 1See the payment failureRetry, duplicate prevention and reconciliation
- 2Break down the QA workManual work → AI assistance → human check
- 3Change the delivery assumptionsDevOps, cloud and testing maturity affect the opportunity
Select a capability to inspect its workflow, proof and control.
Test design with an independent quality check
- Workflow
- Use requirements, API contracts and existing tests to draft candidates.
- Proof
- Compile, run repeatedly and test relevant fault detection. Track reviewer effort and later rework.
- Control
- Protect approved assertions and test oracles from silent changes.
Meta's TestGen-LLM and ACH demonstrate filtered generation and mutation-guided testing. [E04][E05]
Diagnosis that gives reviewers inspectable evidence
- Workflow
- Join failure output, scoped changes and dependency context; propose a cause.
- Proof
- Measure classification accuracy on a reviewed sample and total time to resolution.
- Control
- Require supporting evidence and allow the workflow to abstain.
Google's AutoDiagnose is an implementation reference, with distinct evaluation and deployment populations. [E06]
Test maintenance with reproducible evidence
- Workflow
- Join failure history, test code and fixture versions; draft a repair for a confirmed test defect.
- Proof
- Reproduce the failure, repeat the repaired test and run relevant regression. Track acceptance, review effort and recurrence.
- Control
- Preserve domain assertions. Any temporary quarantine needs an owner, expiry and replacement coverage.
FlakyGuard is a deployment reference with distinct reproducibility and repair-acceptance denominators. [E07]
Evaluation of the complete AI application
- Workflow
- Test representative tasks, rare cases and failure conditions against versioned expectations.
- Proof
- Score task success, groundedness and policy compliance by scenario; calibrate model judges with humans.
- Control
- Keep a held-out set and re-evaluate when models, prompts, retrieval or tools change.
NIST supplies the lifecycle framework; evaluation tools help implement the evidence loop. [A01][A02][T03]
Permissions enforced outside the model
- Workflow
- Route requested actions through an identity-aware policy gateway.
- Proof
- Test denied actions, credential scope, approval requirements and bounded recovery.
- Control
- Use short-lived credentials, tool allowlists and separate approval for high-impact actions.
OWASP agentic guidance and OSFI's July 2026 bulletin inform this proposed boundary. [A03][A05][R01]
Production evidence becomes the next regression suite
- Workflow
- Sample traces, user corrections, incidents and drift; curate new evaluation cases.
- Proof
- Monitor task success and failures by cohort alongside latency, cost and escalation rates.
- Control
- Apply privacy-aware retention and trigger a tested fallback when agreed floors fail.
Ongoing measurement connects risk management to release decisions. [A02][R01][T03]
Authored capability map. It is a proposed design, not a measured maturity ranking.
Inspect the architecture. Test the economics.
Understand the uncertainty.
Overview · all connections have equal emphasis.
Moving pulses show direction. Play flow follows specific paths through the architecture.
Pan horizontally to explore the full diagram. A text description is available below.
Inspect a component: Select a component in the diagram
Context service. Scope retrieval to permitted sources and preserve the input version. Repository text remains task data, even when it contains instructions. [A03][R01]
QE agent runtime. Record the model and prompt version; bound retries, memory and spend. The runtime proposes work within the gateway’s authority. [A01][R01]
Action gateway. Enforce identity, resource scope and approvals outside the model. An injected instruction cannot grant the agent a production permission. [A03][R01]
Independent checks. Run generated tests in a sandbox. Validate repeatability and fault detection against an approved oracle before accepting a candidate. [E04][E05]
Evaluation service. Compare application versions on held-out cases and denied-action scenarios. Calibrate model judges against reviewed human labels. [A01][T03][T06]
Evidence store. Join the task, run, review and outcome. Measure verification and correction effort alongside generation time and quality. [D02][E01]
Research behind this design & diagram description
NIST frames assurance across the lifecycle. OWASP and OSFI identify agent identity, tool authority and traceability as control concerns; Meta separates generation from validation.
Design implicationSeparate context, generation, policy enforcement, independent evaluation and release authority. Preserve a shared evidence record across those boundaries.
Engineers and delivery systems supply tasks and approved context. The runtime proposes actions through a policy gateway. Independent sandbox checks and application evaluations supply evidence to the release authority. Traces and decisions feed the evidence store; curated failures become regression cases.
Inspect methods and source documents →Scope & cost assumptions
$45.0killustrative net economic impact3.3% capacity released · 0.45% of addressable spend
| Capacity equivalent | $330000 |
|---|---|
| Uncaptured capacity | −$165000 |
| AI and pilot cost | −$100000 |
| Quality allowance | −$20000 |
| Extra human effort | $0 |
| Net economic impact | $45000 |
Illustrative assumptions, not a forecast. Capacity is effort equivalent; cash savings need Finance-approved capture.
Adoption 50% · net task saving 20% · capacity captured 50%.
Activity share 55% · eligibility 60% · AI/pilot cost 1.0% · quality allowance 0.2%.
Formulas, assumptions and accounting treatment ↗METR estimates · Change in task completion time with AI · Dots show estimates; lines show reported confidence intervals.
Early 2025: 16 developers, 246 issues in familiar repositories. This setting showed a slowdown; it does not represent all software work.
Different cohorts and tool periods; no pooled estimate. The follow-up intervals cross zero and selection bias limits interpretation.
Primary sources: July 2025 RCT ↗ · February 2026 follow-up ↗
Architecture in motionExplore the platform in 3D ↗Orbit the model. Follow a candidate. See where an unsafe action or failed quality check stops.Four scenarios · interactive demo + short film
Strategy for executive leaders.
Architecture for the teams who deliver it.
Our Banking Client: strategic vision
14 selected slides. Open a full-size presentation with optional audio.
Technical · Guided presentationOur Banking Client: engineering blueprint
14 selected slides. Open a full-size presentation with optional audio.
Executive · Guided presentationQuality engineering in the AI era
11 selected slides. Open a full-size presentation with optional audio.
Technical · Guided presentationAI assurance architecture
11 selected slides. Open a full-size presentation with optional audio.
New to the terminology? Open the AI × QE dictionary → Plain-language definitions, examples and a guide to the architecture components.
What the research supports
Read the findings, their methods and the limits of each claim.
02 / EconomicsHow capacity becomes value
Separate task efficiency, released capacity and recognized savings.
03 / GovernanceControls built into the workflow
Connect pilot controls to the Canadian banking supervisory context.
04 / DeliveryA measured path to production
Scope, baseline, pilot and validate each capability before scale.