AI × QE / A quality transformation briefing

From manual QA to AI-assisted delivery.

Start with the work your teams do every day. See where AI can help, what the platform needs, and how to prove the value.

Illustrated workflow connecting requirements, AI assistance, test execution, human review and release
One platform. Every QA workflow.Explore a proposed system for requirements, testing, evidence and human decisions.
30 sources · findings with caveats4 audience decks · vision & architecture1 fintech case · explicit assumptions

A path through the conversation

Start with the audience. Go deeper together.

The fintech story / Illustrative case

75 QA staff.
One payment release.

What happens when the payment provider accepts a transfer, but the customer sees a timeout? Use this familiar problem to explain the platform, the workflows and the human decisions.

Follow AI from requirement to release →
  1. 1See the payment failureRetry, duplicate prevention and reconciliation
  2. 2Break down the QA workManual work → AI assistance → human check
  3. 3Change the delivery assumptionsDevOps, cloud and testing maturity affect the opportunity

The expanded quality mandate

Start with AI-assisted QE. Extend to assurance of AI.

Browse the document library →

Select a capability to inspect its workflow, proof and control.

01 / AI for QEImprove software delivery
Shared foundation Versioned context · independent checks · accountable owners · evidence
02 / QE for AIAssure AI behavior and actions

Test design with an independent quality check

Workflow
Use requirements, API contracts and existing tests to draft candidates.
Proof
Compile, run repeatedly and test relevant fault detection. Track reviewer effort and later rework.
Control
Protect approved assertions and test oracles from silent changes.

Meta's TestGen-LLM and ACH demonstrate filtered generation and mutation-guided testing. [E04][E05]

Diagnosis that gives reviewers inspectable evidence

Workflow
Join failure output, scoped changes and dependency context; propose a cause.
Proof
Measure classification accuracy on a reviewed sample and total time to resolution.
Control
Require supporting evidence and allow the workflow to abstain.

Google's AutoDiagnose is an implementation reference, with distinct evaluation and deployment populations. [E06]

Test maintenance with reproducible evidence

Workflow
Join failure history, test code and fixture versions; draft a repair for a confirmed test defect.
Proof
Reproduce the failure, repeat the repaired test and run relevant regression. Track acceptance, review effort and recurrence.
Control
Preserve domain assertions. Any temporary quarantine needs an owner, expiry and replacement coverage.

FlakyGuard is a deployment reference with distinct reproducibility and repair-acceptance denominators. [E07]

Evaluation of the complete AI application

Workflow
Test representative tasks, rare cases and failure conditions against versioned expectations.
Proof
Score task success, groundedness and policy compliance by scenario; calibrate model judges with humans.
Control
Keep a held-out set and re-evaluate when models, prompts, retrieval or tools change.

NIST supplies the lifecycle framework; evaluation tools help implement the evidence loop. [A01][A02][T03]

Permissions enforced outside the model

Workflow
Route requested actions through an identity-aware policy gateway.
Proof
Test denied actions, credential scope, approval requirements and bounded recovery.
Control
Use short-lived credentials, tool allowlists and separate approval for high-impact actions.

OWASP agentic guidance and OSFI's July 2026 bulletin inform this proposed boundary. [A03][A05][R01]

Production evidence becomes the next regression suite

Workflow
Sample traces, user corrections, incidents and drift; curate new evaluation cases.
Proof
Monitor task success and failures by cohort alongside latency, cost and escalation rates.
Control
Apply privacy-aware retention and trigger a tested fallback when agreed floors fail.

Ongoing measurement connects risk management to release decisions. [A02][R01][T03]

Authored capability map. It is a proposed design, not a measured maturity ranking.

Explore the assurance system.

Inspect the architecture. Test the economics.
Understand the uncertainty.

Pan horizontally to explore the full diagram. A text description is available below.

Proposed enterprise AI and quality engineering platform: delivery interfaces, governed context and agent execution, independent assurance, release and evidence feedback DELIVERY EXPERIENCE Engineer / QE IDE · pull request · test workbench Delivery systems Repository · CI · issue tracker AI applications Product workflows · agents Work enters through existing team workflows GOVERNED EXECUTION Context service Classify + minimize Permitted retrieval Versioned input snapshot QE agent runtime Task plan + model router Prompt / model versions Memory + retry limits Action gateway Agent identity + policy Tool / resource allowlists Budget + approval rules task + scope approved context INDEPENDENT ASSURANCE Sandbox + deterministic checks Build · repeatability · mutation · fixtures Evaluation service Held-out cases · judges · human labels allow application versions Repository text and tool output cannot grant permissions. EVIDENCE + FEEDBACK Evidence store Task / run / artifact / decision IDs Quality · effort · cost · incidents Curated regression corpus Accepted failures + reviewed expectations New cases challenge the next release run traces Release authority Reviewer + existing CI / change gates Approve · hold · rollback The generating agent cannot approve itself checks + evaluation evidence Proposed logical architecture · dashed paths carry context, evidence or feedback
Inspect a component: Select a component in the diagram
Proposed logical architecture · select a component to inspect its role and research basis. [A01][A03][R01][E04]
Research behind this design & diagram description
Finding

NIST frames assurance across the lifecycle. OWASP and OSFI identify agent identity, tool authority and traceability as control concerns; Meta separates generation from validation.

Design implication

Separate context, generation, policy enforcement, independent evaluation and release authority. Preserve a shared evidence record across those boundaries.

Diagram in words

Engineers and delivery systems supply tasks and approved context. The runtime proposes actions through a policy gateway. Independent sandbox checks and application evaluations supply evidence to the release authority. Traces and decisions feed the evidence store; curated failures become regression cases.

Inspect methods and source documents →

$45.0killustrative net economic impact3.3% capacity released · 0.45% of addressable spend

Annual value bridge $ thousands · per $10M addressable QA spend
Base assumptions, illustrative annual values
Capacity equivalent$330000
Uncaptured capacity−$165000
AI and pilot cost−$100000
Quality allowance−$20000
Extra human effort$0
Net economic impact$45000
Net benefit vs. adoption All other current assumptions held fixed · $ thousands

Illustrative assumptions, not a forecast. Capacity is effort equivalent; cash savings need Finance-approved capture.

Activity share 55% · eligibility 60% · AI/pilot cost 1.0% · quality allowance 0.2%.

Formulas, assumptions and accounting treatment ↗

METR estimates · Change in task completion time with AI · Dots show estimates; lines show reported confidence intervals.

Early 2025: 16 developers, 246 issues in familiar repositories. This setting showed a slowdown; it does not represent all software work.

Different cohorts and tool periods; no pooled estimate. The follow-up intervals cross zero and selection bias limits interpretation.

Primary sources: July 2025 RCT ↗ · February 2026 follow-up ↗

Three-dimensional assurance platform with a gateway, verification chambers and independent release authorityArchitecture in motionExplore the platform in 3D ↗Orbit the model. Follow a candidate. See where an unsafe action or failed quality check stops.Four scenarios · interactive demo + short film

Start with your perspective

Visual briefings for
strategy and architecture.

Strategy for executive leaders.
Architecture for the teams who deliver it.

Executive · Guided presentation

Our Banking Client: strategic vision

14 selected slides. Open a full-size presentation with optional audio.

Technical · Guided presentation

Our Banking Client: engineering blueprint

14 selected slides. Open a full-size presentation with optional audio.

Executive · Guided presentation

Quality engineering in the AI era

11 selected slides. Open a full-size presentation with optional audio.

Technical · Guided presentation

AI assurance architecture

11 selected slides. Open a full-size presentation with optional audio.

New to the terminology? Open the AI × QE dictionary → Plain-language definitions, examples and a guide to the architecture components.

Go beneath the slides

A research base you can inspect.

Verification log →