Evidence and economics
External results are useful for selecting hypotheses. A local comparison is needed to establish the value of a particular workflow.
Estimates a future state. Useful for strategic scenarios; cannot establish current productivity.
Describes reported experience or associations. Useful for adoption barriers; selection and measurement bias remain.
Can isolate an effect within its design. Transfer still depends on tasks, people, tools and the environment.
Shows implementation mechanics and local outcomes. Usually lacks a comparable untreated group.
Describe available functions and integration contracts. Do not establish comparative effectiveness.
Apparently conflicting results can both be valid
The field experiments reported by Cui and colleagues found a pooled increase in completed tasks. METR’s early-2025 trial found experienced developers took longer on familiar repositories. The studies differ in task setting, participant experience, tool configuration and outcome definition. Pooling those percentages would create a metric that neither study measured. [E01][E03]
METR estimates · Change in task completion time with AI · Dots show estimates; lines show reported confidence intervals.
Early 2025: 16 developers, 246 issues in familiar repositories. This setting showed a slowdown; it does not represent all software work.
Different cohorts and tool periods; no pooled estimate. The follow-up intervals cross zero and selection bias limits interpretation.
Primary sources: July 2025 RCT ↗ · February 2026 follow-up ↗
The 2026 METR update does not provide a clean before/after series: recruitment selection and time measurement changed. Treat its estimates with the uncertainty and limitations visible in the chart. [E02]
Industrial QE cases show the importance of filtering
Meta’s test-generation work combines candidate generation with checks before developer review. Its mutation-guided work tests whether a candidate detects a relevant modeled fault. Google’s AutoDiagnose illustrates contextual failure diagnosis; Uber’s flaky-test work illustrates how repair depends on understanding the execution context. These are useful engineering patterns rather than universal business-case inputs. [E04][E05][E06][E07]
Valid syntax
Available dependencies
Stable outcomes
Relevant environment
Meaningful assertions
Targeted mutation tests
Maintainer judgment
No weakened oracle
A test that merely builds and passes can still preserve a defect or weaken an assertion. Review test meaning, fault detection and the independence of the expected result.
The value model must include the work around generation
Measure net effort per eligible task, including context preparation, generation, review, correction, repeated execution and control administration. Pair it with escaped defects, false dismissals, rework and elapsed delivery time. Keep active human effort separate from agent runtime and queue time.
Released capacity becomes recognized value through an explicit mechanism: additional delivery, reduced external spend, avoided hiring or another Finance-approved route. An improvement in task throughput is not the same unit as a reduction in labor cost.
Scope & cost assumptions
$45.0killustrative net economic impact3.3% capacity released · 0.45% of addressable spend
| Capacity equivalent | $330000 |
|---|---|
| Uncaptured capacity | −$165000 |
| AI and pilot cost | −$100000 |
| Quality allowance | −$20000 |
| Extra human effort | $0 |
| Net economic impact | $45000 |
Illustrative assumptions, not a forecast. Capacity is effort equivalent; cash savings need Finance-approved capture.
Adoption 50% · net task saving 20% · capacity captured 50%.
Activity share 55% · eligibility 60% · AI/pilot cost 1.0% · quality allowance 0.2%.
Formulas, assumptions and accounting treatment ↗The simulator is an authored scenario model. Its defaults are assumptions, not industry benchmarks. See the full savings model and pilot measurement design for definitions and decision gates.
What the current evidence does not settle
There is no general bank-wide QE savings rate in this review. Evidence remains limited on sustained escaped-defect reduction, long-term test maintainability, simultaneous-agent human effort, rare operational failures and the economics of assurance itself. Those are explicit measurement questions for the pilot and subsequent production validation.
Generated tests: the quality filter in practice
Pan horizontally to explore the full diagram. A text description is available below.
Research behind this design & diagram description
In Meta’s 86-component Kotlin evaluation, 75% of target test classes gained at least one generated test that built; 57% had one that also passed reliably; 25% had one that also increased line coverage.
Design implicationTreat generation as a candidate factory. Require independent checks for stability and added test value; use fault detection where possible because coverage alone does not establish semantic quality.
The bars have a common zero-to-100-percent axis. The qualifying share drops from 75% for build success to 57% for reliable passing and 25% for added coverage. These percentages describe test classes, not individual generated tests.
Inspect methods and source documents →Canonical TestGen claim and economics
75% of target test classes had at least one generated test that built; 57% had one that also passed; 25% had one that also increased coverage. These are cumulative class-level yields, not percentages of individual generated tests. Cohort: 86 Kotlin test classes (31 Stories, 55 Reels), §3.3. In separate Instagram/Facebook test-a-thons, engineers accepted 73% of recommendations; 11.5% of all classes to which the tool was applied were improved. The cohorts and denominators differ. Neither acceptance rate is a predicted result for a bank pilot. [E04]
The canonical scenario model separates exact calculations, additional human effort and Finance-approved cash capture.