QA The Other WayQuality assurance for the age of AI-written tests. QA The Other Way.
Strategy

What agentic QA actually saves, and what it costs

The honest ledger has three lines: authoring time you no longer spend, review time you now must spend, and repair time for the tests that rot. The savings are real when the first line beats the other two.

10 min readQA The Other Way
A stopwatch, illustrating the cost of review and repair.
AI-generated editorial illustration.
Illustrative video for this article. English on-screen text; select subtitles for this language.

Every pitch for AI testing quotes the hours you will save. Almost none of them itemize the hours you will spend. This article does the arithmetic slowly, with assumptions printed next to every number, because the only thing worse than no automation is automation whose cost you never measured.

For a team that already has QA engineers

The repetitive half of the job is regression: writing the same login-checkout-settings flows and then maintaining them every time the UI moves. Agents attack exactly that half. They explore the app, draft the flows, and re-draft them when things change. When the drafts are useful, QA engineers can spend less time typing repetitive flows and more on review, risk analysis and exploratory testing. That is a possible capacity gain to measure, not a promised change in every team.

Run the ledger per sprint. Take the hours your team previously spent writing and repairing regression tests, call it A. Take the hours they now spend reviewing agent-written tests, call it R, and the hours repairing the drafts that came back wrong, call it F. Net capacity gained is A minus R minus F. The automation wins when A is bigger than the sum of the other two, Two factors are worth checking in your pilot: whether the review surface makes intent clear, and whether the drafts need extensive repair. A readable step timeline may help your team; measure review time instead of assuming it beats a code diff.

Notice what did not appear in the formula: headcount. The realistic outcome for a QA team is not fewer people. It is the same people covering more of the product, spending their judgment on risk and edge cases instead of on the fourth rewrite of the checkout flow this quarter.

For a company with no QA engineers at all

The comparison here is not agent versus QA hire. It is agent-written tests reviewed by a developer versus the current state, which is usually no tests and a hope. That changes the question completely. The agent's drafts do not compete with a professional's tests; they compete with nothing, and nothing is expensive in a way that never appears on a budget line: the broken login a customer finds before you do, the release you delayed because nobody trusted it.

The cost side is honest too. A founder or developer now spends R hours per cycle reviewing tests instead of writing them. That time is real. It may be smaller than authoring from scratch, but that is a pilot result to measure. Familiarity with pull-request review helps, though test-oracle review is a separate skill. Break-even depends on review and repair time, tool and execution costs, and the expected cost of escaped defects. Catching one bug does not automatically prove the investment paid off.

The discount rate on every AI claim

Three failure modes eat the savings, and each maps to something testable. Wrong oracles: tests that assert on the toast instead of the outcome, covered in A green checkmark is not proof. Brittle locators: drafts that photograph the DOM instead of naming controls, covered in Your locator is a contract. Unsafe exploration: an agent reading hostile pages without a fence, covered in An agent with a browser needs a fence. A vendor that cannot tell you how its review surface, its selector quality, and its run boundaries work is asking you to take the savings on faith.

The worksheet

Write down, for your last four weeks: hours spent authoring regression tests, hours repairing them, regressions that escaped to users, and what one escaped bug cost you in real terms. Then pilot an agent on one critical flow for two weeks and record review hours and repair hours. Include tool fees, compute and environment setup when converting hours into a money estimate; the capacity formula alone is not a cash-savings formula. The tools that make this loop cheap are explicit about the shape: AnyTest takes a URL, its agents explore and draft the end-to-end tests, and a human reviews each step before the suite is kept. If your pilot's R plus F is clearly below your old A, you have your answer, in your own numbers, which are the only ones worth trusting.

Nobody else's benchmark will make this decision for you. The arithmetic will.

Common questions

Does agentic QA replace QA engineers?

Not by itself. It can help with repetitive authoring and maintenance, while QA engineers still review intent, risk and edge cases. AnyTest states in its own FAQ that it complements QA rather than replacing judgment.

How do I measure whether AI-written tests save my team time?

Track authoring hours before, then review hours plus repair hours after. Net gain is authoring minus review minus repair. Measure it on your own flows for at least two weeks; vendor benchmarks are not evidence about your product.

Is agentic QA worth it for a company with no QA team?

The comparison is against having no tests at all. An agent drafts coverage and a developer reviews it on their schedule. Whether it pays depends on review, repair, tool and execution costs compared with the value of earlier defect detection. Measure those in a bounded pilot.

Sources