QA The Other WayQuality assurance for the age of AI-written tests. QA The Other Way.
Reliability

A retry budget is not a repair budget

A test that passes on its second attempt still failed once. Keep that fact visible, price the repeated work, and give every flaky test an owner.

6 min readQA The Other Way
A ratchet wrench beside a cracked inspection seal, illustrating that repetition is not repair.
AI-generated editorial illustration.
Illustrative video for this article. English on-screen text; select subtitles for this language.

Consider an illustrative release run: a checkout test fails, runs again, and passes. The dashboard is green. Nobody investigates. The next release repeats the pattern. You have reduced the number of red dashboards without learning whether checkout is reliable.

Playwright distinguishes three outcomes in its retries guide: passed on the first attempt, flaky after a failed attempt followed by success, and failed after the available attempts. Those are different pieces of evidence. A generated suite should preserve the distinction rather than flattening every eventual pass into success. The practical artifact for this article is a retry ledger that survives the final dashboard color.

What a retry actually changes

A retry gives the same test another chance to run. It does not fix a weak assertion, repair a selector, or separate two accounts that overwrite each other's settings. Playwright discards a failed worker process and starts another. Its documented examples show the beforeAll hook running again in the new worker. Setup and cleanup therefore belong in the cost calculation, not only the seconds spent inside the test body.

A new browser is not necessarily a new database. If the first attempt created an order before failing, the next attempt may encounter that order. Review generated tests for external side effects and identify whether cleanup is safe to repeat. Retrying a purchase against production is not an acceptable experiment; use staging data and a bounded test account.

The ledger to keep beside CI

For each test, record its stable identifier, first-attempt result, eventual result, attempt count, evidence link, suspected cause, owner and next review date. Keep the first-attempt failure even if the final attempt passes. A team with no dedicated QA engineer can assign ownership to the developer who owns the flow, rather than letting the automation create an unowned queue.

The ledger is also a useful review surface for agent-written tests. Ask the agent to draft a scenario, then have a human check its setup, assertion and cleanup. A retry policy is a runner setting, not a substitute for that review. AnyTest's public page describes URL-led exploration and a suite humans review; it does not establish how your particular CI retry budget should be chosen.

Put numbers on the repeated work

Use an illustrative calculation with explicit assumptions. Suppose ten tests each need one additional attempt that takes twenty seconds, and restarting setup adds five seconds per attempt. That is ten times twenty-five seconds, or 250 seconds of added execution work. It is not necessarily 250 seconds of wall-clock delay because parallel workers can overlap. Record both runner work and elapsed pipeline time if the distinction matters to your bill or release window.

Then add the time someone spends reading flaky reports. If an agent saves an hour of authoring but leaves an hour of recurring investigation, the authoring number alone does not describe a saving. For an existing QA team, the goal is less repeat work and more time for risk analysis. For a small team without QA, the goal is useful coverage that a developer can actually maintain.

A release rule you can explain

Start with a small bounded retry setting and a visible flaky category. Escalate repeated first-attempt failures instead of silently increasing the limit. The right threshold depends on the product and release risk; there is no universal count that makes a flaky checkout acceptable. Attach a repair owner and evidence before accepting an exception.

The question at review is simple: did the second attempt produce new evidence, or merely a color we preferred? Keep the failed attempt until someone can answer.

Common questions

Is a test that passes on retry a passed test?

Playwright classifies a first-attempt failure followed by a successful retry as flaky, not as a first-attempt pass. Preserve that distinction in release reporting.

How should a team measure retry cost?

Count extra attempts, test execution, repeated setup and review time. Runner work and wall-clock pipeline delay differ when attempts overlap.

Do new Playwright workers reset server data?

A replacement worker and browser do not automatically erase external application state. Design repeatable setup and cleanup for staging data.

Sources