QA The Other WayQuality assurance for the age of AI-written tests. QA The Other Way.
Test Design

A green checkmark is not proof

AI agents can now write your end-to-end tests. They inherit the oldest bug in testing: asserting on what is easy to see instead of what must be true. The fix is an independent oracle.

8 min readQA The Other Way
A receipt beside a caliper, illustrating independent measurement.
AI-generated editorial illustration.
Illustrative video for this article. English on-screen text; select subtitles for this language.

A test clicks Submit. A toast appears saying "Saved." The assertion checks the toast. The test goes green. Three weeks later a customer reports that their settings never actually persist.

Nothing about this story is new, and nothing about it requires AI. It is the standard way end-to-end tests lie: the assertion watches the user interface, and the user interface is the thing under test. When an AI agent writes the tests, the same failure mode arrives at industrial scale, because the agent also learns what to assert by watching the interface.

The oracle problem, in one paragraph

Every test has two halves. The first half does something. The second half decides whether the result is right. That second half is the oracle, and it is the half that matters. A weak oracle checks the nearest visible signal: a toast, a spinner stopping, a URL change. A strong oracle checks the outcome from an independent direction: the record in the database, the response of the API the page calls, the state after a fresh reload.

Playwright's own best-practices guide makes the adjacent point about actions: tests should verify behavior the end user can see, and avoid coupling to implementation details like CSS classes. The same principle applied to assertions gives you the rule for the AI era. Assert on the outcome a user would check, through a path the bug cannot fake.

Why AI-written tests drift toward weak oracles

An agent that explores your app and writes tests learns the app from its surface. It clicks, it watches what changes, and it encodes exactly that: click this, then this changes. A toast is a convenient, deterministic-looking signal, so the agent asserts on the toast. The agent is not careless. It simply has no access to your intent, only to your pixels.

Playwright's auto-waiting makes this feel safer than it is. Actionability checks confirm an element is visible, stable, and receiving events before the click lands. Auto-retrying assertions wait for the expected condition. Both reduce timing-related failures. Neither asks whether the condition was the right one. A test can be perfectly stable and completely wrong.

Three oracle upgrades that work today

Reload and read. After the save, navigate away and back, or reload, and assert the value is still there. This checks persistence the way a user would notice its absence.

Assert one layer down. Playwright can wait for the network response the UI depends on. Pair the UI assertion with expect(response).toBeOK() on the API call that actually saves, or query the API directly after the UI flow and compare the stored value.

Check the side effect out of band. For a signup, assert on the outbox, the audit log, or the admin list, not the welcome banner. The banner is designed to appear; the record exists only if the system worked.

The review question that changes everything

If you let agents write tests and humans review them, spend review time on one question per test: what does this assert, and could it pass while the feature is broken? Reviewing the happy path is reading the test. Reviewing the oracle is testing the test.

This is also the honest way to use a tool like AnyTest. Its agents explore your app from a URL, build end-to-end tests, and leave them for your review. The review surface shows each step and its outcome. Used well, that review is where the oracle check happens: not "did the agent click the right things" but "did the test it wrote prove anything". The vendor's own material says humans still decide. This is the decision that matters.

A green checkmark tells you the test ran. It never tells you the test was right.

Common questions

What is a test oracle?

The part of a test that decides pass or fail. An assertion on a success message is a weak oracle. An assertion on independently verified state, like data after a reload or an API query, is a strong one.

Does a passing end-to-end test prove the feature works?

No. It proves the specific conditions the test asserted were true. If the assertion watches only the UI, the feature can be broken underneath while the test stays green.

How do I audit tests written by an AI agent?

For each test ask what it asserts and whether it could pass while the feature is broken. Prefer tests that verify persisted state through a second path: reload, API query, or an out-of-band record.

Sources