Test Automation

Rethinking Self-Healing Test Automation

Self-healing tests sound magical — but most implementations just mask bad selectors. Here is what self-healing actually does, why that is usually the wrong trade, and what we think it should mean instead.

How self-healing test automation works

The mechanics are less mysterious than the marketing. A self-healing locator does roughly this:

  1. At authoring time, the framework records more than the selector you wrote. It captures a fingerprint of the element — tag, id, classes, text content, attributes, position in the DOM, sometimes its neighbours.
  2. At run time, it tries your selector first. If the selector resolves, nothing else happens.
  3. If the selector finds nothing, the framework searches the page for the element whose fingerprint most closely matches the one it recorded, scoring candidates on attribute overlap, text similarity and structural proximity.
  4. If the best candidate scores above a threshold, it acts on that element and logs a repair. The test continues. The suite stays green.

Some products do the scoring with a model rather than a weighted heuristic. That changes the accuracy, not the shape of what is happening.

What that trade actually costs

Read step four again. A test was written to assert something about a specific element. That element could not be found. The framework chose a different element, and reported success.

Three things follow, and each of them is worse than a red build.

A failure that should have been loud became a log line. Selectors break for reasons. Sometimes it is a cosmetic refactor and the heal is correct. Sometimes a developer removed the element and the healer latched onto the one beside it. Both produce the same green tick, and the second one means your regression suite has silently stopped testing the thing it was built to test.

The repair is invisible until it is load-bearing. Heal logs get read for about a fortnight after rollout. After that the suite is green and nobody opens them. Six months later a meaningful share of your assertions are running against elements chosen by a scoring function rather than by a person, and nobody can say which ones.

It removes the pressure that would have fixed the real problem. Selector churn is a symptom. The underlying causes — markup without stable hooks, tests coupled to styling classes, a UI that no one owns the structure of — are real problems with real fixes. A healer makes them painless enough to never address, which means they compound.

The honest framing is that self-healing is not reliability. It is deferral, and the interest rate is high.

The exception worth naming

There is one case where it genuinely earns its place: a suite you have inherited, written against an application you do not control, which you need green this week while you do the real work.

That is triage, and triage is legitimate. The failure is treating the triage measure as the architecture — leaving the healer switched on for two years and calling the result a stable suite.

If you adopt it on those terms, set an expiry date and treat the heal log as a backlog rather than a receipt.

What self-healing should mean

The useful version of the idea is not “find a different element.” It is “make the element impossible to lose.”

That is a different investment, and it is almost entirely unglamorous.

Locate by role and accessible name, not by structure. A button found by its role and the text a user reads survives a refactor that moves it three levels up the DOM. A CSS path does not. This is why getByRole exists and why it should be the default rather than the fallback.

Give the application stable test hooks, and treat them as an interface. A data-testid is a contract between the application and its suite. It belongs in code review like any other interface, and removing one should be as deliberate as changing an API signature.

Never couple a test to styling. A class name exists to apply a visual rule. The day it is renamed by a designer who has never seen your suite is the day you learn how many tests depended on it.

Let the build go red. This is the one teams resist. A broken selector is information: something changed and nobody told the suite. Suppressing that signal to keep a dashboard green trades a small recurring cost for an unbounded one.

How we build it

Our suites are built on role-and-name locators first, explicit test hooks where the semantics genuinely are not available, and no healing layer at all. When a locator breaks, the build goes red, somebody looks at it, and the fix takes minutes because the failure points at exactly one element.

The maintenance burden people expect from this never materialises, for a reason worth stating plainly: most selector churn comes from selectors that were brittle when they were written. Fix the authoring and the churn largely disappears — which means the healer was solving a problem the team created and would otherwise have had to notice.

Where we do want adaptive behaviour is in exploration rather than regression: walking an unfamiliar application, reproducing a vague bug report, finding the states nobody documented. That is a genuinely good use of a model driving a browser, and it has nothing to do with keeping a pipeline green.

We wrote about what changes when an agent addresses the page through the accessibility tree instead of a selector in Playwright MCP does not heal your tests, and this is part of how we approach QA and test automation generally.

A shorter version of this piece was first published on LinkedIn.

Building something that has to hold up?

Bring the process, not a specification. We will tell you honestly whether an agent is the right answer.