Test Automation

Playwright MCP does not heal your tests. It removes the thing that breaks.

We argued a while back that most self-healing test automation is not healing anything. It fuzzy-matches a selector that has broken, finds something close enough, and carries on — which means the suite stays green while quietly testing a different element than the one it was written for. The failure is not fixed. It is hidden.

Playwright MCP is interesting because it does not answer that argument. It makes it beside the point.

It does not look at your page the way your tests do

Playwright MCP is a Model Context Protocol server that lets a model drive a real browser. The part that matters for testing is how it perceives the page. Its README is direct about it: the server “uses Playwright’s accessibility tree, not pixel-based input,” and needs “no vision models” because it “operates purely on structured data.”

So when an agent calls browser_snapshot, what comes back is not a screenshot and not HTML. It is the accessibility tree — the same structure a screen reader consumes. Every interactive element carries a role, an accessible name, a state, and a reference. The agent then passes that reference to browser_click or browser_type.

Read that again with a self-healing argument in mind. There is no CSS selector. There is no XPath. There is nothing written six months ago that a frontend refactor can invalidate, because the locator is derived at the moment of use from what the page actually exposes.

That is not self-healing. It is not having the wound.

Which moves the problem rather than solving it

The catch is immediate and it is the whole of this article: the agent can only act on what the accessibility tree exposes.

A button implemented as an unlabelled <div> with a click handler has no role and no accessible name. It is not in the tree. The agent cannot see it, cannot reference it, and cannot click it. An icon button with no aria-label appears as a button with no name, which the model has to guess at from position and context — and guessing is the behaviour we objected to in fuzzy selector matching.

You can escape this. Playwright MCP ships a vision capability behind --caps=vision, which adds coordinate tools like browser_mouse_click_xy. It will click your unlabelled div. It will also be clicking a position on a screen, which is the most brittle locator ever devised, and you have travelled a long way to arrive back where you started.

So the honest summary is that accessibility-tree automation converts your accessibility debt into test debt, at a one-to-one rate. Teams that never did the semantic work do not get the benefit. Teams that did get tests with nothing to break.

The non-determinism nobody mentions

There is a second cost, and it is the one that should decide where you use this.

A test that re-derives its target on every run is a test whose behaviour can change without the code changing. That is tolerable when an agent is exploring an application. It is not tolerable in a regression suite, where the entire value proposition is that a failure means something specific changed.

If an agent resolves “the submit button” differently this week than last because the page gained a second button, your suite has not healed. It has silently changed what it asserts, and you will find out when something real gets through.

Determinism is not a nice-to-have in regression testing. It is the product.

Where we would actually use it

Our line on this is narrow and we think defensible.

Use MCP-driven agents for exploration. Walking an unfamiliar application, finding the states nobody documented, reproducing a bug from a vague report. The agent does not need a maintained suite to do any of that, and it is genuinely faster than a person clicking.

Use them to generate locators, then freeze them. Let the agent resolve the element by role and accessible name once, then write that into a deterministic Playwright locator that is checked into the repository and reviewed like code. You get the derivation benefit without handing your regression gate to a model.

Do not put an agent in the critical path of a release. A pipeline gate should fail for one reason only: the application changed. Anything that can fail for a second reason — model variance, a snapshot that read differently, a provider timeout — is not a gate, it is a suggestion.

What this is really telling you

The useful thing about Playwright MCP is not that it automates a browser. Plenty of things do. It is that it has made your accessibility tree load-bearing, and in doing so it has turned a thing most teams treat as a compliance checkbox into testing infrastructure.

Which means the highest-return investment in test stability is not a self-healing library, an AI locator service, or an agent. It is semantic HTML: real buttons, labelled controls, correct roles, accessible names that match what the user reads.

Do that and the agent can see your application, screen reader users can use it, and your selectors stop breaking because you stopped writing them. Skip it and no amount of healing — fuzzy, agentic, or otherwise — gets you a suite you can trust.

We build and maintain suites on this basis as part of our QA and test automation work, and we wrote about the accessibility side of it in WCAG and software testing.

Building something that has to hold up?

Bring the process, not a specification. We will tell you honestly whether an agent is the right answer.