A few months ago, I was working on a feature at work that went through quite a few rounds of review before it got merged.

At work, unit tests and smoke tests sometimes aren’t enough to get my PRs through. I also attach dev-test results, which is basically proof that I actually ran the thing and it works. I’ve noticed people review my PRs sooner when the dev-test results are clear. Reviewers are human too, they like seeing things work.

The annoying part is the timing. If I dev-test too early, someone asks for changes and I have to test everything again. If I wait until the code settles, the PR just sits there without any evidence and nobody is in a hurry to look at it. It’s a bit of a chicken-and-egg problem.

Most of the time I just live with it, because redoing a quick check is no big deal. But this feature was different. One round of dev-testing meant going through dozens of forms spread over a few different pages, each time with a different mix of options to pick and fields to fill in. I did it once and it was really boring. I definitely didn’t want to do it again after every review.

So I tried to get AI to do it for me.

My first idea was to give an AI agent the Playwright MCP server, describe the steps, and let it click through the app. It kind of worked, but it was slow. Every click is a round trip to the model, and the model has to look at the page and think about what to do next. At one point I watched it spend a long time trying to figure out that it had to close a popup, one of those that only show up sometimes. I would’ve closed it in a second.

It also didn’t do things exactly the same way every time. Even with very detailed instructions, it sometimes took a slightly different path. For a test I want to repeat after every code change, that’s not great. And of course, every run costs tokens.

Part of the reason is the app itself. It looks fine to a person, but the AI reads the HTML, and the HTML is mostly divs and icon buttons with no labels. Some things that sit right next to each other on the screen are in completely different places in the markup.

That’s what made me change my approach. Instead of asking the AI to click for me, I asked it to write a Playwright script that does the clicking. The AI only has to figure out the page once. After that it’s just a script, so it runs the same way every time and it’s way faster.

To help it, I gave it two things the page itself couldn’t tell it. First, a description of what’s on each screen and how to get from one to the next. Second, the frontend source code, so it could pick selectors from the actual markup instead of guessing.

What the AI ended up writing was a set of building-block functions, one for each thing you can do in the app. To give you a feel, imagine it’s an app for managing orders and invoices (it isn’t exactly, but it’s close enough in shape):

Task<Order> CreateOrder(string customer, decimal amount, bool expressShipping, bool giftWrap);
Task AdjustOrder(Order order, decimal byAmount);
Task CancelOrder(Order order, string reason);

Each of these hides the messy parts, like paging through a table to find the right row, or dealing with dropdowns. Once I had them, building a scenario was just a matter of calling them in a different order with different values:

var alice = await CreateOrder("Alice", amount: 120, expressShipping: true, giftWrap: false);
var bob = await CreateOrder("Bob", amount: 80, expressShipping: false, giftWrap: true);
// ...a few dozen more like these

await AdjustOrder(alice, byAmount: -20);
await CancelOrder(bob, reason: "Found it cheaper elsewhere");

It reads pretty much like the checklist I would’ve written by hand anyway. And if I need a different scenario, I just call the same functions differently.

(And yes, I know Playwright has codegen, which records your clicks into a script. But a recording only gives you one path with fixed values. I needed the same flow with lots of different combinations, plus that popup that only shows up sometimes. That’s why building blocks worked much better for me.)

Of course the script didn’t work the first time. This is where the MCP server was actually useful. The AI runs the script, reads the error (for example, “this selector matches 2 elements”), checks the page, fixes the script and runs it again. I mostly just watched.

There are about 1,500 lines of code behind those building blocks, and to be honest, I didn’t read most of it. I could just watch it run in the browser and see if it did the right thing.

Once the script was working, a full round of dev-testing was just one command. When review comments came in, I fixed the code and ran the script again. So testing early wasn’t a problem anymore.

What I like about this is that the change was pretty small. I didn’t switch tools or anything. I just asked the AI for something different: a script, instead of doing the work itself. It still handled the boring parts like selectors and waits, and I got code I can run whenever I want without asking the AI again.

Some time after I’d finished all this, I came across Playwright’s own Test Agents. They overlap quite a bit with what I did. One agent writes the tests, another runs them and fixes the ones that fail. They’re made for building a test suite though, and what I needed was one long dev-test run that builds on itself step by step, so I’m not sure they would’ve been a good fit. Still, it’s nice to see Playwright going in the same direction.

Letting the AI click isn’t bad. For a quick one-off check, it’s probably the easiest option. Which approach makes sense really depends on the situation.

In my case, I knew I’d be running the same test again after every round of review. The AI didn’t know that. So it was never going to suggest writing a script on its own (or maybe it would, after spending half a day on it, who knows). I asked it to click, so it clicked.

The AI and the tools were the same either way. The AI was really good at the tactical part, like finding selectors and fixing broken steps. But the strategy, how to approach the whole thing, came from me. Mostly because the context that mattered was in my head, not in the prompt.

This is just one example, of course. Maybe the AI would’ve picked the same strategy if I’d told it everything I knew. But that’s the tricky part. Who provides the context? Usually me. How would I know I’ve given it all the context it needs, when I’m not always sure which parts matter? And how would the AI know? It doesn’t know what it doesn’t know. I don’t have a good answer to that yet.

For now, I guess the strategy is still on me.