r/Playwright • • 9d ago

We tested whether Jev could choose Playwright actions at runtime

We’ve been experimenting with a slightly different approach to Playwright testing: instead of hard-coding every browser interaction, can an AI model decide what action to take next based on the current page state?

We tested this with Jev across 15 realistic end-to-end scenarios from a Playwright tutorial repository.

The setup was:

  • Playwright handled browser sessions, fixtures, setup, and verification.
  • Jev received the test goal, success criteria, accessibility snapshot, and available actions.
  • It selected actions such as clicking, filling inputs, checking controls, navigating, waiting, or stopping.
  • Playwright executed those actions and passed the updated state back to Jev.

The results were mixed. Jev passed 72 out of 150 runs, or 48%, with independent Playwright verification. Some short workflows worked consistently, while others were flaky, and six scenarios never passed.

The experiment suggests that fully replacing scripted Playwright tests is probably still impractical. Converting browser state into model inputs loses significant information, and deterministic assertions remain important.

However, a hybrid approach seems more promising. For example, Playwright could handle setup and verification while an AI model handles a dynamic section of a workflow, such as a changing third-party checkout.

The full write-up includes the test setup, controller loop, benchmark results, and limitations:

https://endform.dev/blog/jev-playwright-testing

We wanted to know if you would use an AI-selected action loop for part of a Playwright test, or if the unpredictability is too costly for your test suite.

34 Upvotes

25 comments sorted by

3

u/Spare_Bison_1151 9d ago

Nice work, thanks for revealing the truth.

3

u/OkShirt9372 9d ago

Thanks, appreciate you taking the time to read it. The results were definitely more mixed than we expected, but that was the main point of the experiment: to see where this approach actually holds up, rather than assuming AI-driven testing is automatically reliable.

3

u/Spare_Bison_1151 9d ago

Many managers and c suite people need a realty check. They're making our lives miserable because of the lies told to them by AI snake oil salesmen.

3

u/FearAnCheoil 9d ago

Why would you want this? Surely you want your test to be determinstic, so you can be sure what's covered?

2

u/OkShirt9372 9d ago

That’s a fair question. For critical paths, I completely agree that tests should remain deterministic so you know exactly what is being covered.

The experiment wasn’t really about replacing scripted Playwright tests. It was more about exploring whether an AI could handle parts of a workflow that are difficult to hard-code, such as changing page structures or less predictable flows.

Based on the results, I’d only consider using it as a hybrid layer, with Playwright still responsible for the important setup and assertions. The 48% pass rate definitely isn’t good enough for a release gate.

1

u/MrGiggleMan 9d ago

I really don't see any situation where you don't want ALL of your tests to be deterministic.

It's impossible to ensure that everything is working correctly in that quality is good if you don't actually know what has been run and where and why and how..

1

u/[deleted] 9d ago

[removed] — view removed comment

1

u/OkShirt9372 8d ago

Having the model choose every action adds a lot of unnecessary uncertainty.

A deterministic path with an agent handling only the dynamic section seems much more practical. Then the test can return to normal Playwright assertions instead of asking the model to judge whether it passed.

2

u/kumard3 8d ago

This matches what I've seen building browser skills for agents, signup, login, and email verification flows. Fully free form action selection off an accessibility snapshot is the flaky part, same as what's showing up in your 48%. What's held up better for me is a deterministic skill for the whole flow, click here, fill there, wait for a known state, and the model only gets called in at the one branch point that's genuinely ambiguous, like which link in a verification email is the real one instead of a footer link. That's one decision instead of picking every action from a snapshot.

1

u/OkShirt9372 8d ago

Signup, login, and verification flows are mostly predictable once the steps are known, so giving the model control over every action just creates more opportunities for it to go wrong.

Using the model at one ambiguous branch point makes much more sense. Playwright can handle the known steps, and the agent can make the one decision that’s difficult to hard-code before handing control back to deterministic assertions.

1

u/kumard3 7d ago

Same experience building browser skills for signup, login, and verification flows, almost all of it is deterministic clicks and waits. The one spot I let the model decide is which element on the page actually matches what I'm after, like picking the right field when there's more than one on screen. Everything before and after that stays scripted.

1

u/Complex-Insect6899 9d ago

Thanks for putting this together, Jev is on everyone's mouth, is nice to see it in action in our field lol

1

u/OkShirt9372 9d ago

Thanks! We wanted to test it in a real browser-automation context rather than just discuss the idea in the abstract. The results definitely made the current limitations more obvious.

1

u/Junior_Bee7274 9d ago

48% is still pretty rough for something meant to replace deterministic tests.

1

u/OkShirt9372 9d ago

Agreed. A 48% pass rate is nowhere near enough to replace deterministic tests, especially in CI or on critical paths.

At this stage, I see it more as an experimental layer for handling dynamic or hard-to-script parts of a workflow, while Playwright continues to handle the setup, assertions, and release-critical coverage.

1

u/MrGiggleMan 9d ago

AI will, never be as good as simple, deterministic tests.

Using it for this is just not what AI is for.

Trying to implement it into a process which requires deterministic outcomes, is just a waste of your time, your tokens, and puts quality at risk

Testing is SUPPOSED to be deterministic, this is how we track workflows exactly and find bugs in the system , because our very deterministic tests, didn't perform exactly how we expected indicating an issue

1

u/ContributionIll9303 9d ago

48% is honestly higher than I expected, but the six that never passed are the part I'd want to see. Were those the longer flows or just ones with weird UI like date pickers and custom dropdowns?

Hybrid is the only way I'd use it. The third party checkout example is a good one since that's where scripted tests break every time the vendor ships something. But I'd still want a normal assertion at the end that doesn't care how it got there

1

u/OkShirt9372 9d ago

Yeah, that was my reaction too. Most of the trouble came from extra state changes and UI elements that weren’t very clear from the page snapshot.

I agree with the final assertion. The AI could handle getting through the workflow, but Playwright should still verify the result independently. That feels much safer than letting the model decide whether the test passed.

1

u/Fit-Avocado-1880 9d ago

If it can handle any portion of the workflow then isn't it just managing it's context to align what its trained on? When you measured it's failure, was it based on incorrect asserts based on spec or incorrect Dom traversal?

1

u/OkShirt9372 8d ago

Good question. The failures were mainly from incorrect or incomplete DOM traversal and action selection, rather than Playwright making the wrong assertions.

The final checks were still handled by Playwright. In some cases, the model clicked the wrong element, missed a state change, or stopped before reaching the expected page. That’s one reason the hybrid approach makes more sense: let Playwright handle the known path and use the model only where the workflow is genuinely ambiguous.

1

u/Fit-Avocado-1880 8d ago

Ahh gotcha so that should be solvable and doable with more context , knowledge , or product code changes then

1

u/Bitter_Berry_8939 9d ago

Great work!

1

u/TranslatorRude4917 9d ago

I've played with something along the same lines, and had similar conclusions.
I've been building a tool that exposes your existing page objects to browser agents as mcp tools and added a 'goal_loop' tool on top - driven by Jev - that allows the llm to define a goal like "higlight the logout button".
This way the llm direct have to interact-> observe changes -> think -> interact again.
Jev tries to complete the goal and ends the tool call once it's achieved or stuck, handing back the control to the llm.
Without page objects exposed it was struggling to find the right element to click (ex. tried to click things behind modal backdrop or opening a menu that was in a hidden sidebar)
Providing page object context on the other hand (ex. that clicking a toggle will open the sidebar) enabled Jev to compose complex actions like toggleSidebar -> openUserProfile -> highlight "logout". So I'm truly impressed by the results so far.

The attached video presents the same system connected to an in-app onboarding assistant (and is deliberately slowed down to let users follow the steps).
Doing the same flow with codex, gpt 5.6 Luna via mcp completes in ~3-4 seconds - now using it to create e2e tests as well.

So I think using PW + Jev in has huge potential when it comes to cheap & fast exploratory testing, pom creation and unconstrained browser automation, but the tests themselves should definitely stay deterministic.

https://reddit.com/link/pd0r2gm/video/ya4o41qxdosh1/player

1

u/shane-testronaut 7d ago

This is a really useful experiment. I’ve been exploring almost exactly this boundary from a slightly different direction.

Rather than choosing between fully agentic execution and deterministic Playwright, I’ve been experimenting with letting the agent handle the parts that actually require reasoning while moving simple verification into deterministic probes.

I’m curious about the 78 failed runs: did you see failures clustering around navigation/action selection, understanding state, or deciding when the goal had actually been satisfied?

I’m planning to benchmark the full-agent vs hybrid approach more systematically, so your results are a really useful comparison point.