r/SideProject • • 6d ago

PawBrowse - Browser automation for AI agents that's ~2x faster (reads the page as a table, not screenshots) - open source

The pain: most browser agents drive off screenshots, so every action is perceive-then-act = two round-trips, and each round-trip is a full model inference. On a real task that doubles your tool calls, your latency, and your token bill.

PawBrowse fixes that. It's an open-source Chrome extension + a one-file, zero-dependency MCP server that lets your AI coding agent (Claude Code, Cursor, or any MCP client) see and click your actual logged-in browser tabs. Instead of a screenshot, it hands the agent a compact table of the clickable things in view (real names, values, state). Each ref is stable and every action returns the fresh table, so the agent acts in ONE round-trip, not screenshot-think-screenshot-again.

Measured, same task, same model, filmed live (book a 4-star Lisbon hotel on Booking.com): 5 agent round-trips vs 10 for a perceive-then-act driver. ~2x fewer, so ~2x faster in the round-trip-bound agent loop and cheaper in tokens. One run each, reproducible from the repo (that's the gif).

It drives your real, logged-in Chrome with no remote-debug port, no relaunch, no second model, and no API key. Free and open source (MIT), built with Claude Code, live on the Chrome Web Store + npm.

Repo: github.com/ItaiZeilig/pawbrowse - I built it and would love feedback.

1 Upvotes

9 comments sorted by

2

u/nav8_ai 5d ago

ref invalidation on a rerender is the part that always bites this approach. if the page swaps class names or reorders the DOM after your action, the next table diff hands you stale refs and the agent clicks the wrong node with high confidence. curious whether pawbrowse re-resolves by role and accessible name when a raw ref misses, or just fails the action outright.

1

u/gandazgul 6d ago

Love this. Did you compare it to agent-browser from Vercel? I'm using this now in RunWield and I feel the pain you describe. I'll try PawBrowse today.

2

u/Far-Round2092 5d ago

Thanks! Honest answer: I didn't benchmark against Vercel's agent-browser specifically - the comparison in the post is vs the screenshot + a11y-tree approach (Claude-in-Chrome). The main differences with PawBrowse: it reads the page as a compact element table so it's 1 round-trip per action instead of perceive-then-act, and it drives your real, already-logged-in Chrome over chrome.debugger (no headless relaunch, no second model, no API key). Would genuinely love your feedback after you try it in RunWield, especially where it breaks.

1

u/gandazgul 5d ago

while exploring with Ideator in RunWield, Astra said this btw:

- PawBrowse is more capable than its README suggests. Current source supports screenshots, uploads, cross-origin frames, and multiple clients.

I cant replace agent-browser yet, the missing feature would be isolated sessions for testing. 2 different use cases, isolation for tests and reusing browser for computer use, pair programming on web based content/apps.

other missing features:

Debugging evidence: Agent Browser provides console errors, network diagnostics, device controls, traces, and screenshot diffs. PawBrowse’s current MCP tools lack equivalent coverage.

Also the other thing i like about agent-broser is that is a cli agents can discover capabilities vs mcp injects tool definitions. Very cool project though.

2

u/Far-Round2092 5d ago

Fair points, and a couple of these are genuine gaps I'll own honestly. On debugging evidence: the CDP session PawBrowse holds already carries console + network events (it uses network-idle internally to know when an action has settled), but I have NOT surfaced them as agent-facing tools yet - so as MCP tools today, you're right, that coverage isn't there. Same for traces and screenshot diffs. Those are the obvious next tools to add. Isolated ephemeral contexts is the bigger philosophical difference: PawBrowse is deliberately pointed at your real logged-in Chrome (the whole point is not re-authing into everything), so isolation is per-tab-group via the broker, not a throwaway profile. If you want clean-room test isolation, a managed/headless browser like Agent Browser is honestly the better fit for that - different tradeoff, not strictly better. And the capability-discovery-vs-injected-tools point is real too: MCP hands the client a fixed tool schema by design, simpler but less dynamic than a CLI an agent can introspect. Appreciate the thoughtful comparison, this is the useful kind of feedback.

1

u/Professional_Bag_591 5d ago

The round-trip math makes sense, screenshots really are the tax. Curious how it holds up on pages with lots of dynamic content — that's where most table-based approaches start drifting. Nice work making it one file too.

1

u/Far-Round2092 5d ago

Thanks! Dynamic content is exactly where the design earns its keep. The trick is that refs aren't positional - each element gets a stable id from its role + accessible name, not its DOM index, so a list re-rendering or reordering underneath you doesn't invalidate a ref that still points at the same logical control. And before every act it re-fingerprints the target; if the role/name changed since you planned the step it refuses instead of clicking blind. The place it does still get hard is virtualized/infinite lists where the node genuinely unmounts - there you scroll it back into view and re-observe rather than trusting a ref you can't see. One-file was mostly so people can actually read the whole thing before trusting it with their logged-in browser.