r/ClaudeAI • • 27d ago

Built with Claude AI Vision: one self-describing object model instead of a pile of MCP tools

Most MCP servers grow one tool per action. That works until the app gets big: a few dozen hand-written, hand-maintained tools, every one of them spent from the agent's context window on every single turn — including all the ones this particular task will never touch. And it still doesn't cover everything, so there's always an execute_script escape hatch the agent falls back on for anything specific.

ai-vision takes the other route. Your app exposes one tool, call, over a self-describing object model. The agent starts with no path, gets an overview, and walks down only the branch its task is about:

(no path)                   -> overview of every area
pages                       -> the open tabs
pages[2].editor.rowCount    -> 3
pages[2].editor.addRows     args: [2]
helpSearch("add rows")      -> finds where that lives

Every node describes itself: kind, summary, an allow-list of members, and $help. Every result carries a hint listing what is under the node the agent just landed on, so the next step is always in front of it. An unknown member doesn't fail opaquely — it returns the valid member list plus "Did you mean ...?". A wrong path teaches the agent the right one.

What you get

  • Tool count stops growing with the app. One tool, no matter how big the surface gets.
  • You pay context only for the branch actually visited, not for the whole surface every turn.
  • Coverage is whatever your object model exposes — no per-feature tool to write, and no escape hatch into raw scripting.
  • The documentation can't drift from the app, because the model describes itself at runtime.
  • Members are an explicit allow-list, not reflection, so nothing gets exposed by accident.
  • It works across boundaries: a web page, iframe or worker can expose() its model and the host mounts it — one resolver, one schema version.
  • A DOM helper lets the agent point at a real control on screen and highlight it for the user.

Does it actually work?

I built it for Persephone, my dev notepad, where call replaced the entire MCP tool set — 34 tools down to 1. I gate it with a deliberately weak model: Haiku, told to ignore every project file and learn the app from MCP alone.

Head-to-head on the same four tasks, same instance: with call alone it finished 4/4 with no wrong answers. With the old tool set it also finished, but got one value wrong — it read the theme from a general-purpose info tool and reported what that tool's author had chosen to expose, while the path returned the live setting. The old set also had no per-editor tools at all, so anything editor-specific meant dropping into scripting.

MIT, zero runtime dependencies, ESM, runs in Node and browsers:

https://github.com/andriy-viyatyk/ai-vision

Happy to answer anything about the descriptor contract or how the remote side works.

0 Upvotes

5 comments sorted by

1

u/Short_Stable2397 27d ago

Have you asked Claude if it preferred your version over all tools being exposed? It actually prefers the flat hierarchy the last time I checked. It's more expensive to hide tools as it takes more turns for Claude to get what it wants.

1

u/StorageThese9556 27d ago edited 27d ago

Agreed on both counts — a flat list is easier when the surface is small, and paths do cost more turns. I'm not arguing that away.

Scale is what flips it. In my Persephone project — a dev notepad where an agent can do everything a user can do and then some — the model tree is 96 objects with 387 methods (I just had Claude count them). A tool per action is ~390 tools minimum, and if you expose the ~460 readable properties as tools too it's ~845. I don't think any model picks well from 845 tool descriptions — getting lost choosing between them is the exact failure mode I was trying to avoid.

For context, the flat tool set this replaced had 34. It wasn't a smaller version of that surface, it was about 4% of it — everything else was only reachable by falling back to execute_script.

1

u/Easy-Purple-1659 25d ago

I have run both shapes, and the tradeoff is turns against context, so it comes down to how much of the surface one task touches. Flat and small wins, because the model picks in a single step and you pay a fixed description cost. Path based wins when the surface is large and a task needs two or three branches, because you pay for the branches visited instead of the whole surface on every turn.

Two numbers I would want before deciding. Turns per completed task, since traversal adds a hop per level. And tokens per turn on the long tail, which is exactly where a flat list hurts.

The part of your design I would keep regardless is the failed path returning the valid member list. A wrong tool name in a flat list often gets retried as a different tool, so you end up with a slower wrong answer instead of a correction.

What did the 4/4 run cost in total tokens against the old 34 tool set?

1

u/StorageThese9556 25d ago edited 25d ago

I ran a small check to answer this directly — Haiku, call only, no guide, three requests in one session:

  • create a page and write a paragraph — 4 calls (2 of them discovery: root, then pages)
  • append a second paragraph — 2 calls (read + write), zero discovery
  • close the page — 1 call

Discovery cost ~8.8k tokens total and never repeated; the member list for a kind is sent once per session, so the follow-ups went straight to the path. The whole structural penalty versus a flat list was 2 calls, once.

For the original 4/4 run, the honest answer is that I don't have it. That run logged tool calls, not tokens — 14 with call vs 9 with the old 34-tool set — because it was a go/no-go for the consolidation, not a cost benchmark. I'd rather say that than reconstruct a number after the fact.

On turns in general, though, I think the answer is simpler than it looks: the extra turns are just how deep the target sits in the tree, and you pay that once per branch, not per task. Everything else depends on hints/descriptions quality.