Fast, cheap browser sub-tasks for Claude and other agents: TypeSafe Jev decisions on top of browser-harness - danielnc/jev-browse

1 points•danielnc•6 days ago•1 comment•

1 comment

danielnc6 days ago
Hi HN, I built this after watching Claude Code drive a browser for a mundane job: triaging user feedback on an internal admin page and updating statuses. It worked, but it was slow and expensive. Every click cost a full model turn in which the agent re-read the page and its own context before deciding what to do next.

jev-browse hands that loop to a small decision model instead. Your agent calls one function, fast_run(url, goal), from a browser-harness script. On each step, jev-browse snapshots the page's visible controls in one browser call, and asks TypeSafe's Jev (a "System One" model that returns typed choices with probabilities rather than generated text) which operation to perform and on which element, in a single ~0.3s request. When a field needs text that isn't literally in the goal, it asks a small text model (a local Ollama model, Claude Haiku via the CLI, or any OpenAI-compatible endpoint). The agent gets back an outcome, evidence, and a typed reason whenever jev-browse stops.

It's inspired by browser-use's jev-ultrafast demo, which showed that one typed choice per step is enough to drive real sites. jev-browse turns that into a tool agents use day to day: pluggable text backends with a known-answer health check, an interactive find/click/check mode, a gate before commit-like clicks (send, pay, delete), owned-tab isolation so it never touches your own tabs, and "hand back instead of guess" for iframes, uploads, auth walls and low-confidence steps.

Numbers, on my machine only, with outcomes checked by code rather than a model: on public-site tasks (Wikipedia navigation, a Google Flights search, a filter-and-open listing), Claude with jev-browse was 1.6–3.0× faster and 2.1–5.3× cheaper than the same Claude driving the browser itself, with the same pass rate. Almost all of the saving comes from fewer agent turns. Called from a plain script with no agent, a Wikipedia task takes about 4s and a fraction of a cent. The benchmark code is in the repo if you want to reproduce it or break it.

Caveats: page text is sent to TypeSafe (and to the text model if one is needed), so don't use it on sensitive pages. It doesn't handle canvas-heavy UIs, shadow DOM or cross-origin iframes yet; it hands those back. And "done" is a claim for your agent to verify, not proof.

MIT licensed. Setup is a copy-paste prompt for Claude Code or Codex. I'd love feedback, especially from people running agents against real web apps.

Read the full thread on Hacker News →

Related stories