Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .cursor/rules/project.mdc
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,10 @@ Agent ─stdio─ thin MCP client (index.ts) ─IPC─► daemon (daemon.ts, own
- `autoReSnapshot`: on ref failure (virtualized feed), snapshot the tab and embed `freshRefs` in the response so the agent retries in one step. Do NOT auto-retry the click (non-idempotent). `browser_scroll` returns `refsMayBeStale: true` as a hint.
- `isNew`: each snapshot tracks `role|name` fingerprints per tab; elements appearing since the last snapshot are tagged `isNew: true` so the agent can focus on what changed (big token saver after an overlay/dropdown opens). State lives in `lastSnapshotFingerprints` Map, cleared on tab close.

## Native accessibility tree (`browser_snapshot source:"native"`)

`handlers/ax-snapshot.js` + pure `lib/ax-native.js`. CDP `Accessibility.getFullAXTree` per frame (main + same-process child frames, grafted under their `Iframe` node) → same tree shape as the DOM snapshot. Refs are bound into the SAME page registry (`__bcDom.registry`) by structural path: main world `DOM.resolveNode` + `Runtime.callFunctionOn(pathOfTarget)` → isolated world walks the path (`-1` = enter shadow root, `-2` = enter iframe document, else element-child index), verifies the tag, registers, stores the smart-selector fallback. No DOM mutation. Gotchas: batch `callFunctionOn` PER FRAME (Chrome rejects mixed JS contexts); AX ids repeat across frames (namespaced on merge); user-agent shadow nodes (e.g. a date input's sub-fields) are unreachable: no ref, and omitted from the compact tree (`unreachableNodes`); scoping resolves in the page, then finds the node in a `DOM.getDocument({pierce:true})` dump with the same path grammar. CDP failure → DOM snapshot + `nativeUnavailable`. Default `source:"dom"` never touches the debugger.

## Concurrency (extension/lib/tab-concurrency.js — pure, unit-tested)

- `TabMutexMap`: same-tab serializes, cross-tab parallelizes.
Expand Down
24 changes: 23 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@ It already has your browser open right there. It just can't see it.
- **Batches.** `browser_batch` runs a list of tool calls in one round-trip and stops at the first failure — a click → type → Tab → wait → read sequence is one call instead of five.
- **Console-style JavaScript.** `browser_evaluate` accepts code as you'd type it in DevTools: top-level `await`, several statements, the last expression's value is returned, DOM nodes come back as readable descriptions — and page CSP doesn't block it.
- **Refs that stay right.** Snapshot/find refs resolve through one shared page runtime: the registered element first, then the first *visible* selector match (across open and closed shadow roots and same-origin iframes), then a verified fallback that only re-binds when role, tag and name identify one element — an ambiguous match is reported as gone instead of clicked. Hidden duplicates are skipped.
- **Native Chrome accessibility tree.** `browser_snapshot { source: "native" }` reads the tree Chrome itself computes (CDP `Accessibility.getFullAXTree`): exact roles, accessible names (`aria-labelledby`, `<label>`, native widget semantics), states (`checked`, `expanded`, `invalid`, heading `level`…) and `aria-hidden`/`inert` exclusion — what a screen reader sees. Its refs work with every ref tool, including inside closed shadow roots and same-origin iframes. Opt-in; the default `source: "dom"` needs no debugger.
- **Shadow DOM everywhere.** `snapshot`, `text`, `find`, `click_text`, `wait` and every locator see web components (open and closed roots, slots), so sites like caniuse read like any other page.
- **Frozen tabs don't freeze the agent.** Every page call has an 8 s budget; a tab that stops answering is reported as `TAB_WEDGED` in seconds, later calls fail fast after a 1.5 s probe, and `browser_navigate` / `browser_tabs reload` replace the frozen tab in place (the result carries the new `tabId`).
- **Coordinates when you need them.** Click, hover and wheel-scroll at `x`/`y`, triple-click, ctrl/shift-click, key sequences with `repeat`, and zoomed `region` screenshots that tell you how image pixels map to those coordinates.
Expand Down Expand Up @@ -246,6 +247,27 @@ The model is **tab-first**: the agent always says _which_ tab to act on. It neve
If a ref is stale but the element still exists, it's found automatically via a robust selector + text/role scan (response carries `via: "fallback"`). If the element was scrolled away entirely (virtualized feeds), the response carries **`freshRefs: [...]`** with a fresh snapshot inline — retry with one of those new refs in the same step, no separate snapshot needed.
4. **Verify** — snapshot or read text again after the action.

### Native accessibility tree

`browser_snapshot` has two sources. Both return the same shape (`ref`, `role`, `name`, `value`, state flags, `href`, `children`) and their refs are interchangeable with every ref tool.

| | `source: "dom"` (default) | `source: "native"` |
|---|---|---|
| Built from | A walk of the DOM with ARIA rules re-implemented in the extension | Chrome's accessibility engine, via CDP `Accessibility.getFullAXTree` |
| Roles / names / states | Approximation (explicit `role`, tag map, `aria-label`, `<label>`, text) | Exactly what assistive technology gets: name computation, native widget roles, `checked` / `expanded` / `invalid` / heading `level`, `aria-hidden` and `inert` honoured |
| Debugger | Not used (no banner) | Attached (yellow "being debugged" banner, same as trusted input) |
| Cost | ~tens of ms on typical pages | Same, plus Chrome's tree computation: roughly 2 s per 40k accessibility nodes on a very large page |
| Custom clickable `<div tabindex>` | Listed | Listed only when it has an accessible name (Chrome calls it `generic`) |
| If unavailable | — | Falls back to the DOM tree and adds `nativeUnavailable: "<reason>"` |

```
browser_snapshot { tabId: 15, source: "native" }
→ { source: "native", tree: [ { ref: "s4k2-3", role: "textbox", name: "Search query", value: "abc", required: true }, … ] }
browser_snapshot { tabId: 15, source: "native", selector: "form", compact: false } // scoped, full tree incl. text
```

How refs are bound: every accessibility node carries a `backendDOMNodeId`; the extension resolves it to its element (closed shadow roots and same-origin iframes included) and registers it in the same page-side ref registry the DOM snapshot uses — no attribute or other mutation of the page. Because of that, `click` / `type` / `hover` / `select` / `scroll` / `drag` / `fill_form` act on the exact element (two buttons both named "Save" stay distinct) and keep the stale-ref fallback and `isNew` behaviour. Limits: at most 1500 refs per snapshot (`refLimited: true` when hit); cross-origin (out-of-process) iframes are reported in `skippedFrames`; controls with no page element to act on — Chrome-internal parts such as the sub-fields and picker button inside a date input, or an element that vanished mid-snapshot — are left out of the compact tree and listed without a `ref` in the full tree, counted in `unreachableNodes`.

### Safe Observe → Act workflow

For automation that must fail safely when a page changes, use the Browser Controller 2.0 agent API. `browser_observe` captures one compact semantic state in a single page execution and returns session-, tab-, document-, and snapshot-owned refs:
Expand Down Expand Up @@ -341,7 +363,7 @@ See [`agent-config/`](agent-config/) for manual installation or to customize the
| Tool | What it does |
|------|-------------|
| `browser_observe` | Compact atomic semantic observation with snapshot/document identity, geometry, state, and dynamic allowed actions |
| `browser_snapshot` | Accessibility tree with element refs. Compact mode (default) returns only interactive elements; `filter` / `depth` / `ref` (subtree) / `maxChars` (default 20k) keep it small. Traverses open + closed shadow DOM, slots and same-origin iframes. |
| `browser_snapshot` | Accessibility tree with element refs. Compact mode (default) returns only interactive elements; `filter` / `depth` / `ref` or `selector` (subtree) / `maxChars` (default 20k) keep it small. Traverses open + closed shadow DOM, slots and same-origin iframes. `source: "native"` returns Chrome's own accessibility tree instead of the DOM-derived one ([details](#native-accessibility-tree)). |
| `browser_screenshot` | Capture a tab as an image over CDP — `maxWidth` / `scale` / `jpeg` to cut tokens, `fullPage` for the whole page, `region` to zoom; reports the pixel → x/y mapping |
| `browser_text` | Extract text from page or element (incl. shadow DOM); `mode:"article"` = main content only; `offset` paging |
| `browser_find` | Query elements by natural language ("search input", "Save button") — tokenized, role-aware, shadow DOM + same-origin iframes |
Expand Down
2 changes: 1 addition & 1 deletion agent-config/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,7 @@ Debug (tabId required, per-tab, capped 200 entries): `browser_console`, `browser
## Pattern

1. `browser_tabs { action: "list" }` → pick a `tabId`
2. `browser_snapshot { tabId }` to see the page and get refs (refs are tab-scoped)
2. `browser_snapshot { tabId }` to see the page and get refs (refs are tab-scoped). Add `source: "native"` for Chrome's own accessibility tree (exact accessible names, roles and states; uses the debugger) when the default tree misses or mislabels a control
3. Use `{ tabId, ref }` with interaction tools
4. Re-snapshot after navigation/DOM changes to refresh refs
5. `browser_wait { tabId, selector }` before interacting with dynamic content
Expand Down
2 changes: 1 addition & 1 deletion agent-config/cursor/rules/browser-controller.mdc
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ A `"Tool X disabled"` error means you must activate it via `browser_tools {actio
- `browser_select` - Select from `<select>`. Params: `tabId`, `ref` OR `selector`, `value` OR `label` OR `index`

### Reading (all require `tabId`)
- `browser_snapshot` - Accessibility tree with refs. Params: `tabId`, `selector`, `compact`
- `browser_snapshot` - Accessibility tree with refs. Params: `tabId`, `selector`, `compact`, `source` (`"dom"` default | `"native"` = Chrome's own computed tree via the debugger)
- `browser_screenshot` - Capture a tab as image (activates the tab to capture). Params: `tabId`, `format`, `quality`
- `browser_text` - Get text content. Params: `tabId`, `selector`, `maxLength`
- `browser_find` - Find elements by natural language. Params: `tabId`, `query`, `limit`
Expand Down
2 changes: 2 additions & 0 deletions agent-config/skills/browser-automation/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,8 @@ The server hides tool definitions until needed (`BROWSER_CONTROLLER_PROGRESSIVE=

For large pages, scope with a selector: `browser_snapshot { tabId, selector: "main" }`.

When the default tree mislabels or misses a control (custom widgets, `aria-labelledby` chains, `aria-hidden` overlays), ask for Chrome's own accessibility tree: `browser_snapshot { tabId, source: "native" }`. Same output shape and the refs work with every tool; it attaches the debugger (yellow banner) and falls back to the default tree, with `nativeUnavailable`, if it can't.

Use `browser_text { tabId }` to extract raw text when you need full content.

### Reading efficiently (don't re-scan everything)
Expand Down
Loading
Loading