Guide

Links and accessibility and extract

Three other ways to read one page. The hrefs the scrape dropped and the fields of a form and structured data from a prompt

A scrape gives you the words. Sometimes you need something else from the same page.

shell
browsers links https://example.com -p playwright
text
https://iana.org/help/example-domains

The scrape of that page says "Learn more" and drops where it goes. links is the other door: every unique href, in page order. Cloudflare and Playwright have it.

For a model, browsers_links pages them. 500 per call unless you pass limit, 5 000 at most, and when more are left the answer ends with the offset to ask for:

text
[1200 more links; call again with offset 500]

Every call reads the page again. A total that changes between calls means the page changed, not the package.

Accessibility tree

A form in markdown is a mess. A form as its accessibility tree is a list of fields with their state:

shell
browsers accessibility https://httpbin.org/forms/post --root fieldset
text
- group "Pizza Size"
  - Legend
    - StaticText "Pizza Size"
      - InlineTextBox
  - paragraph
    - none
      - radio "Small" checked=false
      - none
  - paragraph
    - none
      - radio "Medium" checked=false
      - none
  - paragraph
    - none
      - radio "Large" checked=false
      - none

One node per line, indented, with the name, value and states like checked, disabled and expanded. --root keeps the subtree of the first element that matches, the boring generic nodes too. Leave it out and you get the whole page with those trimmed. --all keeps them anyway. A root that matches nothing is an error, not an empty tree.

Cloudflare only, for now.

Extract

shell
browsers extract https://example.com --prompt "Extract the heading and the link" -p cloudflare
json
{
  "heading": "Example Domain",
  "link": "https://iana.org/help/example-domains"
}

Cloudflare answers at once. Hyperbrowser runs it as a job and waits up to two minutes. A job still running after that comes back with its ID instead of a timeout. The tool takes a JSON schema too, when "extract the heading" is too much freedom.

It's a model reading the page for you, so the answer is only as good as the prompt. Treat it that way.