Links and accessibility and extract
A scrape gives you the words. Sometimes you need something else from the same page.
Links
browsers links https://example.com -p playwright
https://iana.org/help/example-domains
The scrape of that page says "Learn more" and drops where it goes. links is the other door: every unique href, in page order. Cloudflare and Playwright have it.
For a model, browsers_links pages them. 500 per call unless you pass limit, 5 000 at most, and when more are left the answer ends with the offset to ask for:
[1200 more links; call again with offset 500]
Every call reads the page again. A total that changes between calls means the page changed, not the package.
Accessibility tree
A form in markdown is a mess. A form as its accessibility tree is a list of fields with their state:
browsers accessibility https://httpbin.org/forms/post --root fieldset
- group "Pizza Size"
- Legend
- StaticText "Pizza Size"
- InlineTextBox
- paragraph
- none
- radio "Small" checked=false
- none
- paragraph
- none
- radio "Medium" checked=false
- none
- paragraph
- none
- radio "Large" checked=false
- none
One node per line, indented, with the name, value and states like checked, disabled and expanded. --root keeps the subtree of the first element that matches, the boring generic nodes too. Leave it out and you get the whole page with those trimmed. --all keeps them anyway. A root that matches nothing is an error, not an empty tree.
Cloudflare only, for now.
Extract
browsers extract https://example.com --prompt "Extract the heading and the link" -p cloudflare
{
"heading": "Example Domain",
"link": "https://iana.org/help/example-domains"
}
Cloudflare answers at once. Hyperbrowser runs it as a job and waits up to two minutes. A job still running after that comes back with its ID instead of a timeout. The tool takes a JSON schema too, when "extract the heading" is too much freedom.
It's a model reading the page for you, so the answer is only as good as the prompt. Treat it that way.