Crawl and search
Crawl
browsers crawl https://example.com --maxPages 5 -p playwright
ℹ Crawling via playwright...
--- https://example.com ---
This domain is for use in documentation examples without needing permission. …
A crawl starts at one page and follows links, ten pages unless you say --maxPages, one link deep unless you say --maxDepth. Cloudflare, Hyperbrowser and Playwright crawl. Playwright does it right there in your Chromium. The other two start a job on their side and poll it.
A slow crawl keeps its job
A big site takes a while, and a tool call can't wait forever. Cloudflare and Hyperbrowser get two minutes. A job still running after that doesn't turn into a timeout, it comes back with its ID:
[provider=cloudflare] Crawled 0 pages.
Job ID: 5b1f… (status: running)
The job is still running. Call browsers_crawl with this jobId and provider cloudflare to wait for it again.
Pass that jobId back, with the same provider (and the same browser if you used Kitesurf), and without url. It waits for the same job. No second crawl, no second bill. In the CLI that's --job:
browsers crawl --job <job-id> -p cloudflare
A job ID belongs to the provider that made it, so naming the provider isn't optional there. Picking one on its own, the package could land on a provider that has never heard of your job.
Shared budget
For a model, every page's text shares one maxChars, 20 000 by default. Ten pages don't get 20 000 each. That's deliberate. A crawl is the easiest way to drown a context window, and this one can't.
Search
browsers search "browser automation"
Awesome Browser Automation - GitHub
https://github.com/angrykoala/awesome-browser-automation
A curated list of awesome browser automation tools and resources. …
Hyperbrowser is the one with a search route, so search goes there. Title, URL and a snippet per result. It's a web search that happens to live in a browser package, and it's handy exactly when a model needs a URL before it can scrape one.