Guide

MCP and Pi and OMP

Hand the browsers to a model. One MCP server and two extensions that call the same executors and answer with the same text

MCP

shell
browsers mcp
jsonmcp.json
{
  "mcpServers": {
    "browsers": { "command": "browsers", "args": ["mcp"] }
  }
}

A stdio server. It reads the same environment as the CLI, so whichever keys you export are the providers a model gets. No key at all still leaves Playwright.

Pi and OMP

shell
pi install git:github.com/agntn/browsers
omp install @agntn/browsers

Both extensions register the same tools as the MCP server, with the same names, schemas and descriptions, and a terminal view of each call.

The tools

browsers mcp · tools/list12 tools · same on MCP, Pi and OMP
  • browsers_scrapeRead-only/open-world network fetch: scrape content from a URL using a cloud browser provider.
  • browsers_sessionCreate a cloud browser session.
  • browsers_releaseRelease/destroy a cloud browser session.
  • browsers_providersRead-only/idempotent local/env status: list browser-as-a-service providers and which ones are currently configured via environment variables.
  • browsers_screenshotTake a screenshot of a URL using a cloud browser provider and return the image, or write it to `path` for captures too large to return.
  • browsers_extractExtract structured data from a URL using AI.
  • browsers_crawlCrawl a website following links.
  • browsers_pdfGenerate a PDF from a URL and write it to `path`.
  • browsers_linksExtract the unique links of a webpage in page order.
  • browsers_accessibilityRead a webpage's accessibility tree: the roles, names, values and states (checked, disabled, expanded) of its headings, links, buttons and form fields, one node per indented line.
  • browsers_searchWeb search via browser provider.
  • browsers_capabilitiesRead-only: report which library operations a browser provider supports, including scrape, screenshot, element screenshot, navigate, evaluate, sessions, CDP, and stateless modes.
descriptions from src/tool-contract.ts, as a model reads them

Hover a row for the whole description. It's the exact text a model gets in tools/list, and it's written to steer: which providers can do what, when to pass path, what to do with a jobId.

Same executors, same text

MCP, Pi and OMP don't have three implementations. They all call src/tool-operations.ts, so a fix lands once and a model reads the same answer on every host. The panel on the home page runs browsers_capabilities from that file at build time, which is as close to "what a model sees" as a web page gets.

Bounds a model can live with

  • browsers_scrape stops at 20 000 characters unless you pass maxChars, 200 000 at most, and says how much it cut.
  • browsers_crawl shares that budget across every page it read.
  • browsers_links returns 500 per call unless you pass limit, 5 000 at most, and names the next offset.
  • browsers_accessibility stops at the same 20 000 by default.
  • browsers_screenshot returns the image, or writes it to path when it's too big to return.

What a model can't do

Click. Log in. Type into a form. The tools never grew navigate or evaluate, and a session ID from browsers_session only goes to browsers_release. A model can read the web through this package. Driving it is your code's job, with the CDP URL a session gives you.

Need a plain fetch that isn't a browser? @agntn/web. Need last year's version of the page? @agntn/archives.