For AI agents: a documentation index is available at /llms.txt
Skip to main content

Browser Agent

The browserless_agent tool turns any MCP-compatible AI assistant into a browser-using agent. It exposes a snapshot-based protocol so the model can reason about what's on the page and drive it step by step. Sessions are one-shot by default; opt in on the first call to keep a browser session across tool calls.

Use it when a single scrape isn't enough — filling forms, logging in, clicking through multi-step flows, paginating search results, interacting with consent dialogs, or any task that needs the browser's state to survive between turns.

Prerequisites
  • Set up the Browserless MCP Server in your MCP client (Claude Desktop, Cursor, VS Code, Windsurf, Claude Code, etc.)

How it works​

The agent follows the classic ReAct loop — Reason → Act → Observe:

  1. goto — navigate to a URL
  2. snapshot — capture every interactive and informational element on the page (buttons, links, inputs, headings, images with alt text) as a compact list with stable selectors
  3. plan — the model reads the snapshot and decides what to do
  4. act — click, type, select, scroll, etc., using selectors taken directly from the snapshot
  5. re-snapshot whenever the page changes (navigation, form submission)
  6. Repeat until the task is done, then close

The browser keeps cookies, local storage, and navigation history across commands. For this loop to span tool calls, set keepSessionAlive: true on the first call, then pass the returned sessionId on later calls. That keeps the same browser state available — you can log in once and keep interacting with authenticated pages.

Session lifetime​

This contract applies to MCP server v1.33.0 and later. Earlier versions kept sessions alive by default. Check your server version when upgrading a self-hosted or pinned client.

These are top-level browserless_agent parameters, alongside commands:

ParameterTypeBehavior
sessionIdOptional stringPass the ID returned by a kept-alive call to continue that browser. Omitting it starts a new browser, not a continuation.
keepSessionAliveOptional boolean, no schema defaultOmitted on a new session: one-shot. Set true on the first call for multi-step work. On later calls, echoing sessionId keeps it alive without repeating the flag. Explicit false force-closes even when continuing.
sessionIdkeepSessionAliveAfter a successful command batch and download drain
OmittedOmittedClose the browser and free its concurrency slot.
OmittedtrueKeep the new browser for later calls.
Returned IDOmitted or trueKeep that browser for further calls.
Omitted or returned IDfalseClose the browser, including a continued session.

Exceptions: keepSessionAlive is ignored for profile creation (createProfile) and attached sessions (attachSessionId), whose lifetimes are managed separately. If download collection fails, the session is retained for recovery instead of being auto-closed: retry getDownloads with the sessionId supplied in the error, then close the session when finished. Kept sessions still have the unchanged 15-minute idle backstop; keeping a session alive is not indefinite retention.

One-shot example​

When one batch finishes the task, omit both lifetime parameters. This call reads the page and automatically closes after the batch and download drain succeed:

{
"rationale": "Reading the example page",
"commands": [
{ "method": "goto", "params": { "url": "https://example.com" } },
{ "method": "text" }
]
}

Multi-call example​

If you need to inspect a snapshot before deciding what to do next, opt in before that first call finishes. When unsure whether more calls are needed, keep the session alive.

1. Start and inspect:

{
"rationale": "Inspecting the example page",
"keepSessionAlive": true,
"commands": [
{ "method": "goto", "params": { "url": "https://example.com" } },
{ "method": "snapshot" }
]
}

2. Continue: replace SESSION_ID_FROM_FIRST_CALL with the returned ID. No keepSessionAlive flag is needed; echoing the ID keeps the browser alive.

{
"rationale": "Reading the page text",
"sessionId": "SESSION_ID_FROM_FIRST_CALL",
"commands": [{ "method": "text" }]
}

3. Finish and force-close: explicit false overrides the keep-alive implied by sessionId.

{
"rationale": "Reporting completion",
"sessionId": "SESSION_ID_FROM_FIRST_CALL",
"keepSessionAlive": false,
"commands": [
{
"method": "reportOutcome",
"params": { "success": true, "reason": "completed" }
}
]
}

You can also end a kept session with an explicit close command, optionally preceded by reportOutcome.

Hard-coded multi-call clients

Update scripts that open a browser and expect to reuse it later: set keepSessionAlive: true on the first call, then echo the returned sessionId on every continuation. Passing the ID only on the second call is too late if the first call was one-shot; it cannot restore the closed browser's page, login, or progress.

Available methods​

MethodPurpose
goto { url, waitUntil? }Navigate to a URL. Defaults to waitUntil: "domcontentloaded".
back { waitUntil? }Go back in browser history.
forward { waitUntil? }Go forward in browser history.
reload { waitUntil? }Reload the current page.

Observation​

MethodPurpose
snapshot { maxElements?, full?, targetId? }Capture the page as a list of interactive + informational elements with selectors. The first snapshot is complete; later ones return only what changed since the previous snapshot. Pass full: true to get the complete element list again. targetId snapshots a background tab without switching to it.
text { selector? }Extract the text of a specific element (or the whole page).
html { selector? }Get the raw HTML of a section (or the whole page).
screenshot { fullPage?, selector?, clip?, type?, quality?, toDisk? }Capture the page, an element, or a region as an image. selector, clip, and fullPage are mutually exclusive. toDisk: true saves the file and returns a file reference instead of inline image data (see Files for how references work per transport).
evaluate { content }Run arbitrary JavaScript in the page context (IIFE syntax).

Interaction​

MethodPurpose
click { selector }Click an element.
type { selector, text }Type text into an input.
select { selector, value }Choose an option in a <select>.
checkbox { selector, checked? }Toggle a checkbox (preferred over click for checkboxes).
hover { selector }Hover over an element.
scroll { selector?, direction? }Scroll the page or a specific element.

Secrets and CAPTCHAs​

MethodPurpose
loadSecret { ref, selector? }Inject a credential from a secrets vault (for example an op:// reference) into an input. The value is resolved server-side and typed into the field, so it never enters the model's context. Use this instead of type for vault-managed usernames and passwords.
solve { type?, timeout?, wait? }Solve a CAPTCHA on the page. Auto-detects the type when omitted; supports Cloudflare, reCAPTCHA (v2/v3), GeeTest, DataDome, Akamai, and more. Experimental and cloud-only: self-hosted Enterprise deployments lack the solver backend.

Files​

MethodPurpose
uploadFile { selector, files }Attach files to an <input type="file">. Each file comes from a local path (stdio mode), a download handle, or inline base64 content. Combined size is capped (10MB default, 50MB max).
getDownloadsList the files the page has downloaded, returned as file references you can pass back to uploadFile or save locally.

File references are transport-specific. In stdio (local) mode they're plain filesystem paths. Over HTTP each file comes with two references: a single-use download URL for saving the bytes, and a browserless-download://<id> handle for uploadFile. The download URL is consumed on first use, so don't pass it as an upload handle; use the browserless-download:// handle instead.

Tabs​

The session isn't limited to a single page. Sites that open results in new tabs (external links, OAuth popups, print views) stay reachable:

MethodPurpose
getTabsList every open tab with its targetId, URL, and title.
switchTab { targetId }Make another tab the active one.
createTab { url?, activate?, waitUntil? }Open a new tab. Pass activate: false to open it in the background and keep the current tab active. waitUntil applies only when activating; defaults to "domcontentloaded".
closeTab { targetId }Close a specific tab.

Waiting​

MethodPurpose
waitForSelector { selector, timeout? }Wait for a DOM element to appear.
waitForNavigation { timeout? }Wait for a page navigation to complete.
waitForTimeout { time }Wait for a fixed duration (ms).
waitForRequest { url?, method?, timeout? }Wait for a network request matching a URL glob.
waitForResponse { url?, statuses?, timeout? }Wait for a network response matching a URL glob and/or status codes.

Session​

MethodPurpose
liveURL { timeout?, interactable?, quality?, type?, resizable? }Open a shareable live view of the browser that a human can watch or interact with.
closeEnd the browser session.

Reporting​

MethodPurpose
reportOutcome { success, reason? }Optionally report the task verdict before close. success is a boolean; reason is optional and must be one of completed, blocked_by_site, captcha, login_required, timeout, or other. The last report wins and does not replace the page result returned by a batch.

Batching​

When multiple actions share the same page state (for example filling a form), they can be batched into a single tool call using the commands array. This is significantly faster than invoking the tool once per step.

Continue a browser kept alive by an earlier call, using selectors from its snapshot:

{
"sessionId": "SESSION_ID_FROM_FIRST_CALL",
"commands": [
{ "method": "type", "params": { "selector": "input#email", "text": "user@example.com" } },
{ "method": "type", "params": { "selector": "input#password", "text": "secret" } },
{ "method": "click", "params": { "selector": "input#remember-me" } },
{ "method": "click", "params": { "selector": "button[type='submit']" } }
]
}

Commands run sequentially, so goto can be followed by text or snapshot in the same batch, as in the examples above. Batch interactions that use the same page state, but don't reuse selectors from an old snapshot after navigation or a form submission. Capture a fresh snapshot first; if you need another tool call to decide the next interaction, keep the session alive as shown in the multi-call example.

Error recovery​

The agent protocol returns structured errors with codes the model can reason about:

CodeMeaningTypical recovery
SELECTOR_NOT_FOUNDElement didn't exist or wasn't found in timeTry a deep selector (< selector) — the element may be in shadow DOM. Otherwise re-snapshot.
NAVIGATION_TIMEOUTPage didn't reach the requested ready stateRetry with a longer timeout or waitUntil: "domcontentloaded".
TIMEOUTGeneric operation timeoutRetry with a longer timeout.
BROWSER_CRASHEDSession is goneThe MCP client transparently reconnects — the agent just re-navigates.
INVALID_PARAMSBad argumentsNot retryable.
RATE_LIMITEDToo many requestsWait and retry.

When a selector isn't found, the error response also includes a fresh snapshot so the model can re-plan without an extra round trip.

When to use the agent vs. Smart Scraper​

  • Use browserless_smartscraper when you need the content of a single page. It picks the best strategy (direct fetch → proxy → headless browser → CAPTCHA solving) and returns the content in your preferred format.
  • Use browserless_agent when the task requires multiple turns, state, or interaction — logging in, filling forms, paginating, navigating menus, reacting to what's on the page.

For the Cloud-only, draft-preview Agentic Checkout flow, which is not yet available on the hosted fleet, keep this browser session open and use the dedicated browserless_link_checkout flow. Payment credentials never enter browserless_agent commands.

Example tasks​

Ask your AI assistant:

Find a French ratatouille recipe on Allrecipes with a 4-star rating or higher and at least 15 reviews. Note the variety of vegetables included and the overall cooking time.

Open Google Flights, search for a round-trip flight from New York to Tokyo departing March 10, 2026 and returning March 24, 2026, and list the five cheapest options.

Search for a local text-to-speech repository on GitHub updated in the last 10 days with at least 200 stars and summarize its main objective.

Go to FedEx. Calculate the shipping rates for a 10-pound package from Miami, FL to San Francisco, CA, and summarize prices and delivery dates.

FAQ & Troubleshooting​

How is this different from the /agent/run REST endpoint?

They split the reasoning differently. With browserless_agent, your AI model does the thinking: it reads snapshots, plans, and drives the browser step by step through MCP tool calls. With /agent/run, Browserless's managed agent does the thinking: you POST a task in natural language, it browses autonomously, and you poll for the structured result. Use this tool when you already have an MCP-connected assistant in the loop; use /agent/run when you want to submit a task from any HTTP client and collect the answer.

What is the Browserless browser agent MCP tool?

The browser agent tool gives AI models direct control of a browser session, allowing them to navigate, click, type, and extract information from web pages autonomously through natural language instructions.

How does the browser agent differ from REST API tools?

REST API tools call individual Browserless endpoints for specific tasks. The browser agent provides a live browser session where the AI can perform multi-step interactions, making it better suited for complex, dynamic workflows.

Can the browser agent handle login flows?

Yes. The browser agent can navigate to login pages, fill in credentials, handle multi-factor authentication prompts, and maintain session state across subsequent page interactions.

For sites you log into repeatedly, use Authenticated Profiles to save the login once and reuse it across sessions instead of logging in every time.

Next steps​

Was this page helpful?