# Owl Browser API Reference

Complete documentation for all 180 automation tools. Navigate, interact, extract data, solve CAPTCHAs, and more.

### Interactive API Playground

The Docker server includes up-to-date API documentation with an interactive playground to test all tools in real-time. Access it at `http://localhost:8080` when running the server.

[**Python SDK** \  
Async-first library for Python 3.12+](/content/python-sdk/index.html)  [**Node.js SDK** \  
TypeScript library for Node.js 18+](/content/node-sdk/index.html)

## Total Tools

180

## Categories

35

## AI-Powered

7

## CAPTCHA Solvers

5

## Browse by Category

### Ad Blocking

Get ad blocker statistics and manage resource blocking

1 tool

### Advanced Mouse

Mouse movement and advanced cursor operations

1 tool

### AI Intelligence

AI-powered page analysis, natural language automation, and smart queries

7 tools

### CAPTCHA Solving

Detect, classify, and solve various types of CAPTCHAs

5 tools

### Clipboard

Read and write to the system clipboard

3 tools

### Console & Debugging

Read console logs and debug browser behavior

2 tools

### Content Extraction

Extract text, HTML, markdown, and structured data from pages

11 tools

### Context Management

Create, manage, and close isolated browser contexts (sessions)

5 tools

### Cookie Management

Read, set, and delete browser cookies

6 tools

### Demographics & Context

Get location, weather, time, timezone, and demographic information

5 tools

### Dialog Handling

Handle JavaScript alerts, confirms, and prompts

4 tools

### Download Management

Manage file downloads from the browser

6 tools

### Element State

Check visibility, enabled state, attributes, and dimensions

10 tools

### File Operations

Upload files to web pages

1 tool

### Frame Handling

Work with frames and iframes in web pages

3 tools

### FTP Client

List, download, and upload files via FTP/FTPS

2 tools

### HTTP Client

Make HTTP requests, manage sessions, and handle cookies

5 tools

### Input Control

Clear inputs, focus, blur, select text, and keyboard operations

5 tools

### IPC Testing

Run and manage IPC communication tests

7 tools

### JavaScript Evaluation

Execute JavaScript code in the page context

1 tool

### License Management

Manage browser license and activation

5 tools

### Navigation

Navigate pages, go back/forward, reload, and control browser history

7 tools

### Network Interception

Intercept, block, mock, and monitor network requests

10 tools

### Page Interaction

Click, type, select, drag and interact with page elements

13 tools

### Profile Management

Create, load, and save browser profiles with fingerprints

5 tools

### Proxy Management

Configure proxies, manage connections, and stealth features

4 tools

### Screenshots & Visual

Capture screenshots, highlight elements, and display overlays

3 tools

### Scroll Control

Scroll pages, navigate to elements, and control viewport position

4 tools

### Server Management

Restart the browser, read server logs, and manage the server

2 tools

### Tab Management

Manage browser tabs, popups, and windows

8 tools

### Video Recording

Record browser sessions, manage live streams, and capture video

11 tools

### Viewport & Page Info

Get page information and control viewport size

3 tools

### Wait Utilities

Wait for elements, network idle, conditions, and timeouts

7 tools

### WebMCP

Discover and call page-declared tools via WebMCP (W3C draft)

4 tools

### Zoom Control

Control page zoom level

3 tools

## Popular Tools

`go`
One-shot browser navigation tool. Creates a new context, navigates to the URL, waits for page load, extracts the HTML, and closes the context. Optimiz...

`browser_create_context`
Create a new isolated browser context with its own cookies, storage, and optional proxy configuration. Each context acts as an independent browser ses...

`browser_navigate`
Navigate the browser to a specified URL. This is a non-blocking operation that starts navigation and returns immediately. Use browser\_wait\_for\_network...

`browser_click`
Click on an element using CSS selector, XY coordinates, or natural language description. Supports semantic element finding using AI - describe what yo...

`browser_type`
Type text into an input field with human-like keystroke simulation. Target the field using CSS selector, coordinates, or natural language (e.g., 'emai...

`browser_screenshot`
Capture a PNG screenshot with configurable modes. 'viewport' (default) captures the current visible area, 'element' captures a specific element by CSS...
