Selenium MCP Server
Provides 66 tools for browser automation using Selenium WebDriver, enabling AI agents to navigate, interact, capture screenshots, assert conditions, manage cookies, handle tabs, execute HTTP requests, record sessions, and more.
README
🎭 Selenium MCP Server
Model Context Protocol server for browser automation using Selenium WebDriver.
66 tools for navigation, interaction, screenshots, assertions, cookies, tabs, HTTP API, recording, batch execution, iframes, and raw WebDriver access — all with strict hexagonal architecture.
Why this MCP?
Most Selenium MCPs expose basic WebDriver commands. This one adds the layers that make AI agents truly reliable:
| Feature | Status |
|---|---|
Page snapshot with stable element refs (capture_page → e1, e2, ...) |
✅ |
| Persistent per-domain selector hints | ✅ |
| HTTP API tools (GET/POST/PUT/PATCH/DELETE) | ✅ |
| Batch multi-step execution (up to 20 steps in 1 call) | ✅ |
| Built-in test assertions | ✅ |
Raw WebDriver API access (run_selenium) |
✅ |
| Recording for test generation | ✅ |
| Dialog handling (alert/confirm/prompt) | ✅ |
| Cookie management | ✅ |
| Stealth mode | ✅ |
| Strict hexagonal architecture (Domain/Ports/Adapters) | ✅ |
| Session tracing (NDJSON) | ✅ |
| Selenium Manager auto-provisioning | ✅ |
Quick Start
One-click install
From npm
npm install -g mcp-selenium-webdriver
Or run directly:
npx mcp-selenium-webdriver
Prerequisites
- Node.js 18+
- Chrome, Firefox, or Edge installed
- Selenium Manager provisions the matching driver automatically — no manual setup needed.
# Verify everything is ready
npx mcp-selenium-webdriver doctor
MCP Client Configuration
VS Code Copilot (.vscode/mcp.json):
{
"servers": {
"selenium-mcp": {
"command": "npx",
"args": ["mcp-selenium-webdriver"],
"type": "stdio"
}
}
}
Claude Desktop:
{
"mcpServers": {
"selenium-mcp": {
"command": "npx",
"args": ["mcp-selenium-webdriver"]
}
}
}
Cursor (.cursor/mcp.json):
{
"mcpServers": {
"selenium-mcp": {
"command": "npx",
"args": ["mcp-selenium-webdriver"]
}
}
}
Environment Variables
| Variable | Default | Description |
|---|---|---|
SELENIUM_BROWSER |
chrome |
Browser: chrome, firefox, edge |
SELENIUM_HEADLESS |
false |
Run browser in headless mode |
SELENIUM_STEALTH |
false |
Hide automation indicators |
SELENIUM_GRID_URL |
— | Selenium Grid hub URL |
SELENIUM_MCP_OUTPUT_MODE |
stdout |
Output: stdout or file |
SELENIUM_MCP_OUTPUT_DIR |
auto | Custom output directory |
SELENIUM_MCP_SAVE_TRACE |
false |
Save session trace as NDJSON |
SELENIUM_MCP_UNRESTRICTED_FILES |
false |
Bypass workspace sandbox |
HTTPS_PROXY / HTTP_PROXY |
— | Corporate proxy |
Pass them in your MCP config:
{
"mcpServers": {
"selenium-mcp": {
"command": "npx",
"args": ["mcp-selenium-webdriver"],
"env": {
"SELENIUM_HEADLESS": "true",
"SELENIUM_STEALTH": "true"
}
}
}
}
CLI Flags
npx mcp-selenium-webdriver [flags]
npx mcp-selenium-webdriver doctor # preflight diagnostics
| Flag | Description |
|---|---|
doctor |
Run diagnostics (Chrome, driver, proxy, cache) and exit |
--headless |
Run browser headless |
--stealth |
Enable stealth mode |
--save-trace |
Save session trace JSON |
--output-mode=stdout|file |
Set output mode |
--output-dir=<path> |
Custom output directory |
--grid-url=<url> |
Selenium Grid hub URL |
--allow-unrestricted-file-access |
Bypass workspace sandbox |
Tools (66)
🧭 Navigation (7)
| Tool | Description |
|---|---|
navigate_to |
Navigate to a URL. Starts browser automatically. |
go_back |
Navigate back in history. |
go_forward |
Navigate forward in history. |
refresh_page |
Refresh the current page. |
scroll_page |
Scroll the page or element into view. |
get_current_url |
Get the current URL. Read-only. |
get_title |
Get the page title. Read-only. |
📸 Page Analysis (5)
| Tool | Description |
|---|---|
capture_page |
Capture page state as element list with stable refs (e1-e300). Read-only. |
get_page_source |
Get the full HTML DOM. Read-only. |
get_visible_text |
Get visible text, optionally scoped to a CSS selector. Read-only. |
get_visible_html |
Get visible HTML (scripts removed by default). Read-only. |
take_screenshot |
Take a screenshot (viewport, full-page, or element). |
🖱️ Elements (9)
| Tool | Description |
|---|---|
click_element |
Click an element by CSS selector or ref. |
hover_element |
Hover over an element. |
double_click |
Double-click an element. |
right_click |
Right-click (context menu) an element. |
select_option |
Select a dropdown option by value, text, or index. |
drag_drop |
Drag one element to another. |
teach_selector |
Teach a preferred CSS selector for a domain. |
iframe_click |
Click an element inside an iframe. |
iframe_fill |
Fill an input inside an iframe. |
⌨️ Input (3)
| Tool | Description |
|---|---|
input_text |
Type text into an input field. |
key_press |
Press a keyboard key, optionally with modifiers (ctrl, alt, shift, meta). |
file_upload |
Upload a file through a file input. |
🖲️ Mouse (3)
| Tool | Description |
|---|---|
mouse_move |
Move mouse to coordinates. |
mouse_click |
Click at coordinates (left, right, middle). |
mouse_drag |
Drag from one position to another. |
📑 Tabs (4)
| Tool | Description |
|---|---|
tab_list |
List all open tabs. Read-only. |
tab_select |
Switch to a specific tab. |
tab_new |
Open a new tab. |
tab_close |
Close a tab. |
✅ Verification (4)
| Tool | Description |
|---|---|
verify_element_visible |
Verify an element is visible. Read-only. |
verify_text_visible |
Verify text is visible. Read-only. |
verify_value |
Verify an input has the expected value. Read-only. |
verify_list_visible |
Verify multiple texts are visible. Read-only. |
🌐 Browser Control (7)
| Tool | Description |
|---|---|
wait_for |
Wait for a condition (element visible, URL, title...). |
execute_javascript |
Run JavaScript in the browser context. |
resize_window |
Resize the browser window. |
dialog_handle |
Handle dialogs (alert, confirm, prompt). |
console_logs |
Get browser console logs with filtering. Read-only. |
network_monitor |
Monitor network requests / toggle offline mode. |
pdf_generate |
Generate a PDF from the current page. Read-only. |
🔄 Browser Lifecycle (5)
| Tool | Description |
|---|---|
start_browser |
Start the browser explicitly with options. |
close_browser |
Close the browser and end the session. |
reset_session |
Reset the session (close + restart). |
set_stealth_mode |
Enable/disable stealth mode. |
browser_status |
Get session status (URL, title, state). Read-only. |
🍪 Cookies (3)
| Tool | Description |
|---|---|
get_cookies |
Get all cookies. Read-only. |
add_cookie |
Add a cookie. |
delete_cookie |
Delete a cookie by name. |
⏺️ Recording (4)
| Tool | Description |
|---|---|
start_recording |
Start recording actions for test generation. |
stop_recording |
Stop recording and return the action log. |
recording_status |
Check recording status. Read-only. |
clear_recording |
Clear all recorded actions. |
💾 Selector Hints (4)
| Tool | Description |
|---|---|
selector_hint_save |
Save a preferred CSS selector for a domain. |
selector_hint_get |
Get saved hints for a domain. Read-only. |
selector_hint_list |
List domains with hints. Read-only. |
selector_hint_delete |
Delete a hint. |
🌍 HTTP API (5)
| Tool | Description |
|---|---|
http_get |
HTTP GET request. Read-only. |
http_post |
HTTP POST request with body. |
http_put |
HTTP PUT request. |
http_patch |
HTTP PATCH request. |
http_delete |
HTTP DELETE request. |
⚡ Power Tools (3)
| Tool | Description |
|---|---|
batch_execute |
Execute up to 20 tool steps in a single call. |
run_selenium |
Execute raw Selenium WebDriver API code (access to driver, By, until, Key). |
browser_generate_locator |
Generate a robust locator for an element by description. Read-only. |
Architecture
Hexagonal architecture (Ports & Adapters):
src/
├── domain/
│ ├── models/ Pure types & value objects
│ ├── ports/ 18 interfaces (IBrowserPort, INavigationPort...)
│ └── services/ 16 use cases
├── adapters/
│ ├── selenium/ 14 Selenium WebDriver implementations
│ ├── mcp/ MCP SDK server & tool registry
│ ├── persistence/ File system & selector hints storage
│ └── logging/ Console & trace (NDJSON) loggers
├── infrastructure/ DI container (tsyringe), bootstrap, config
├── index.ts Entry point
└── cli.ts CLI (doctor, flags)
Key design decisions:
- Dependency Injection via
tsyringewithSymboltokens — every port is swappable - Zod would be used for runtime validation (planned)
- Selenium Manager auto-provisions chromedriver matching your installed Chrome version
- 16-phase selector engine for
capture_page(ID → testId → role → label → ... → positional index) - Stealth mode injects scripts to hide
navigator.webdriverand automation indicators
Example Usage
Ask your AI agent:
Navigate to https://example.com, capture the page, click the first link, take a screenshot, verify "Example Domain" is visible.
The agent chains the tools automatically:
navigate_to → capture_page → click_element → take_screenshot → verify_text_visible
Or use raw WebDriver API:
run_selenium with:
const el = await driver.findElement(By.css('h1'));
const text = await el.getText();
return text;
Batch execution:
batch_execute with:
steps: [
{ tool: "navigate_to", args: { url: "https://example.com" } },
{ tool: "take_screenshot", args: { name: "before" } },
{ tool: "click_element", args: { selector: "a" } },
{ tool: "take_screenshot", args: { name: "after" } }
]
Development
git clone <this-repo>
cd mcp-selenium-webdriver
npm install
npm run build
npm run doctor
Scripts
npm run build # Compile TypeScript
npm run dev # Run with tsx (hot reload)
npm run typecheck # Type-check without emitting
npm test # Run all tests
npm run test:coverage # Run tests with coverage
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。