browser-agent

browser-agent

An MCP server that gives Claude a real, persistent Chrome browser with logged-in sessions, enabling automation of sites that block headless browsers. It supports 50+ tools, cross-session knowledge, and recipe replay for complex workflows.

Category
访问服务器

README

browser-agent

An MCP server that gives Claude a real, persistent, logged-in Chrome browser.

Most browser MCPs use headless Playwright. They get blocked by Cloudflare, lose sessions between conversations, and fail on sites that detect automation. browser-agent runs actual headful Google Chrome with persistent profiles — so Claude can log into sites once and use them forever, across every conversation.


Why this exists

Headless browsers are fingerprinted and blocked. Real Chrome isn't.

When you need Claude to operate on LinkedIn, Gmail, Google Search Console, Discord, WhatsApp Web, YouTube Studio, or any site that requires a real logged-in session — headless fails. browser-agent solves this by giving Claude a real Chrome instance with:

  • Your actual cookies and login state (persistent across restarts and conversations)
  • A residential proxy so it looks like a real user from a real ISP
  • A cross-session knowledge base so Claude learns how each site works and never rediscovers the same traps
  • A recipe system so proven multi-step flows can be saved and replayed

Features

50+ MCP tools across 8 categories

Navigation & Interaction

  • browser_navigate — go to a URL, wait for load
  • browser_snapshot — compact role+name outline with clickable refs (preferred over screenshots)
  • browser_click, browser_click_text, browser_click_xy — click by ref, text, or coordinates
  • browser_type, browser_fill — type into inputs
  • browser_scroll, browser_press, browser_back — keyboard and scroll
  • browser_do — batch an entire flow into one tool call (saves 10x round-trips)

Reading & Scraping

  • browser_read_text — extract readable text from a page or selector
  • browser_elements — full element tree with refs
  • browser_eval — run arbitrary JavaScript
  • browser_console — read browser console logs
  • browser_aria — native accessibility tree

Network & Downloads

  • browser_fetch — fetch a URL in the browser's authenticated session (bypasses CORS)
  • browser_request — make raw HTTP requests with the browser's cookies
  • browser_download — download a file
  • browser_list_downloads — list downloaded files
  • browser_network — inspect network requests

Profile Management

  • browser_profile — switch to a named profile (each has its own Chrome, cookies, proxy)
  • browser_profiles — list all registered profiles
  • browser_add_profile — create a new profile with optional custom proxy
  • browser_free_profiles — stop idle profile browsers to reclaim RAM
  • browser_export_session, browser_import_session — backup/restore login sessions

Anti-Detection

  • browser_solve_cloudflare — handle Cloudflare challenges
  • browser_solve_captcha — solve captchas via 2captcha API
  • Stealth JS injected on every page (disables navigator.webdriver, spoofs fingerprints)
  • WebRTC forced through proxy (no IP leak)
  • --disable-blink-features=AutomationControlled + real Chrome flags

Knowledge & Recipes

  • browser_knowledge — recall what Claude has learned about a site (shared across sessions)
  • browser_learn — save a tip, recipe, or gotcha for a domain
  • browser_save_recipe — save a proven multi-step flow
  • browser_run_recipe — replay a saved recipe
  • browser_list_recipes — see all saved recipes
  • browser_maintenance — review the knowledge backlog

Live View

  • browser_login_view — open noVNC so you can manually log into a site by hand
  • browser_login_done — confirm login complete, returns page state
  • browser_view / browser_view_stop — watch the browser live in your browser
  • browser_screenshot — capture a screenshot

Session Control

  • browser_status — current profile, proxy, idle state, page title
  • browser_restart — restart Chrome (keeps profile)
  • browser_stop — shut down Chrome and VNC
  • browser_pages, browser_switch_page — manage multiple open tabs
  • browser_close_own_tabs — close tabs opened this session
  • browser_idle_autoclose — toggle auto-close after inactivity
  • browser_wait — wait N milliseconds

Architecture

Claude Code (stdio MCP)
    │
    ▼
server.py  ─── FastMCP (Python 3.11)
    │
    ├── browser.py      Chrome CDP session management
    ├── profiles.py     Multi-profile registry + port allocation
    ├── proxy.py        Residential proxy + local SOCKS5 relay
    ├── config.py       Per-profile configuration
    ├── stealth.js      Anti-fingerprint JS injected on every page
    ├── knowledge/      Cross-session site knowledge (JSON per domain)
    └── recipes/        Saved replayable flows (JSON)
    │
    ▼
Real headful Chrome (Xvfb display)
    │
    ├── Per-profile user data dir  (logins survive restarts)
    ├── Residential proxy          (real ISP egress IP)
    └── noVNC live view            (watch/control via browser)

Each profile gets its own:

  • Chrome user data directory (isolated cookies, localStorage, extensions)
  • Xvfb display number
  • CDP port
  • VNC port
  • Proxy (can override global proxy per profile)

Built-in knowledge base

The repo ships with cross-session knowledge for 25+ domains accumulated from real production use:

gmail · google search · google drive · youtube · youtube studio · linkedin · discord · whatsapp web · upwork · fiverr · freelancer · metricool · chatgpt · gemini · google labs · hostinger panel · spaceship · flowcv · sendpulse and more.

Each entry contains proven recipes, known gotchas, and site-specific tips — so Claude doesn't rediscover the same traps in future conversations.


Setup

Requirements

  • Linux (tested on AlmaLinux 9 / Ubuntu 22+)
  • Python 3.11+
  • Google Chrome (/usr/bin/google-chrome)
  • Xvfb (xorg-x11-server-Xvfb)
  • x11vnc + websockify (for live view)
  • A residential SOCKS5 proxy (required for production use; SSH tunnel supported for dev)

Install

git clone https://github.com/ahsaanfarooq/browser-agent
cd browser-agent
python3.11 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Configure

cp proxy.env.example proxy.env
# Edit proxy.env: set PROXY_SOURCE and proxy credentials

Key settings in config.py:

Variable Default Description
DISPLAY :1 Xvfb display number
CDP_PORT 9222 Chrome DevTools Protocol port
HOST_IP — Your server IP (for live view URLs)
SCREEN 1920x1080x24 Virtual display resolution

Add to Claude Code

{
  "mcpServers": {
    "browser-agent": {
      "command": "/path/to/browser-agent/venv/bin/python",
      "args": ["/path/to/browser-agent/server.py"]
    }
  }
}

Log in once, use forever

# Ask Claude:
"Open the browser view and navigate to gmail.com so I can log in"

# Claude calls browser_login_view("https://gmail.com")
# You see a live noVNC view, log in manually
# Claude calls browser_login_done() — session is now persisted

After logging in once, Claude can use that account in every future conversation without re-authenticating.


How Claude uses it

Token-lean — browser_snapshot returns a compact text outline with inline refs. Claude reads once and acts — no repeated round-trips, no expensive screenshots for navigation.

Batch flows — browser_do([...steps]) sends a whole sequence in one tool call with wait_for conditions. A 10-step flow = 1 API call.

Gets smarter over time — after solving a tricky interaction, Claude calls browser_learn() to save it for all future sessions. browser_knowledge("site.com") is auto-called on first visit each session, so Claude inherits everything learned before.


Example use cases

  • Content operations — manage 30+ WordPress sites, GSC, social accounts; check rankings, schedule posts, handle outreach
  • Freelance automation — drive Upwork, Fiverr, LinkedIn with real sessions (no API required, no rate-limit blocks)
  • Research — read Gmail threads, Google Docs, Discord servers, WhatsApp Web, paywalled pages
  • Monitoring — check dashboards, analytics, server panels that have no API
  • Data extraction — scrape sites protected by Cloudflare, login walls, or bot detection
  • Workflow automation — fill forms, upload files, click through multi-step flows on any site

Security notes

  • Proxy is fail-closed: if the proxy is down, Chrome refuses to launch. Your real IP is never used.
  • WebRTC is forced through the proxy (prevents IP leaks via browser_eval).
  • VNC ports bind 127.0.0.1 only. Live view is accessed via SSH tunnel or an authenticated dashboard.
  • Chrome runs with --no-sandbox (required on Linux servers without user namespaces). Run on a dedicated VPS, not your personal machine.

License

MIT


Author

Ahsaan Farooq — ahsaanfarooq.tech

Built for real production automation across 30+ sites. Running daily since early 2026. Contributions welcome.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选