Wafle-Scraper
Enables AI agents to scrape web pages safely with incognito browsing, rate limiting, and ethical CAPTCHA handling.
README
Wafle-Scraper — Universal MCP Server for Safe Web Scraping
Wafle-Scraper is an MCP server that lets AI agents (OpenCode, Claude, Cursor) extract data from the web safely and responsibly.
Safety First
| Rule | Enforcement |
|---|---|
| No private data | Every browser session is incognito — no cookies, no localStorage, no saved passwords |
| Only what you ask | The scraper never mines extra data beyond your explicit request |
| No localhost | Internal/private IPs are blocked by default |
| Rate limited | Minimum 2 seconds between requests — never floods servers |
| CAPTCHA = human only | No automated CAPTCHA solving. If one appears, you solve it interactively |
| User-Agent rotation | Each request looks like a real browser |
Quick Start
pip install wafle-scraper
playwright install chromium
wafle-scraper
Configuration
OpenCode / Claude Desktop / Cursor
{
"mcpServers": {
"wafle-scraper": {
"command": "wafle-scraper",
"description": "Web scraping & browser automation — incognito, audited, safe"
}
}
}
CLI Options
wafle-scraper # MCP stdio mode (default for agents)
wafle-scraper --http --port 8000 # HTTP SSE mode
wafle-scraper --version # Show version
MCP Tools
| Tool | Description | Safety |
|---|---|---|
scrape_url |
Extract text from a static URL (requests + BeautifulSoup) | ✅ Read-only, no JS |
scrape_browser |
Navigate a page in isolated incognito browser and extract text | ✅ Incognito, no cookies |
scrape_reddit |
Fetch public posts from a subreddit via official API | ✅ API, no scraping |
extract_emails |
Find email addresses on a public page | ✅ Only what you ask |
browser_interact |
Click, type, scroll, extract, screenshot in live browser | ✅ You control the actions |
Browser Backend (Playwright)
- Incognito always:
storage_state=None, fresh context per session - No permissions: No camera, mic, location access
- Anti-detection: Rotating UA, viewport, locale, timezone
- Natural delays: Human-like timing between actions
- Gradual scroll: Loads lazy content naturally
CAPTCHA Handling
Wafle-Scraper does NOT solve CAPTCHAs automatically. When a CAPTCHA is detected:
- The scraper pauses
- Prompts you to open the URL in your browser
- You solve the CAPTCHA manually
- Type
doneand the scraper continues
This is the only ethical and reliable approach without paid services.
Security
- Blocked:
localhost,127.0.0.1, private IPs,file://,chrome:// - Rate limiting (configurable, default 2s min interval)
- Scope enforcement — only processes what you explicitly request
- User-Agent rotation
- Browser isolation — Playwright contexts are fully sandboxed
Requirements
- Python 3.10+
- Playwright with Chromium installed (
playwright install chromium) - Windows, macOS, Linux
Installation from Source
git clone https://github.com/creandoaldia/wafle-scraper.git
cd wafle-scraper
pip install -e .
playwright install chromium
License
MIT
Why "Wafle-Scraper"?
Part of the WAFLE ecosystem (Web AI Framework for Language Ecosystems). Wafle-Scraper gives WAFLE agents the ability to read the live web — safely, transparently, and under your control.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。