MCP WebScout
A Model Context Protocol server that provides web search capabilities via DuckDuckGo and advanced content extraction using Crawl4AI and LLM-powered analysis. It enables users to perform web-wide searches and fetch processed website data through automated browser interaction and intelligent summarization.
README
MCP WebScout
A Model Context Protocol (MCP) server providing web search (DuckDuckGo) and intelligent content extraction with LLM-powered analysis.
Features
- search: Search the web using DuckDuckGo
- fetch: Advanced web fetching with Crawl4AI and LLM extraction
System Requirements
| Requirement | Version | Notes |
|---|---|---|
| Python | >= 3.10 | Required runtime environment |
| pip | latest | Package manager (included with Python) |
| Playwright | latest | Required by Crawl4AI for browser automation |
| DeepSeek API Key | - | Required for LLM extraction mode |
| Proxy (optional) | - | Required for users in mainland China |
Python Dependencies (14 packages)
| Package | Version | Purpose |
|---|---|---|
| mcp | >=1.0.0 | MCP protocol implementation |
| duckduckgo-search | >=3.0.0 | DuckDuckGo search API |
| requests | >=2.32.0 | HTTP requests |
| beautifulsoup4 | >=4.12.0 | HTML parsing |
| openai | >=1.30.0 | OpenAI API client for DeepSeek |
| crawl4ai | >=0.5.0 | Advanced web scraping |
Quick Start
Get started in 5 steps:
1. Clone and Setup Environment
git clone <repository>
cd mcp-webscout
python -m venv .venv
On Windows:
.venv\Scripts\activate
On macOS/Linux:
source .venv/bin/activate
2. Install Dependencies
pip install -e ".[dev]"
3. Install Playwright Browsers
playwright install chromium
4. Configure Environment Variables
cp .env.example .env
Edit .env and add your configuration:
# Required for LLM extraction
DEEPSEEK_API_KEY=sk-your-actual-key-here
# Required for mainland China users
PROXY_URL=http://127.0.0.1:7890
USE_PROXY=true
5. Verify Installation
# Run tests
pytest tests/ -v
# Test the server
python -m mcp_webscout --help
Detailed Configuration
For detailed environment setup instructions, see ENV_SETUP.md.
Usage
As a Command
mcp-webscout
As a Python Module
python -m mcp_webscout
With Claude Desktop
Add to your claude_desktop_config.json:
Basic Configuration
{
"mcpServers": {
"webscout": {
"command": "mcp-webscout"
}
}
}
With Environment Variables (Recommended)
{
"mcpServers": {
"webscout": {
"command": "mcp-webscout",
"env": {
"DEEPSEEK_API_KEY": "sk-your-key-here",
"PROXY_URL": "http://127.0.0.1:7890",
"USE_PROXY": "true",
"DEFAULT_MAX_LENGTH": "5000",
"PYTHONUTF8": "1"
}
}
}
}
Windows Configuration
{
"mcpServers": {
"webscout": {
"command": "python",
"args": ["-m", "mcp_webscout"],
"env": {
"DEEPSEEK_API_KEY": "sk-your-key-here",
"PROXY_URL": "http://127.0.0.1:7890",
"USE_PROXY": "true",
"PYTHONUTF8": "1"
}
}
}
}
Tools
search
Search the web using DuckDuckGo.
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
| query | string | Yes | Search query |
| max_results | integer | No | Maximum results (1-10, default: 5) |
Returns:
Formatted search results with titles, URLs, and snippets.
Example:
{
"query": "Python programming",
"max_results": 3
}
fetch
Advanced web fetching with Crawl4AI and LLM extraction.
Parameters:
| Name | Type | Required | Description |
|---|---|---|---|
| url | string | Yes | URL to fetch |
| mode | string | No | Extraction mode: simple, llm (default: simple) |
| prompt | string | No | Custom extraction prompt for LLM mode |
| max_length | integer | No | Maximum characters (default: 5000) |
| use_proxy | boolean | No | Use proxy (default: true) |
Returns:
Fetched and optionally extracted content.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。