TaobaoScraper MCP Server

TaobaoScraper MCP Server

An MCP server for scraping product data from Taobao/Tmall and JD.com, providing 8 tools for scraping, task management, notifications, and system control.

Category
访问服务器

README

TaobaoScraper MCP Server

An MCP (Model Context Protocol) server for scraping product data from Taobao/Tmall and JD.com (Jingdong). Provides 8 tools that can be used directly in Claude Desktop or Claude Code.

Platform: Windows (requires Chrome browser) Language: Python 3.10+

Features

  • Multi-platform scraping - Taobao, Tmall, JD.com in one tool
  • 8 MCP tools - Scraping, task management, notifications, system control
  • Hot search words - Discover trending keywords and market insights
  • Excel export - Auto-saves results as .xlsx files
  • Notification push - WeChat (ServerChan / PushPlus) and Email alerts
  • Dual transport - stdio (local) and SSE (remote) protocols
  • Trilingual GUI - Chinese / English / Korean interface (optional)

Architecture

Claude Desktop / Claude Code
        |
   MCP Server (stdio or SSE)
        |
   FastAPI Backend (:8000)
        |
   Selenium + Chrome (:9222)

Quick Start

1. Prerequisites

  • Python 3.10+
  • Google Chrome browser
  • Windows OS

2. Install

git clone https://github.com/jhongjun1981/taobao-scraper-mcp.git
cd taobao-scraper-mcp

# Install MCP server dependencies
pip install -r requirements_mcp.txt

# Install API backend dependencies
pip install -r requirements_api.txt

3. Configure

cp .env.example .env
# Edit .env to set your SCRAPER_API_KEY

4. Start Chrome with debug port

# Close all Chrome instances first, then:
START_GUI.bat
# Or manually:
chrome.exe --remote-debugging-port=9222 --user-data-dir="%LOCALAPPDATA%\Google\Chrome\Debug Profile"

5. Login to Taobao (first time only)

Open Chrome and login to your Taobao account. Cookies will persist in the debug profile.

6. Start API Backend

python run_api.py

7. Configure MCP Server

For Claude Code - Add to .claude/settings.json:

{
  "mcpServers": {
    "taobao-scraper": {
      "command": "python",
      "args": ["run_mcp.py"],
      "cwd": "/path/to/taobao-scraper-mcp"
    }
  }
}

For Claude Desktop - Add to claude_desktop_config.json:

{
  "mcpServers": {
    "taobao-scraper": {
      "command": "python",
      "args": ["run_mcp.py"],
      "cwd": "C:\\path\\to\\taobao-scraper-mcp"
    }
  }
}

SSE mode (remote):

python run_mcp.py --transport sse --port 8001
{
  "mcpServers": {
    "taobao-scraper": {
      "type": "sse",
      "url": "http://your-server:8001/sse"
    }
  }
}

Tools Reference

Scraping Tools

scrape_products

Scrape product data from e-commerce platforms.

Parameter Type Default Description
keyword string required Search keyword
platform string "taobao" "taobao" / "jd" / "multi"
pages int 3 Number of pages to scrape
sort_by string "sale" Sort method
tmall_only bool false Tmall products only
exact_match bool false Exact keyword match
price_min float 0 Min price filter
price_max float 0 Max price filter
wait_for_result bool true Wait for completion
timeout int 300 Timeout in seconds

Example: "Scrape running shoes from both Taobao and JD"

scrape_hotwords

Scrape trending/related search keywords for market analysis.

Parameter Type Default Description
keyword string required Base keyword
wait_for_result bool true Wait for completion
timeout int 120 Timeout in seconds

Task Management Tools

list_tasks

List recent scraping tasks with status (pending/running/completed/failed).

Parameter Type Default Description
limit int 20 Max tasks to return

get_task

Get detailed status of a specific task. Set include_result=true to get full results.

Parameter Type Default Description
task_id string required Task ID
include_result bool false Include result data

cancel_task

Cancel a running or pending task.

Parameter Type Default Description
task_id string required Task ID to cancel

Export Tools

list_files

List all exported Excel data files with filename, size, and modification time.

Notification Tools

send_notification

Send push notifications via WeChat or Email.

Parameter Type Default Description
channel string required "wechat" / "email" / "test"
title string "" Notification title
content string "" Notification body
wx_type string "server_chan" "server_chan" or "pushplus"

System Tools

system_status

Check system health or restart Chrome.

Parameter Type Default Description
action string "check" "check" or "restart_chrome"

Environment Variables

Variable Default Description
SCRAPER_API_KEY changeme-your-secret-key API authentication key
SCRAPER_API_URL http://localhost:8000 FastAPI backend URL
SCRAPER_CHROME_PORT 9222 Chrome debug port
MCP_SSE_HOST 0.0.0.0 SSE listen address
MCP_SSE_PORT 8001 SSE listen port
MCP_HTTP_TIMEOUT 30 HTTP request timeout

Data Output

Scraped data is automatically saved as Excel files containing:

  • Product title, image URL, price
  • Monthly sales volume, review count
  • Shop name, shop type, shop region
  • Product URL, scrape timestamp

Important Notes

  • Taobao requires login - You must login to your own Taobao account in Chrome first
  • JD works without login - JD scraping does not require authentication
  • One task at a time - Chrome can only handle one scraping task concurrently
  • Windows only - The Chrome automation layer uses Windows-specific APIs
  • Respect rate limits - Excessive scraping may trigger anti-bot protection

License

MIT License - see LICENSE file.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选