Taobao Sourcing Assistant
A local, human-paced MCP server for sourcing products on Taobao/Tmall, extracting per-SKU pricing, variant-linked reviews, and exporting to a comparison spreadsheet.
README
Taobao Sourcing Assistant
A local, human-paced MCP server that removes the drudgery of sourcing products on Taobao/Tmall. You keep all judgment (search intuition, buy decisions, sending supplier messages); the tool drives a real Chrome window to extract — for every product — a price for every SKU variant, specs, images, and reviews linked to the variant bought, then tabulates it into a comparison spreadsheet. Ships with a Claude Skill (sourcing playbook) and Chinese supplier-message templates (drafted by Claude, sent manually by you).
Built on the QR-login + persistent-session approach of
JeremyDong22/taobao_mcp, rebuilt as 6 FastMCP tools with embedded-data + DOM extraction (mtop interception kept as a fallback), per-SKU pricing, variant-linked reviews, xlsx export, and a captcha human-handoff.
What it does NOT do
No headless scraping, no proxy rotation, no captcha-solving service, no auto-messaging, no cloud. Not getting your account flagged is the priority, not speed.
Install (one time)
# from the project root
uv venv --python 3.12
uv pip install -e ".[dev]"
You need Google Chrome installed (the real app, not Chromium, not Comet). The
launcher is pinned to it in config.toml. If Chrome lives somewhere non-standard,
edit [browser] executable_path, or clear it ("") to let Playwright resolve the
chrome channel.
Configure
Edit config.toml (defaults are sensible):
[browser] executable_path— pinned Google Chrome binary (avoids launching Comet/other Chromium).[browser] user_data_dir— the persistent profile (your login lives here; gitignored).[pacing]— random delays +max_products_per_minute(keep it low).[limits]—max_reviews,review_pages.[output] dir— where xlsx +run.logland.
Run
.venv/bin/python server.py # stdio MCP server
npx @modelcontextprotocol/inspector .venv/bin/python server.py # interactive inspect
For Claude Desktop, register it as an MCP server pointing at the full venv python
path and server.py (use absolute paths — /Volumes/...).
First-run login (once per session)
- Call
taobao_initialize_login(or justtaobao_fetch_product— it auto-ensures login). - A visible Chrome window opens to the Taobao QR page.
- Scan the QR with your Taobao app. The server polls and continues automatically.
- The session persists in
user_data_dir— restarts reuse it, no re-scan.
Tools
| Tool | Purpose |
|---|---|
taobao_initialize_login |
Open Chrome, QR login (you scan). |
taobao_session_status |
Login/health (read-only). |
taobao_search |
Keyword → result list for you to pick from. |
taobao_fetch_product |
One product: every SKU variant + price/stock, specs, images. |
taobao_fetch_reviews |
Recent reviews, each tagged with the variant bought. |
taobao_export_xlsx |
3-sheet comparison workbook (Summary / Variants / Reviews). |
The Skill
skill/SKILL.md is the sourcing playbook (search → you pick → fetch → translate →
summarize reviews → normalize price-per-unit → compare → export → flag risks).
skill/supplier_templates.md has Chinese message templates — Claude drafts, you
send via Wangwang.
Troubleshooting
- It launched Comet / the wrong browser — set
[browser] executable_pathto your Google Chrome binary (default:/Applications/Google Chrome.app/Contents/MacOS/Google Chrome). - "login_required" / NotLoggedInError — run
taobao_initialize_loginand scan the QR; keep the window open. - A slider/verification appeared — solve it yourself in the Chrome window; the tool pauses (
human_action_required) and resumes. It logs tooutput/run.log. - Screenshots/automation "page still loading" — the new detail page holds a connection open; this server uses embedded-data + DOM extraction (not screenshot-waits), so this only affects ad-hoc scripts.
SelectorDriftError— Taobao changed its layout; patch the one filesrc/extract/selectors.py.- Wrong price on a multi-model listing — the headline price is the cheapest model; always read the per-SKU price for the exact variant.
补贴后prices may include a 国补 subsidy that needs a mainland ID — verify the real checkout price. - Only a few reviews returned — deep review pagination is shallow (known limit); increase scrolling in
src/extract/reviews.pyif needed. - Reset everything — delete
user_data/chrome_profile/and re-scan the QR.
Risks (don't hide these)
- Scraping Taobao violates its ToS; using your own logged-in account carries account-limitation risk. Keep volume low and human-paced.
- mtop endpoints / selectors drift — budget periodic maintenance (selectors are centralized).
Tests
.venv/bin/python -m pytest -q # parsers, output, MCP contract, drift, evals
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。