spider-data
MCP server providing web search, fetch, and crawl capabilities for AI agents via the spider.cloud API, with caching, cost controls, and robust error handling.
README
spider-data
Web search, fetch, and crawl for AI agents, backed by the
spider.cloud API. Drop-in replacement for built-in
WebSearch and WebFetch tools.
Zero dependencies. One CLI, one MCP server, one skill file.
spider-data search "postgres lock contention"
spider-data fetch "https://example.com/article"
spider-data crawl "https://docs.example.com" --limit 30
Why
Cost-effective. Fetches start on a free direct request. They escalate to the
paid browser fleet only on a confirmed block (403, challenge page, empty render
target), never on a guess. Each URL has a 2 paid-request ceiling. Results are
cached for 24 hours (search) and 7 days (fetch and crawl), revalidation is free,
and re-reads via --section and --grep cost nothing.
Reliable. Every outcome has a stable failure_code and exit code, so an
agent branches on exit 4 instead of parsing prose. Block detection is
structural, not heuristic, so a short page behind a CDN is not misreported as
walled. Rate limits and 5xx responses retry with backoff and honour
Retry-After. 184 tests run on every publish.
Scalable. The cache is shared across runs and processes, so parallel agents on the same topic pay once. A 14 day host memory skips the free request on hosts that consistently return 403. It can only skip a free request, never add a paid one. Long pages return a section outline plus a budgeted excerpt, keeping context windows bounded on large crawls.
Requirements
Node 20+ or Bun. No other dependencies.
Install
git clone https://github.com/MirandaKim1434/spider-data
cd spider-data
npm link
SPIDER_API_KEY_INPUT=<your-key> spider-data auth login
spider-data doctor
auth login stores the key in the OS keychain.
On a machine with Bun but no Node, the bin/ shebang will not resolve. Install a
wrapper instead of linking:
printf '#!/bin/sh\nexec bun "%s/bin/spider-data.js" "$@"\n' "$PWD" > ~/.local/bin/spider-data
chmod +x ~/.local/bin/spider-data
Claude Code
ln -s "$PWD/skill" ~/.claude/skills/spider-data
Disable the built-ins in ~/.claude/settings.json so there is one path to the
web rather than two:
{
"permissions": {
"deny": ["WebSearch", "WebFetch"],
"allow": ["Bash(spider-data:*)"]
}
}
MCP
{ "mcpServers": { "spider-data": { "command": "node", "args": ["/path/to/spider-data/src/mcp.js"] } } }
Exposes web_search, web_fetch, and web_crawl. Use "command": "bun" on a
Bun-only machine. See adapters/ for host-specific notes.
Commands
spider-data search <query> spider-data cache {stats|clear|purge <url>}
spider-data fetch <url> spider-data stats [--since 7d]
spider-data crawl <url> spider-data doctor
spider-data auth {login|status}
Available on all commands: --json for machine-readable output, --refresh to
bypass the cache, --debug to attach a redacted response snippet to errors.
| Command | Flags |
|---|---|
search |
--num <n> (8), --country <XX> (US), --near "<place>" (off by default) |
fetch |
--format md|html|text, --max-chars <n> (40000), --section <n>, --grep <pattern>, --no-escalate, --max-paid-rungs <n> (2), --render, --unblock |
crawl |
--limit <n> (20), --depth <n> |
--section and --grep read from the cache and are free. Run
spider-data --help for the full list.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Success, including a search that matched nothing |
| 2 | Usage error |
| 3 | Cross-site redirect, returned rather than followed |
| 4 | Target refuses anonymous access |
| 5 | Target does not exist |
| 6 | Transient |
| 7 | API key rejected |
| 8 | No credit remaining |
| 9 | Rate limited |
| 10 | Refused by local policy (SSRF guard, robots.txt, scheme) |
Cost tracking
spider-data stats --since 7d
paid: true in the audit log means a request left the machine. Cache hits and
free revalidations are never marked paid. The only function that can mark an
attempt paid is the one that opens the socket.
Security
The API key is never written to process.env, never read from a .env found by
walking up the directory tree, and never included in an error message. A single
redactor guards all writes to stdout, stderr, the cache, and the log.
test/leak.test.js verifies this against a server that reflects request headers
into every response body.
DNS is resolved and each address is checked against the private and reserved ranges before any request, so the API cannot be used as an SSRF relay. robots.txt is honoured by default.
Tests
npm test
npm run leakcheck
184 tests. No network access or API key required.
Documentation
docs/DESIGN.md covers block detection, the escalation ladder,
and the assumptions the package makes rather than verifies.
License
MIT. See LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。