crawlcheck-mcp-server
Enables crawling websites in a real Chromium browser to report JavaScript console errors, uncaught exceptions, failed requests, broken links, and accessibility violations, classifying issues as first- or third-party.
README
crawlcheck-mcp-server
MCP server for crawlcheck. Crawls a website in a real Chromium browser and reports JavaScript console errors, uncaught exceptions, failed requests, broken links, and accessibility violations — classifying console errors and failed requests as first-party (your code) or third-party (embeds and scripts you don't control).
Runs locally over stdio, so crawls use your machine's browser and nobody pays for hosted compute.
Prerequisite
npx playwright install chromium
Without it the tools return: Chromium is not installed for Playwright. Run: npx playwright install chromium.
Configuration
Claude Desktop / Claude Code — .mcp.json:
{
"mcpServers": {
"crawlcheck": {
"command": "npx",
"args": ["-y", "crawlcheck-mcp-server"]
}
}
}
Cursor — same shape in mcp.json.
Tools
| Tool | Purpose |
|---|---|
crawlcheck_scan_site |
Crawl a site. Returns a summary, the top 20 issues, and a scan_id. Never returns the full issue list. |
crawlcheck_check_page |
Check a single page without following links — the fast path for "is this page broken?". |
crawlcheck_get_issues |
Filtered, paginated issues from a completed scan. Reads from cache; never re-crawls, never blocks. |
Typical flow: scan_site for the summary, then get_issues to drill in without re-crawling or
flooding the context window.
Two behaviours worth knowing
Classification is not universal. Only console-error and failed-request carry a
classification, because those are the only issue types for which the engine computes a resource
origin. page-error, broken-link and a11y have no classification field at all — they are
reported as unclassified, which does not mean third-party. crawlcheck_get_issues therefore
defaults to classification: "all"; filtering to "first-party" would silently hide every broken
link and accessibility violation.
One crawl at a time. A second concurrent scan is rejected immediately rather than queued:
A crawl is already running. crawlcheck runs one crawl at a time — wait for the current scan to finish, then retry. If you have a scan_id from an earlier scan, use crawlcheck_get_issues instead; it reads from cache and never blocks.
This is a correctness requirement, not throttling. runCrawl in crawlcheck 0.2.x redirects the
global console.log while crawling (so that progress output can't corrupt this server's JSON-RPC
frames on stdout), and overlapping calls would interleave that redirect. Rejecting beats queueing
because MCP clients time out tool calls anyway, so a queued crawl usually dies waiting.
Exit semantics
summary.exitCode mirrors crawlcheck's process exit code and is only ever 0 or 1. Accessibility
violations are advisory and do not affect it unless strict: true. Third-party console errors and
failed requests are excluded unless include_third_party: true.
Development
npm install # crawlcheck resolves to ../crawlcheck until 0.2.0 is on npm
npm run build # generates src/types/report.v1.d.ts, then compiles
npm test # builds, then runs the unit tests
npm run check:types # regenerate types and fail if they drift from the schema
npm run inspect # MCP Inspector against the built server
src/types/report.v1.d.ts is generated from crawlcheck's published JSON Schema and committed.
Don't hand-edit it — npm run check:types regenerates and fails on any diff, so a crawlcheck schema
bump can't slip through unnoticed. The only hand-written declaration is runCrawl's signature in
src/types/crawlcheck.d.ts, which deliberately omits crawlcheck's internal raw field so that
reaching for it is a compile error.
Verifying
Unit tests cover ranking, digest limits, cache TTL/LRU, single-flight (including that a failed crawl releases the lock), and error mapping. What they can't cover is a real crawl — for that:
npm run build
npx @modelcontextprotocol/inspector node dist/index.js
In Inspector v2.x the server appears in a Servers list as Disconnected — flip the toggle beside
it to start the process before the tools show up.
Concurrency is the one behaviour the Inspector can't exercise, because it sends a single tool call at a time. For that:
node test/live-concurrency.mjs [url]
It drives the built server over stdio and fires a second crawlcheck_scan_site while the first is
still crawling. Expected: the second is rejected with "A crawl is already running", the first
completes, and a third succeeds afterwards — proving the lock releases rather than wedging the
server. Exits non-zero on failure. Excluded from npm test because it launches a browser and hits
the network.
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。