crawlcheck-mcp-server

crawlcheck-mcp-server

Enables crawling websites in a real Chromium browser to report JavaScript console errors, uncaught exceptions, failed requests, broken links, and accessibility violations, classifying issues as first- or third-party.

Category
访问服务器

README

crawlcheck-mcp-server

MCP server for crawlcheck. Crawls a website in a real Chromium browser and reports JavaScript console errors, uncaught exceptions, failed requests, broken links, and accessibility violations — classifying console errors and failed requests as first-party (your code) or third-party (embeds and scripts you don't control).

Runs locally over stdio, so crawls use your machine's browser and nobody pays for hosted compute.

Prerequisite

npx playwright install chromium

Without it the tools return: Chromium is not installed for Playwright. Run: npx playwright install chromium.

Configuration

Claude Desktop / Claude Code — .mcp.json:

{
  "mcpServers": {
    "crawlcheck": {
      "command": "npx",
      "args": ["-y", "crawlcheck-mcp-server"]
    }
  }
}

Cursor — same shape in mcp.json.

Tools

Tool Purpose
crawlcheck_scan_site Crawl a site. Returns a summary, the top 20 issues, and a scan_id. Never returns the full issue list.
crawlcheck_check_page Check a single page without following links — the fast path for "is this page broken?".
crawlcheck_get_issues Filtered, paginated issues from a completed scan. Reads from cache; never re-crawls, never blocks.

Typical flow: scan_site for the summary, then get_issues to drill in without re-crawling or flooding the context window.

Two behaviours worth knowing

Classification is not universal. Only console-error and failed-request carry a classification, because those are the only issue types for which the engine computes a resource origin. page-error, broken-link and a11y have no classification field at all — they are reported as unclassified, which does not mean third-party. crawlcheck_get_issues therefore defaults to classification: "all"; filtering to "first-party" would silently hide every broken link and accessibility violation.

One crawl at a time. A second concurrent scan is rejected immediately rather than queued:

A crawl is already running. crawlcheck runs one crawl at a time — wait for the current scan to finish, then retry. If you have a scan_id from an earlier scan, use crawlcheck_get_issues instead; it reads from cache and never blocks.

This is a correctness requirement, not throttling. runCrawl in crawlcheck 0.2.x redirects the global console.log while crawling (so that progress output can't corrupt this server's JSON-RPC frames on stdout), and overlapping calls would interleave that redirect. Rejecting beats queueing because MCP clients time out tool calls anyway, so a queued crawl usually dies waiting.

Exit semantics

summary.exitCode mirrors crawlcheck's process exit code and is only ever 0 or 1. Accessibility violations are advisory and do not affect it unless strict: true. Third-party console errors and failed requests are excluded unless include_third_party: true.

Development

npm install        # crawlcheck resolves to ../crawlcheck until 0.2.0 is on npm
npm run build      # generates src/types/report.v1.d.ts, then compiles
npm test           # builds, then runs the unit tests
npm run check:types  # regenerate types and fail if they drift from the schema
npm run inspect    # MCP Inspector against the built server

src/types/report.v1.d.ts is generated from crawlcheck's published JSON Schema and committed. Don't hand-edit it — npm run check:types regenerates and fails on any diff, so a crawlcheck schema bump can't slip through unnoticed. The only hand-written declaration is runCrawl's signature in src/types/crawlcheck.d.ts, which deliberately omits crawlcheck's internal raw field so that reaching for it is a compile error.

Verifying

Unit tests cover ranking, digest limits, cache TTL/LRU, single-flight (including that a failed crawl releases the lock), and error mapping. What they can't cover is a real crawl — for that:

npm run build
npx @modelcontextprotocol/inspector node dist/index.js

In Inspector v2.x the server appears in a Servers list as Disconnected — flip the toggle beside it to start the process before the tools show up.

Concurrency is the one behaviour the Inspector can't exercise, because it sends a single tool call at a time. For that:

node test/live-concurrency.mjs [url]

It drives the built server over stdio and fires a second crawlcheck_scan_site while the first is still crawling. Expected: the second is rejected with "A crawl is already running", the first completes, and a third succeeds afterwards — proving the lock releases rather than wedging the server. Exits non-zero on failure. Excluded from npm test because it launches a browser and hits the network.

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选