readability-mcp

readability-mcp

Transforms already-rendered HTML into LLM-friendly Markdown with metadata, using Mozilla Readability, Turndown, and DOMPurify. No outbound requests; ideal for post-JavaScript content extraction.

Category
访问服务器

README

readability-mcp

Turn already-rendered HTML (captured post-JavaScript from a browser or chrome-devtools MCP) into clean, LLM-friendly Markdown + metadata, using Mozilla Readability, Turndown, and DOMPurify.

The key idea: rendering and extraction are decoupled. A real browser (chrome-devtools) owns rendering; this server only transforms the HTML it is handed. The server makes no outbound requests — there is no fetch, no SSRF surface. The optional url is origin context only, used to absolutize relative links; it is never fetched.

Install

npm install readability-mcp
# or run on demand:
npx readability-mcp

Requires Node >= 22. Build from source:

git clone <repo> && cd readability-mcp
npm install
npm run build      # bundles to dist/index.js
node dist/index.js # starts the stdio MCP server

The chrome-devtools handoff

The motivating flow is two hops — each tool does the one thing it is best at:

// 1. In the chrome-devtools MCP, grab the RENDERED document (post-JS):
mcp__chrome-devtools__evaluate_script({
  function: () => document.documentElement.outerHTML,
});

// 2. Hand the returned HTML string to readability-mcp.
//    `url` is OPTIONAL context (origin for absolutizing relative links) — never fetched.
mcp__readability__extract({ html: "<that string>", url: pageUrl });

This matters most for SPAs and JS-augmented pages, where the initial HTML is an empty <div id="root"> and only the post-JS DOM has the content.

MCP client config

Add to your MCP client config (Claude Code, Claude Desktop, etc.):

{
  "mcpServers": {
    "readability": {
      "command": "npx",
      "args": ["-y", "readability-mcp"]
    }
  }
}

Tools

Both tools return MCP structured content (schemaVersion, metadata, diagnostics) validated by a zod outputSchema, plus a human/LLM-readable payload in content[0].text. Nothing throws across the wire — failures become { "isError": true } results.

extract — primary tool

Extracts the main article from rendered HTML and returns Markdown + metadata + diagnostics.

Option Default Description
html (required) — Rendered HTML (post-JS), e.g. document.documentElement.outerHTML.
url — Optional origin. Never fetched; used to absolutize relative links/images.
format markdown markdown | html | text | json. json emits {metadata, content, diagnostics}.
metadataMode none none | yaml | json — prepend a metadata block to the markdown/text payload.
extraction balanced balanced | aggressive | conservative — maps to Readability's scorer knobs.
selectors.include — Restrict extraction to a subtree: "main", "article", ".post".
selectors.exclude — Strip boilerplate before Readability: ["nav", "footer", "[role=banner]"].
maxNodes — Perf/safety cap = Readability maxElemsToParse.
minArticleLength — Semantic alias for Readability charThreshold.
gfm true Tables, strikethrough, task lists.
headingStyle atx atx (#) | setext (underlining).
codeBlockStyle fenced fenced (```) | indented.
images keep keep | drop | src-only (bare URL) | reference (link-ref style).
sanitize true Run DOMPurify on the article HTML.
maxChars — Truncate the payload at a block boundary — never inside a fenced code block.
wordsPerMinute 200 For readingTimeMin.
keepClasses false Retain all classes (default strips non-language classes).
readabilityOverrides — Escape hatch — passed verbatim to new Readability(doc, …). Unstable.

Fallback. If Readability's parse() returns no article (e.g. an app shell or image-only page), a selector cascade salvages the first usable root — article → main → [role=main] → largest text-dense block → body — and reports diagnostics.fallbackUsed: true with extractedNode naming the root that was used.

Metadata cascade. Each metadata field is resolved by priority: JSON-LD → OpenGraph → Twitter → <meta>/<time> → Readability → <title> (first non-empty value wins).

html_to_markdown — fragment path

Converts an arbitrary HTML fragment to Markdown without Readability scoring (e.g. a snippet already isolated via chrome-devtools). Same Turndown + DOMPurify path; reports fallbackUsed: true, extractedNode: "fragment". Shares the format, gfm, headingStyle, codeBlockStyle, images, sanitize, maxChars, wordsPerMinute, selectors, and url options. Metadata is minimal (url, wordCount, readingTimeMin, and a title from the fragment's first heading).

outline — heading pre-check

Returns the document outline (h1–h6 in document order with stable anchor ids) as a cheap "is this worth reading?" / "where's the section about X?" pre-check before paying for full extraction. Runs no Readability, Turndown, or sanitization — a pure heading walk over the normalized DOM.

Option Default Description
html (required) — Rendered HTML (post-JS), e.g. document.documentElement.outerHTML.
url — Optional origin. Never fetched; carried through to metadata.url.

Output shape: structuredContent.outline = [{level, text, anchor}] plus an indented-bullet TOC rendered into content[0].text, and metadata = {title?, url?} (title falls back from <title> to the first <h1>). Anchor precedence: the heading's own id, then a descendant permalink's #fragment, then a slug of the text (deduped -1, -2, … for generated slugs only — author ids are kept verbatim).

Diagnostics

structuredContent.diagnostics exposes: readerable, extractedNode, fallbackUsed, removedNodes (element delta vs. the document), sanitization.{scripts,iframes} (counted across the whole pipeline), and truncated.

Payload size (stdio) — DESIGN §11.1

A full rendered SPA can be several MB as a string, and MCP tool args travel over JSON-RPC on stdio. Mitigations:

  • Scoped capture (recommended) — real pages are large (hundreds of KB of outerHTML), so capture only what you need rather than the whole document. Via chrome-devtools evaluate_script, grab document.head.outerHTML (for metadata) plus a content subtree such as document.querySelector('article')?.outerHTML || document.querySelector('main')?.outerHTML, and pass that to extract with url. url absolutizes relative links within whatever HTML is passed.
  • selectors.include — scope to the article subtree, e.g. "main", so only the relevant DOM is scored and serialized.
  • maxChars — cap the returned payload; truncation lands at a block boundary and never splits a fenced code block.
  • maxNodes — a hard cap on elements parsed (Readability.maxElemsToParse) for very large documents.

Development

npm run typecheck   # tsc --noEmit
npm run build       # vite build -> dist/index.js
npm run lint        # eslint
npm test            # vitest run
npm run test:update-goldens   # UPDATE_GOLDENS=1 vitest run

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选