readability-mcp
Transforms already-rendered HTML into LLM-friendly Markdown with metadata, using Mozilla Readability, Turndown, and DOMPurify. No outbound requests; ideal for post-JavaScript content extraction.
README
readability-mcp
Turn already-rendered HTML (captured post-JavaScript from a browser or chrome-devtools MCP) into clean, LLM-friendly Markdown + metadata, using Mozilla Readability, Turndown, and DOMPurify.
The key idea: rendering and extraction are decoupled. A real browser (chrome-devtools) owns rendering; this server only transforms the HTML it is handed. The server makes no outbound requests — there is no fetch, no SSRF surface. The optional url is origin context only, used to absolutize relative links; it is never fetched.
Install
npm install readability-mcp
# or run on demand:
npx readability-mcp
Requires Node >= 22. Build from source:
git clone <repo> && cd readability-mcp
npm install
npm run build # bundles to dist/index.js
node dist/index.js # starts the stdio MCP server
The chrome-devtools handoff
The motivating flow is two hops — each tool does the one thing it is best at:
// 1. In the chrome-devtools MCP, grab the RENDERED document (post-JS):
mcp__chrome-devtools__evaluate_script({
function: () => document.documentElement.outerHTML,
});
// 2. Hand the returned HTML string to readability-mcp.
// `url` is OPTIONAL context (origin for absolutizing relative links) — never fetched.
mcp__readability__extract({ html: "<that string>", url: pageUrl });
This matters most for SPAs and JS-augmented pages, where the initial HTML is an empty <div id="root"> and only the post-JS DOM has the content.
MCP client config
Add to your MCP client config (Claude Code, Claude Desktop, etc.):
{
"mcpServers": {
"readability": {
"command": "npx",
"args": ["-y", "readability-mcp"]
}
}
}
Tools
Both tools return MCP structured content (schemaVersion, metadata, diagnostics) validated by a zod outputSchema, plus a human/LLM-readable payload in content[0].text. Nothing throws across the wire — failures become { "isError": true } results.
extract — primary tool
Extracts the main article from rendered HTML and returns Markdown + metadata + diagnostics.
| Option | Default | Description |
|---|---|---|
html (required) |
— | Rendered HTML (post-JS), e.g. document.documentElement.outerHTML. |
url |
— | Optional origin. Never fetched; used to absolutize relative links/images. |
format |
markdown |
markdown | html | text | json. json emits {metadata, content, diagnostics}. |
metadataMode |
none |
none | yaml | json — prepend a metadata block to the markdown/text payload. |
extraction |
balanced |
balanced | aggressive | conservative — maps to Readability's scorer knobs. |
selectors.include |
— | Restrict extraction to a subtree: "main", "article", ".post". |
selectors.exclude |
— | Strip boilerplate before Readability: ["nav", "footer", "[role=banner]"]. |
maxNodes |
— | Perf/safety cap = Readability maxElemsToParse. |
minArticleLength |
— | Semantic alias for Readability charThreshold. |
gfm |
true |
Tables, strikethrough, task lists. |
headingStyle |
atx |
atx (#) | setext (underlining). |
codeBlockStyle |
fenced |
fenced (```) | indented. |
images |
keep |
keep | drop | src-only (bare URL) | reference (link-ref style). |
sanitize |
true |
Run DOMPurify on the article HTML. |
maxChars |
— | Truncate the payload at a block boundary — never inside a fenced code block. |
wordsPerMinute |
200 |
For readingTimeMin. |
keepClasses |
false |
Retain all classes (default strips non-language classes). |
readabilityOverrides |
— | Escape hatch — passed verbatim to new Readability(doc, …). Unstable. |
Fallback. If Readability's parse() returns no article (e.g. an app shell or image-only page), a selector cascade salvages the first usable root — article → main → [role=main] → largest text-dense block → body — and reports diagnostics.fallbackUsed: true with extractedNode naming the root that was used.
Metadata cascade. Each metadata field is resolved by priority: JSON-LD → OpenGraph → Twitter → <meta>/<time> → Readability → <title> (first non-empty value wins).
html_to_markdown — fragment path
Converts an arbitrary HTML fragment to Markdown without Readability scoring (e.g. a snippet already isolated via chrome-devtools). Same Turndown + DOMPurify path; reports fallbackUsed: true, extractedNode: "fragment". Shares the format, gfm, headingStyle, codeBlockStyle, images, sanitize, maxChars, wordsPerMinute, selectors, and url options. Metadata is minimal (url, wordCount, readingTimeMin, and a title from the fragment's first heading).
outline — heading pre-check
Returns the document outline (h1–h6 in document order with stable anchor ids) as a cheap "is this worth reading?" / "where's the section about X?" pre-check before paying for full extraction. Runs no Readability, Turndown, or sanitization — a pure heading walk over the normalized DOM.
| Option | Default | Description |
|---|---|---|
html (required) |
— | Rendered HTML (post-JS), e.g. document.documentElement.outerHTML. |
url |
— | Optional origin. Never fetched; carried through to metadata.url. |
Output shape: structuredContent.outline = [{level, text, anchor}] plus an indented-bullet TOC rendered into content[0].text, and metadata = {title?, url?} (title falls back from <title> to the first <h1>). Anchor precedence: the heading's own id, then a descendant permalink's #fragment, then a slug of the text (deduped -1, -2, … for generated slugs only — author ids are kept verbatim).
Diagnostics
structuredContent.diagnostics exposes: readerable, extractedNode, fallbackUsed, removedNodes (element delta vs. the document), sanitization.{scripts,iframes} (counted across the whole pipeline), and truncated.
Payload size (stdio) — DESIGN §11.1
A full rendered SPA can be several MB as a string, and MCP tool args travel over JSON-RPC on stdio. Mitigations:
- Scoped capture (recommended) — real pages are large (hundreds of KB of
outerHTML), so capture only what you need rather than the whole document. Via chrome-devtoolsevaluate_script, grabdocument.head.outerHTML(for metadata) plus a content subtree such asdocument.querySelector('article')?.outerHTML || document.querySelector('main')?.outerHTML, and pass that toextractwithurl.urlabsolutizes relative links within whatever HTML is passed. selectors.include— scope to the article subtree, e.g."main", so only the relevant DOM is scored and serialized.maxChars— cap the returned payload; truncation lands at a block boundary and never splits a fenced code block.maxNodes— a hard cap on elements parsed (Readability.maxElemsToParse) for very large documents.
Development
npm run typecheck # tsc --noEmit
npm run build # vite build -> dist/index.js
npm run lint # eslint
npm test # vitest run
npm run test:update-goldens # UPDATE_GOLDENS=1 vitest run
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。