claude-code-docs
Searches Claude Code documentation via BM25 indexing, returning ranked snippets with category filtering. Enables fast, local search over official docs without full-document scans.
README
claude-code-docs MCP Server
Version: 1.0.0
Runtime: Node.js >=18 (TypeScript, ESM)
Key dependencies: @modelcontextprotocol/sdk, zod, yaml, stemmer
License: Not specified in this package
Problem Statement
Claude Code's documentation is large and frequently updated, but most MCP clients need fast, local search results rather than full-document scans. This server fetches the official docs, chunks them into semantic sections, builds an in-memory BM25 index, and exposes MCP tools that return ranked snippets.
The result is a small, focused MCP server that provides deterministic, query-focused results with minimal client integration surface: a stdio transport, four tools, and a cache-backed indexing pipeline.
Quick Start
- From this directory, install dependencies:
npm install - Build the server:
npm run build - Start the MCP server (stdio transport):
npm start
To use the server from an MCP client, see Client Configuration below.
How It Works
Pipeline overview:
- Fetch official docs from the configured URL and parse
Source:markers into sections. - Synthesize frontmatter (topic/id/category) and chunk each section at semantic boundaries.
- Tokenize and build a BM25 index (with heading-based score boosting).
- Serve MCP tool calls against the in-memory index.
Design properties:
- Two-cache model: a raw content cache (TTL-based) and a serialized index cache (version-gated).
- Fail-open on expected fetch/validation errors by falling back to stale cache when allowed.
- Fail-closed on programmer errors to avoid masking regressions.
- Concurrency-safe index loading with a shared in-flight promise and retry backoff.
Configuration
Environment variables:
| Variable | Default | Purpose | Constraints / Behavior |
|---|---|---|---|
DOCS_URL |
https://code.claude.com/docs/llms-full.txt |
Source documentation URL. | Validated on startup; must be a valid https URL. |
DOCS_TRUST_MODE |
official |
Trust mode controlling source validation and canary policy. | official: pins source to code.claude.com, full canary evaluation (fallback-segment delta + relative-drift checks + absolute fallback-ratio warn). unsafe: accepts any HTTPS URL, structural canaries only (count + size checks). Use unsafe only for local testing or private mirrors. |
RETRY_INTERVAL_MS |
60000 |
Retry backoff for failed index loads. | Validated on startup; must be an integer between 1000 and 600000. |
CACHE_TTL_MS |
86400000 |
Content cache freshness window in milliseconds. | Integer >=0. 0 means the cache is never considered fresh (fetch each load); values > 1 year are capped. |
DOCS_CACHE_MAX_STALE_MS |
0 |
Maximum allowed age for stale cache fallback. | Validated on startup; must be an integer >=0. 0 disables the limit. |
MIN_SECTION_COUNT |
(unset) | Override for the canary's index floor. | Integer >=0. Unset → canary uses its trust-mode default (official: 40, unsafe: 3). 0 disables the index floor. Does NOT affect the content-cache write guard, which is fixed at 40 (CACHE_WRITE_MIN_SECTIONS) and is NOT env-disableable. |
MAX_INDEX_CACHE_BYTES |
52428800 |
Max serialized index size in bytes before writing cache. | Validated on startup; must be an integer >0. If exceeded, index cache write is skipped (server keeps in-memory index). |
MAX_RESPONSE_BYTES |
10485760 |
Max HTTP response size in bytes. | Integer >0. If declared or streamed size exceeds, fetch fails and falls back to stale cache when available. |
FETCH_TIMEOUT_MS |
30000 |
HTTP fetch timeout in milliseconds. | Integer >=0. 0 results in immediate timeout. |
CACHE_PATH |
unset | Override the content cache file path. | Must include a filename (not just a directory). Does not move the index cache. |
XDG_CACHE_HOME |
unset | Base cache directory for defaults. | When set, affects default content and index cache paths. |
Default cache locations:
- macOS content cache:
~/Library/Caches/claude-code-docs/llms-full.txt - macOS index cache:
~/Library/Caches/claude-code-docs/llms-full.index.json - Linux content cache:
$XDG_CACHE_HOME/claude-code-docs/llms-full.txt(or~/.cache/claude-code-docs/llms-full.txt) - Linux index cache: same directory,
llms-full.index.json
Notes:
CACHE_PATHoverrides only the content cache file path. The index cache always uses the default cache directory derived fromXDG_CACHE_HOMEor OS defaults.- Content cache writes use a lock file (
.lock) to coordinate concurrent writers.
Tools
search_docs
Searches the indexed Claude Code docs.
Parameters:
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
query |
string | yes | - | Max 500 chars, trimmed, must be non-empty. |
limit |
integer | no | 5 |
1-20. |
category |
string | no | - | Canonical categories or aliases (see below). |
Canonical categories:
hooks, skills, commands, agents, plugins, plugin-marketplaces, mcp, channels, settings, memory, overview, getting-started, cli, best-practices, interactive, security, providers, gateways, ide, ci-cd, automation, agent-sdk, desktop, integrations, config, operations, troubleshooting, changelog, uncategorized
Aliases:
subagents -> agents, sub-agents -> agents, slash-commands -> commands, claude-md -> memory, configuration -> config, gateway -> gateways
Return shape:
| Field | Type | Description |
|---|---|---|
results[] |
object | Array of matches. |
results[].chunk_id |
string | Chunk identifier. |
results[].content |
string | Full chunk content. |
results[].snippet |
string | Snippet best matching the query. |
results[].category |
string | Derived category. |
results[].source_file |
string | Source URL/path. |
meta |
object | Index provenance attached to each search response. |
meta.trust_mode |
string | Active trust mode: official or unsafe. |
meta.source_kind |
string or null | How content was obtained: fetched, cached, stale-fallback, or bundled-snapshot. Null if no corpus loaded. |
meta.index_created_at |
string or null | ISO timestamp when the BM25 index was built. Null if not yet loaded. |
meta.corpus_age_ms |
integer or null | Milliseconds since the corpus content was obtained (Date.now() - corpus.obtainedAt). Null if no corpus loaded. |
error |
string | Present only on failure. |
reload_docs
Forces a refresh of the docs and rebuilds the index.
Parameters: none.
Return:
- Text message indicating success, chunk count, and any parse warnings.
get_status
Returns a lightweight runtime status snapshot. Use this to check index health, trust configuration, and canary evaluation results without triggering a reload or dumping the full metadata.
Parameters: none.
Return shape:
| Field | Type | Description |
|---|---|---|
trust_mode |
string | Active trust mode: official or unsafe. |
docs_origin |
string | Hostname of the documentation source URL. |
docs_url |
string | Full documentation source URL. |
source_kind |
string or null | How content was obtained: fetched, cached, stale-fallback, or bundled-snapshot. Null if no corpus loaded. |
index_created_at |
string or null | ISO timestamp when the BM25 index was built. Null if not yet loaded. |
corpus_age_ms |
number or null | Milliseconds since corpus content was obtained. Null if no corpus loaded. |
corpus_obtained_at |
string or null | ISO timestamp when corpus content was obtained. Null if no corpus loaded. |
last_load_attempt_at |
string or null | ISO timestamp of the most recent load attempt. Null if never attempted. |
last_load_error |
string or null | Error message from the most recent failed load. Null if last load succeeded. |
warning_codes |
string[] | Active warning codes: fallback_segment_drift, fallback_ratio_high, parse_issues, section_count_drift, stale_corpus. |
is_loading |
boolean | Whether a load/reload is currently in progress. |
dump_index_metadata
Returns structured index metadata useful for debugging ingestion, category mapping, and chunk coverage without dumping the full corpus.
Parameters: none.
Return shape:
| Field | Type | Description |
|---|---|---|
index_version |
string | Serialized index format version. |
built_at |
string | ISO timestamp for the response build time. |
docs_epoch |
string or null | Content hash for the currently loaded docs. |
categories[] |
object | Per-category chunk metadata. |
categories[].name |
string | Canonical category name. |
categories[].aliases |
string[] | Accepted aliases for the category. |
categories[].chunk_count |
integer | Number of chunks in the category. |
categories[].chunks[] |
object | Chunk-level metadata for debugging and inventory building. |
Resources
None.
Transport
The server uses stdio transport via the MCP SDK.
Client Configuration
Example .mcp.json entry:
{
"mcpServers": {
"claude-code-docs": {
"command": "node",
"args": ["/absolute/path/to/claude-code-docs/dist/index.js"]
}
}
}
Tests
Run:
npm test
The suite covers parser, chunker, loader, lifecycle, fetcher, metadata, and cache behavior. Special cases:
tests/integration.test.tsis skipped unlessINTEGRATION=1.tests/corpus-validation.test.tsdepends on a populated content cache.
Known Limitations
- Stdio transport only; no HTTP/SSE transport.
- No background refresh loop; use
reload_docsfor refreshes. - Category filtering is limited to the predefined list above.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。