esp32-docs-mcp

esp32-docs-mcp

An MCP server providing semantic search over ESP32 documentation, including ESP-IDF guides, API references, and Technical Reference Manuals, enabling coding agents to look up ESP32 facts locally with accurate, chip-specific results.

Category
访问服务器

README

esp32-docs-mcp

Semantic search over ESP32 documentation, exposed as an MCP server — so a coding agent can look things up in the ESP-IDF guides, the API reference, and the chip Technical Reference Manuals while it works.

Everything runs locally. Embeddings are computed on-device with MLX (Apple Silicon) and stored in LanceDB; no documentation or query text leaves the machine.

The current corpus is 21,672 chunks — 11,157 from ESP-IDF across nine build targets, 10,515 from ten Technical Reference Manuals.

Why this exists

Pointing a retrieval system at ESP-IDF's .rst sources produces confidently wrong answers, because ESP-IDF documentation is built, not written. The build is where the facts come from:

  • {IDF_TARGET_CONFIG_BOOTLOADER_OFFSET_IN_FLASH} and its siblings are substituted at build time from the chip's soc_caps headers and Kconfig. Parse the source and you embed the placeholder instead of the address.
  • only:: blocks select per-chip branches. Parse the source and mutually exclusive branches for nine different chips collapse into one self-contradictory passage.
  • The entire API reference is generated by doxygen from C headers, and does not exist in the source tree at all.

So this project builds the docs once per chip and indexes the build output. Constants are real values, chip-specific claims are correct per chip, and function signatures, parameters, structs and enums are all present.

Chunks byte-identical across every target are stored once and tagged with every chip they appeared under, so filtering to a chip narrows results to what is true for it without hiding general content. That collapses 43,274 per-target chunks to 11,157 unique (74.2%) while keeping every genuine per-chip difference — a chunk that differs by one substituted constant stays a separate, correctly narrower row.

The Technical Reference Manuals come from Espressif's public LaTeX sources rather than the compiled PDFs, at 99.9% register capture. Registers are indivisible: name, address and every bitfield stay in one chunk. Where Espressif publishes more than one manual for a chip — ESP32-P4 has both a mainline TRM and a "Chip Revision v1.3" one, whose register sets diverge in both directions — both are ingested and each result states which silicon it applies to.

Requirements

  • macOS on Apple Silicon (MLX)
  • Python 3.12+ and uv
  • To build the ESP-IDF half of the corpus: an ESP-IDF checkout installed with the docs feature, plus doxygen
  • To build the TRM half: a clone of Espressif's TRM LaTeX repo

There is no prebuilt index to download. See docs/building-the-corpus.md — budget roughly 20 minutes of docs build per chip and a single embedding pass measured in hours.

Install

git clone https://github.com/ozonejunkieau/esp32-docs-mcp.git
cd esp32-docs-mcp
uv sync

Quickstart

Build the corpus once (full instructions), confirm it, then register the server:

uv run validate_store.py     # row count, doc_type breakdown, search round-trip
uv run mcp_server.py         # speaks stdio; Ctrl-C to stop

If you have just (>= 1.43), every command in the documentation has a recipe — just on its own lists them, grouped, and just --list records what a healthy result from each check looks like.

Register with Claude Code:

claude mcp add esp32-docs -- uv run --directory /path/to/esp32-docs-mcp mcp_server.py

Or in claude_desktop_config.json:

{
  "mcpServers": {
    "esp32-docs": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/esp32-docs-mcp", "mcp_server.py"]
    }
  }
}

The embedding model and database connection load once at startup and are reused, so the first request pays the model load and subsequent queries are fast. docs/usage.md covers what good queries look like and how to read the results.

Tools

esp32_docs_search

Parameter Description
query Natural-language query, e.g. "how does I2S clock configuration work"
doc_type trm, idf, or omit for both. A third value, src, is reserved for the ESP-IDF SoC-header corpus that is landing
chip e.g. esp32p4. Narrows to what's true for that chip; content common to all chips still matches
revision Silicon revision, e.g. v1.3 or mainline. Narrows any content that has a revision — the manuals, and ESP32-P4's SoC register headers. Content without one is unaffected. Omit to see every revision
k Results to return (1–20, default 5)

Returns JSON. Each result carries its text plus source_doc, section_path, file_path and chunk_index for citation, relevance_distance (lower is closer), and three reference lists for follow-up lookups: file_refs (ESP-IDF source paths), doc_refs (other doc pages), symbol_refs (C/C++ symbols).

Two fields matter when accuracy does:

  • revision_scope — null for ESP-IDF content, "all published revisions (…)" for TRM content true of every stepping, or "ONLY revision X …" where it isn't. Applying a revision-specific register definition to the wrong silicon is a hardware bug, and this field says which case you have without you having to reason about a list.
  • source_version — the upstream revision the chunk was built from. Both corpora track moving upstreams, so this is what makes an answer reproducible.

esp32_docs_list_chips

Valid chip values with their coverage: has_idf_docs, has_trm, and revisions. Both coverage flags can be false independently — a chip may have a Technical Reference Manual but no per-chip docs build, or the reverse.

Status and limitations

Both pipelines work end to end and the store is populated. Known limits, stated plainly:

  • Apple Silicon only. Embedding goes through MLX. Nothing else is platform-specific, but there is no fallback backend.
  • You must build the corpus yourself. The LanceDB store is derived data and is not committed; the ESP-IDF half additionally needs a full toolchain install per target, because the docs build runs idf.py set-target to extract the constants.
  • The corpus is a snapshot. ESP-IDF docs change weekly and the TRM sources run ahead of the published PDFs. Every row records source_version and source_commit; refreshing is a manual re-run.
  • Coverage is uneven by design. Nine of the thirteen known chips have an ESP-IDF docs build; ten have a published TRM. esp32_docs_list_chips reports which.
  • Diagrams are not recoverable. Block, timing and bytefield diagrams survive only as their captions. No register bit layout was found to exist only as a figure, but the diagrams themselves are absent.
  • TRM chunks have no doc_refs. Cross-reference extraction is implemented for the ESP-IDF corpus only; it is deferred, not broken.
  • A register census shortfall is not automatically content loss. Some source registers are disabled upstream or tagged for a different chip, and the parser is right to drop them while an independent checker may still count them. Read the named registers before treating a shortfall as a regression — see docs/trm-latex.md.

Documentation

Document For
docs/usage.md Wiring the server in, querying it, reading results, troubleshooting
docs/building-the-corpus.md Building and refreshing the index from source
docs/architecture.md Why the pipelines are shaped this way; what was tried and deleted
docs/trm-latex.md TRM LaTeX ingest — design record and verification
docs/development.md The justfile, the tests, the fixtures
CLAUDE.md Operating guide: invariants, environment traps, commands, current state

Layout

File Purpose
mcp_server.py The MCP server and its two tools
build_idf_docs.sh Builds ESP-IDF docs to XML for one or more chips
sphinx_xml.py Turns a built page's XML into chunks
chunking.py Heading-aware chunk assembly, shared by both corpora
ingest_sphinx_xml.py Chunks one target's build into JSONL
dedup_chunks.py Collapses chunks identical across chips
embed_and_store.py Embeds chunks into LanceDB
embedder.py Qwen3-Embedding-4B via MLX (2560 dims)
schema.py The stored chunk schema
chips.yaml, chip_vocab.py Verified chip vocabulary and doc coverage
latex_parser.py, latex_coverage_check.py TRM LaTeX parsing and macro-coverage reporting — see docs/trm-latex.md
ingest_trm.py Chunks the TRM manuals, deduplicated across silicon revisions
ingest_source.py Chunks the ESP-IDF SoC headers (doc_type = "src")
provenance.py, backfill_provenance.py Record/repair the upstream revision each chunk came from
check_thin_files.py Thin-file and capture-rate checks (xml / latex)
register_census.py, trm_verify.py Check TRM registers survive into chunks
make_trm_fixture.py Synthetic corpora that prove the checkers work
validate_store.py LanceDB row count, samples, search round-trip
justfile Every documented command as a recipe; just --list
tests/ uv run pytest; slow tests need a local corpus and skip without one

License

The code in this repository is MIT — see LICENSE.

Upstream licensing of the corpus

The licence above covers this software, not the documentation it indexes. Building the corpus produces a store containing verbatim documentation text, so redistributing that store means redistributing Espressif's documentation under its own terms — not this project's.

Source Licence
ESP-IDF documentation Apache-2.0
TRM LaTeX sources CC-BY-SA 4.0 (the repo's scripts are Apache-2.0)

The practical consequence is CC-BY-SA is copyleft. Roughly half the corpus — 10,515 of 21,672 chunks — comes from the Technical Reference Manuals, so a built store carries a ShareAlike obligation and cannot simply be republished under a permissive licence. If you publish a store or a derived dataset:

  • Keep the two corpora as separate artifacts. A combined store is more readily argued to be adapted material, where ShareAlike propagates; separable datasets are more clearly a collection.
  • Licence the TRM artifact CC-BY-SA 4.0 and the ESP-IDF artifact Apache-2.0, each attributing Espressif and naming the upstream commit. The store already records source_version and source_commit on every row, so the exact provenance is available per chunk.
  • Preserve attribution and copyright notices in the indexed text. ESP-IDF's COPYRIGHT document carries third-party attributions, including contributor email addresses, that the Apache-2.0 and BSD terms require be retained — they are not incidental content to strip.

Nothing here is legal advice, and whether a mixed store counts as a collection or an adaptation turns on specifics. Get a real opinion before relying on it commercially.

This project distributes no corpus. Everything is built locally from sources you obtain yourself, which is why there is no prebuilt index to download.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选