mcp-semantic-search
Enables semantic search over local markdown specs, RFCs, and docs using Ollama and LanceDB, allowing AI agents to find relevant information by meaning rather than exact keywords.
README
🔍 mcp-semantic-search
Ask your specs a question instead of grepping them.
An MCP server & CLI that gives AI agents local semantic search over markdown docs — running 100% on your machine by default.
Why use this?
Traditional grep misses ideas that don't match exact keywords.
mcp-semantic-search understands intent across your local specs, RFCs, and internal docs.
$ mcp-semantic-search search "how do we stop repeated failed logins"
--- Result 1 [0.688] content ---
file: auth.md
section: Sessions > Credential attempts
After five consecutive bad passwords the account enters a 15-minute
cooling-off period. The counter resets on any successful sign-in.
--- Result 2 [0.541] content ---
file: gateway.md
section: Rate limits > Edge throttling
A single IP address is capped at 20 requests per minute against
/session endpoints. Anything beyond that receives a 429 before it
ever reaches the application.
Notice: Neither chunk contains the words "failed" or "login" —
grep -ri "failed login"returns nothing at all. Semantic search finds both halves of the answer: the account lockout in the auth spec, and the network throttle that backs it up in a different file with entirely different vocabulary.
Key Features
- 🧱 Structure-Aware Chunking — Splits markdown on heading boundaries (
H1–H4) without breaking code blocks or tables. - 🔒 Local & Private — Runs via LanceDB & Ollama. Zero docs leave your machine.
- 🗺️ Document Maps — Generates outline (
toc) chunks so agents can scan doc structures before reading details. - ⚡ Zero-Overhead Indexing — File hashes ensure re-indexing only happens when files actually change.
- 🎯 Targeted Filtering — Filter search by filename, heading, or chunk type (
content,code,table,toc).
Quick Start
1. Requirements & Build
Requires Node.js 22+ and Ollama (or a Gemini API key).
# Pull default model
ollama pull qwen3-embedding:0.6b
# Clone & Build
git clone <repo-url> mcp-semantic-search
cd mcp-semantic-search
npm install && npm run build
2. Run locally (CLI)
Inside the target repository you want to index:
cd ~/projects/my-app
echo '.mcp-search/' >> .gitignore
/path/to/mcp-semantic-search index
/path/to/mcp-semantic-search search "how are expired sessions cleaned up"
Setup as an MCP Server
Add to your project's .mcp.json:
{
"mcpServers": {
"specs": {
"type": "stdio",
"command": "node",
"args": ["/absolute/path/to/mcp-semantic-search/dist/index.js"]
}
}
}
MCP Tools Exposed
index— Indexes your markdown specs directory (setreindex: trueto force rebuild).search— Queries indexed documents (query,file,section,chunk_type,limit,min_score).status— Checks index health, chunk counts, and staleness.
Agent Prompting Tip (
CLAUDE.md): Add this to your project'sCLAUDE.md: "Search specs using thespecsMCP server. Runindexfirst ifstatusshows the index as missing or stale."
Configuration
Set environment variables in your shell or directly inside .mcp.json under the "env" block.
Global Settings
| Variable | Default | Purpose |
|---|---|---|
EMBEDDING_BACKEND |
ollama |
Vector provider: ollama or gemini |
SPECS_DIR |
<project>/specs |
Absolute path to markdown specs |
DB_PATH |
<project>/.mcp-search |
Where LanceDB index is stored |
MIN_SCORE |
0.44 |
Default similarity threshold |
Ollama Backend (Default)
| Variable | Default | Purpose |
|---|---|---|
OLLAMA_BASE_URL |
http://localhost:11434 |
Endpoint for Ollama daemon |
OLLAMA_EMBEDDING_MODEL |
qwen3-embedding:0.6b |
Embedding model to use |
Gemini Backend (Cloud Option)
| Variable | Default | Purpose |
|---|---|---|
GEMINI_API_KEY |
(Required) | Required when EMBEDDING_BACKEND=gemini |
GEMINI_EMBEDDING_MODEL |
gemini-embedding-001 |
Embedding model to use |
GEMINI_EMBEDDING_DIMENSIONS |
768 |
Vector width (128–3072) |
Note: Changing backends automatically triggers a clean index rebuild on the next run.
Example .mcp.json with Custom Config
{
"mcpServers": {
"specs": {
"type": "stdio",
"command": "node",
"args": ["/absolute/path/to/mcp-semantic-search/dist/index.js"],
"env": {
"SPECS_DIR": "/absolute/path/to/my-app/docs",
"EMBEDDING_BACKEND": "gemini",
"GEMINI_API_KEY": "your-api-key-here"
}
}
}
}
CLI Options
mcp-semantic-search search <query> [options]
| Flag | Description | Default |
|---|---|---|
--file <str> |
Filter by matching filename | — |
--section <str> |
Filter by matching section heading | — |
--chunk-type <type> |
content | code | table | toc |
All |
--limit <n> |
Max results to return | 5 |
--min-score <f> |
Similarity floor (0.0–1.0) | 0.44 |
--json |
Output raw JSON instead of plain text | false |
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器