semantic-search-mcp

semantic-search-mcp

Provides semantic code search over codebases using local embeddings with natural language queries. Supports hybrid search, file watching, and respects .gitignore.

Category
访问服务器

README

Semantic Search MCP Server

An MCP server that provides semantic code search using local embeddings. Search your codebase with natural language queries like "authentication middleware" or "database connection pooling".

Features

  • Hybrid search: Combines vector similarity (Jina code embeddings) with FTS5 keyword matching using Reciprocal Rank Fusion
  • 165+ languages: Tree-sitter parsing for Python, TypeScript, JavaScript, Go, Rust, Java, C/C++, Ruby, PHP, and more
  • Incremental indexing: File watcher automatically detects additions, modifications, and deletions
  • Respects .gitignore: Honors your project's .gitignore files (including nested ones)
  • Auto-initialization: Model loads and codebase indexes in the background on server startup
  • Zero external APIs: All embeddings generated locally with FastEmbed

Installation

uv tool install semantic-search-mcp

Or with pip:

pip install semantic-search-mcp

Or run directly without installing:

uvx semantic-search-mcp

Quick Start

Add to Claude Code

Option A: Project-level config (recommended)

After installing with uv tool install or pip install, create .mcp.json in your project root:

{
  "mcpServers": {
    "semantic-search": {
      "command": "semantic-search-mcp"
    }
  }
}

Option B: CLI

claude mcp add semantic-search -- semantic-search-mcp

Option C: Without installing (ephemeral)

If you prefer not to install, use uvx to run in an ephemeral environment:

{
  "mcpServers": {
    "semantic-search": {
      "command": "uvx",
      "args": ["semantic-search-mcp"]
    }
  }
}

Use

The server auto-initializes on startup.

Available Tools

Tool Description
search_code Search codebase with natural language
get_status Get server state, progress, and statistics
pause_watcher Pause file watching (events discarded)
resume_watcher Resume file watching
reindex Start full reindex (runs in background)
cancel_indexing Cancel running indexing job
clear_index Wipe all indexed data
exclude_paths Add paths to ignore (session-only)
include_paths Remove paths from exclusion list

How It Works

Indexing

On startup, the server:

  1. Scans your codebase for supported file types
  2. Parses code into semantic chunks (functions, classes, methods) using Tree-sitter
  3. Generates embeddings for each chunk using Jina's code embedding model
  4. Stores everything in a local SQLite database with vector search support

File Watching

The server monitors your codebase for changes in real-time:

Event Action
File created Parsed, embedded, and added to index
File modified Re-indexed if content hash changed
File deleted Removed from index

Changes are debounced (default 1s) to batch rapid modifications.

What Gets Indexed

Included:

  • Files with code extensions: .py, .js, .ts, .tsx, .jsx, .go, .rs, .java, .c, .cpp, .h, .rb, .php, .swift, .kt, .scala, and more

Excluded:

  • Files matching .gitignore patterns (all .gitignore files in your project are respected)
  • Common non-code directories: node_modules, __pycache__, .venv, build, dist, .git, vendor, etc.
  • Binary files and non-code file types

Configuration

Environment variables:

Variable Default Description
SEMANTIC_SEARCH_DB_PATH .semantic-search/index.db Index database location
SEMANTIC_SEARCH_EMBEDDING_MODEL jinaai/jina-embeddings-v2-base-code Embedding model
SEMANTIC_SEARCH_MIN_SCORE 0.3 Minimum relevance threshold (0-1)
SEMANTIC_SEARCH_DEBOUNCE_MS 1000 File watcher debounce in milliseconds
SEMANTIC_SEARCH_BATCH_SIZE 50 Files per batch (reduce if running out of memory)
SEMANTIC_SEARCH_MAX_FILE_SIZE_KB 512 Skip files larger than this (KB)
SEMANTIC_SEARCH_EMBEDDING_BATCH_SIZE 8 Texts per embedding call (reduce if OOM)
SEMANTIC_SEARCH_EMBEDDING_THREADS 4 ONNX runtime threads (higher = faster on multi-core)
SEMANTIC_SEARCH_USE_QUANTIZED true Use INT8 quantized model (30-40% faster)

Performance

GPU Acceleration

GPU acceleration is auto-detected and used when available:

Platform Provider Installation
NVIDIA CUDA pip install semantic-search-mcp[gpu]
Apple Silicon CoreML Automatic (M1/M2/M3)
AMD ROCm Install ROCm-enabled onnxruntime
Windows DirectML Install DirectML-enabled onnxruntime

Alternative Models

For faster indexing (with quality tradeoffs), you can use a lighter model:

Model Dimensions Speed Best For
jinaai/jina-embeddings-v2-base-code 768 Baseline Code search (default)
BAAI/bge-small-en-v1.5 384 ~10x faster General text
sentence-transformers/all-MiniLM-L6-v2 384 ~32x faster Speed priority

To use an alternative model:

export SEMANTIC_SEARCH_EMBEDDING_MODEL="sentence-transformers/all-MiniLM-L6-v2"

Note: Changing models requires a full reindex (delete .semantic-search/ directory).

UniXcoder (Experimental)

Microsoft UniXcoder is a code-specific model pre-trained on code + AST + comments. It may provide better semantic understanding of code structure, but is substantially slower (~20x slower than Jina).

Model Dimensions Speed Languages
microsoft/unixcoder-base 768 ~20x slower 6 (java, ruby, python, php, js, go)
microsoft/unixcoder-base-nine 768 ~20x slower 9 (+ c, c++, c#)

Installation (requires additional dependencies):

pip install semantic-search-mcp[unixcoder]

Usage:

export SEMANTIC_SEARCH_EMBEDDING_MODEL="microsoft/unixcoder-base-nine"

When to use UniXcoder:

  • You prioritize search quality over indexing speed
  • Your codebase is small to medium sized
  • You have GPU acceleration (CUDA or Apple Silicon MPS)

When to avoid UniXcoder:

  • Large codebases (10,000+ files) - indexing will take hours
  • You need fast initial indexing
  • Running on CPU without GPU acceleration

Claude Code Integration

Skills and commands are automatically installed when the MCP server first starts:

  • Skills~/.claude/skills/ (AI auto-discovery)
  • Commands~/.claude/commands/ (user-invocable slash commands)

To manually reinstall or update:

semantic-search-mcp-install-skills

Available Slash Commands

Command Description
/semantic-search-search <query> Search codebase with natural language
/semantic-search-status Check server status and index stats
/semantic-search-reindex Trigger full codebase reindex
/semantic-search-cancel Cancel running indexing job
/semantic-search-clear Wipe all indexed data
/semantic-search-pause Pause file watcher
/semantic-search-resume Resume file watcher

Requirements

  • Python 3.11+
  • ~700MB disk for embedding model (downloaded on first run, ~150MB with INT8 quantization)
  • ~1GB RAM for embedding model

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选