Turbovec MCP Server
Provides persistent long-term memory (semantic RAG) for AI coding assistants, enabling them to store and semantically search code and documentation across chat sessions without token limits.
README
<img width="1984" height="576" alt="vnuw9pe8htgfwu87qgfbw" src="https://github.com/user-attachments/assets/f781a60d-11b9-4198-b80b-fca16a927184" />
Turbovec MCP (Long-Term Memory RAG for AI)
Turbovec MCP Server is a Model Context Protocol (MCP) implementation that acts as a persistent, long-term memory (Semantic RAG) for AI coding assistants like Zoo Code, Claude Desktop, and Cursor.
By running this local server, your AI assistant gains the ability to "read", "remember", and "semantically search" through vast amounts of code and documentation across different chat sessions, completely bypassing token limitations.
The Problem it Solves
- Context Window Limits: When working on large projects, pasting hundreds of files into the AI chat will exceed token limits or cause the AI to hallucinate.
- AI Amnesia (Stateless Chats): Whenever you start a new chat tab, the AI forgets everything you discussed in the previous session (e.g., project architecture, specific coding guidelines).
- Literal Search vs. Semantic Search: Standard file search (CTRL+F) requires exact keyword matches. This server allows the AI to search by meaning (e.g., searching for "user authentication" will find
login_handler).
Key Features & Advantages
- Persistent Local Memory: Data is safely saved to your local disk (
metadata.jsonandindex.bin). It never expires and survives across system restarts. - Intelligent Text Chunking: Automatically breaks down large documents into overlapping semantic chunks (1000 chars) before embedding, ensuring context is never lost.
- Flawless MCP Stdio Communication: Strictly intercepts and suppresses rogue C-level progress bars (like
tqdmfromsentence-transformers) that normally corrupt JSON-RPC streams, ensuring a stable connection. - 100% Local Privacy: Runs entirely on your machine using the
all-MiniLM-L6-v2embedding model. No data is sent to external cloud APIs for indexing.
Architecture
- Protocol: FastMCP (running over
stdio). - Embedding Model:
sentence-transformers(all-MiniLM-L6-v2) generating 384-dimensional vectors. - Vector Database:
turbovec(TurboQuantIndex) for ultra-fast, locally persisted similarity search. - Storage Layer: Local JSON mapping for metadata, allowing automated fallback and index rebuilding if the
.binfile is lost. - Modular Codebase: The project is cleanly separated into
main.py(entry point),vector_db.py(database logic), andtools.py(MCP tool definitions).
How It Works: AI & MCP Interaction Flow
The Turbovec MCP Server acts as an invisible bridge between your AI client and a persistent local memory database. Here is the step-by-step logic of how they interact:
- User Prompt: The user asks a question or assigns a task in their AI Client (e.g., Zoo Code, Claude Desktop, Cursor).
- LLM Tool Call: The LLM evaluates the prompt and determines it needs past context or codebase knowledge, triggering an MCP tool (like
search_knowledgeoradd_knowledge). - MCP Execution: The AI Client forwards this tool request to the Turbovec MCP Server running locally via the standardized JSON-RPC protocol over
stdio. - Vector Database: The MCP server interacts with the
turboveclocal vector database to embed the query, search for semantic matches, or store new text chunks. - Context Return & Generation: The retrieved data is returned to the AI Client and passed back to the LLM. The LLM seamlessly incorporates this retrieved memory into its final context-aware response to the user.
<img width="3600" height="1771" alt="627eriqdfvafasf" src="https://github.com/user-attachments/assets/dd4e4422-7d2f-4f5a-bb6b-a79a00fb9020" />
Installation
You can run Turbovec MCP Server locally via Python or using Docker.
Option A: Local Python Setup
-
Clone the repository:
git clone https://github.com/henny-bee/Turbovec-MCP-Server.git cd turbovec-mcp-server -
Create a Virtual Environment (Recommended):
python -m venv venv # Windows .\venv\Scripts\activate # Mac/Linux source venv/bin/activate -
Install Dependencies:
pip install -r requirements.txt -
Verify it Runs:
# Windows .\venv\Scripts\python.exe main.py # Mac/Linux ./venv/bin/python main.pyYou should see a success message:
Turbovec MCP Server is successfully running
Option B: Docker Setup
-
Clone the repository:
git clone https://github.com/henny-bee/Turbovec-MCP-Server.git cd turbovec-mcp-server -
Run with Docker Compose:
docker-compose up -d
Alternatively, you can build and run it directly using the provided Dockerfile.
Testing
The project uses pytest for testing to ensure the database and tools work correctly.
-
Install development dependencies:
pip install -r requirements-dev.txt -
Run the tests:
pytest tests/
Zoo Code / AI Editor Integration
To use this server in your AI coding assistant (like Zoo Code, Cursor, or Claude Desktop), add it to your MCP configuration settings (usually found in Settings > MCP Servers, or mcp_settings.json).
Configuration
Add the following block to your mcpServers configuration:
{
"mcpServers": {
"turbovec-mcp": {
"command": "python",
"args": ["C:/absolute/path/to/turbovec-mcp-server/main.py"],
"env": {
"PYTHONUNBUFFERED": "1"
}
}
}
}
Troubleshooting Tip: If you encounter a
ModuleNotFoundError(e.g., missingnumpyorturbovec), it means the editor is using the system Python instead of the virtual environment. To fix this, change"command": "python"to the absolute path of your virtual environment's Python executable (e.g.,"C:/path/to/turbovec-mcp-server/venv/Scripts/python.exe"on Windows, or"/path/to/venv/bin/python"on Mac/Linux).
Custom Instructions (Recommended)
To ensure your AI assistant seamlessly and proactively uses the memory server without asking for permission, we highly recommend adding the following to your AI's Custom Instructions or System Prompt:
You are connected to a long-term memory system via the Turbovec MCP.
You must proactively use `search_knowledge` and `add_knowledge` automatically to save and retrieve important project context, architectural decisions, and code snippets.
Do not ask for permission to save or search memory; execute these operations seamlessly in the background to ensure context is preserved across our sessions.
Available MCP Tools
Once connected, the AI will have access to the following tools:
add_knowledge(title, content): Embeds and saves raw text into memory.add_file_knowledge(file_path): Reads a local file, chunks it, and saves it into memory.search_knowledge(query, top_k): Performs a semantic search to retrieve context from the database.delete_knowledge(title_or_id): Removes a specific piece of knowledge from the database.optimize_memory(): Performs hard-deletion and garbage collection of the vector database to optimize memory usage.clear_memory(): Completely wipes the local database and vector index.
Sponsored by
<a href="https://www.iseekaigo.com/"> <img src="logo-iseekaigo-line.png" alt="ISEEKAIGO" height="50"> </a>
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。