linked-docs

linked-docs

Enables AI assistants to intelligently search and reference documentation using hybrid semantic + keyword search via MCP protocol.

Category
访问服务器

README

Linked Documentation System

A production-ready RAG system with dual interfaces: REST API + MCP Protocol
Enables web applications and AI assistants to intelligently search and reference documentation using hybrid semantic + keyword search.

🎯 What Is This?

This is a complete RAG (Retrieval Augmented Generation) system that provides two ways to access powerful documentation search:

  1. REST API Server (main.py) - FastAPI-based HTTP server for web applications and integrations
  2. MCP Server (mcp_server.py) - Model Context Protocol server for AI assistants (Cursor, Claude Desktop)

Both servers share the same hybrid search engine, enabling accurate documentation retrieval whether you're building a web app or empowering an AI assistant.

Key Features

  • Intelligent Hybrid Search: Combines semantic understanding (FAISS embeddings) with keyword matching (BM25)
  • Smart Ranking: Title/metadata boosting, multi-chunk document expansion, relevance scoring
  • Multi-Format Support: PDF, Markdown, and web documentation (via built-in scraper)
  • MCP Native: Works seamlessly in Cursor, Claude Desktop, and other MCP-compatible tools
  • Enterprise Ready: Access control, audit logging, local-first architecture
  • Fast: <200ms search latency, optimized chunking and indexing
  • Zero Cost: Runs 100% locally, no API keys or cloud dependencies

🏗️ How It Works

┌─────────────────────────────────────────────────────────┐
│              AI Assistant (Cursor/Claude)               │
└────────────────────┬────────────────────────────────────┘
                     │ MCP Protocol (JSON-RPC over stdio)
                     ▼
┌─────────────────────────────────────────────────────────┐
│                   mcp_server.py                         │
│  ┌─────────────────────────────────────────────────┐   │
│  │  Tools: search_documentation(), list_sources()  │   │
│  └─────────────────────────────────────────────────┘   │
└────────────────────┬────────────────────────────────────┘
                     │
    ┌────────────────┼────────────────┐
    ▼                ▼                ▼
┌─────────┐   ┌──────────┐   ┌──────────────┐
│ Hybrid  │   │ Access   │   │ Audit Logger │
│ Search  │   │ Control  │   │              │
└────┬────┘   └──────────┘   └──────────────┘
     │
     ├─ Semantic Search (FAISS + embeddings)
     │   • Title/metadata boosting
     │   • Multi-chunk document expansion
     │
     └─ Keyword Search (BM25)
         • Exact term matching
         • Traditional ranking

Project Structure

LinkedDocsMCP/
├── mcp_server.py          # Main MCP server (stdio interface)
├── main.py                # FastAPI server (for testing/debugging)
├── download_docs.py       # Web documentation scraper CLI
├── connectors/            # Document format handlers
│   ├── pdf.py            # PDF extraction
│   └── markdown.py       # Markdown parsing
├── indexing/             # Search engine core
│   ├── chunker.py        # Semantic text chunking
│   ├── embedder.py       # Sentence transformers
│   ├── vector_store.py   # FAISS vector database
│   ├── keyword_search.py # BM25 implementation
│   └── hybrid_search.py  # Combined search with boosting
├── schemas/              # Data models
│   ├── config.py         # Settings & configuration
│   └── document.py       # Document schemas
├── core/                 # Cross-cutting concerns
│   ├── access_control.py # Permission system
│   └── audit.py          # Query logging
└── data/                 # Local storage
    ├── sources/          # Your documents (PDF, MD)
    └── vector_store/     # Indexed vectors & metadata

🚀 Quick Start (5 Minutes)

Prerequisites

  • Python 3.10+
  • Cursor or Claude Desktop (for MCP integration)

1. Install Dependencies

pip install -r requirements.txt

First run: Downloads ~80MB embedding model (one-time)

2. Add Documentation

Option A: Download from web

# Download Factorio wiki (example)
python download_docs.py https://wiki.factorio.com/Tutorials --crawl --max 20

Option B: Add your own files

# Copy PDFs or Markdown files
copy your-docs.pdf data/sources/
copy your-guide.md data/sources/

3. Set Up MCP in Cursor (or other LLM service)

Add to your Cursor MCP config (~/.cursor/mcp.json or C:\Users\<USER>\.cursor\mcp.json):

{
  "mcpServers": {
    "linked-docs": {
      "command": "python",
      "args": ["C:/full/path/to/LinkedDocsMCP/mcp_server.py"]
    }
  }
}

4. Restart Cursor & Use!

In Cursor's chat:

What are the different enemy types in Factorio?

The AI will automatically search your documentation and provide accurate, cited answers! ✨

Key Features Explained

Hybrid Search

Combines two complementary approaches:

Semantic Search (70% weight)

  • Uses sentence-transformers (all-MiniLM-L6-v2 model)
  • Understands meaning: "authentication setup" matches "configuring auth"
  • Converts text to 384-dimensional vectors
  • Fast similarity search with FAISS

Keyword Search (30% weight)

  • Uses BM25 algorithm (same as Elasticsearch)
  • Exact term matching: great for technical terms, code, etc.
  • Traditional ranking with document length normalization

Smart Ranking Enhancements:

  • Title Boosting: Documents whose titles match the query get 3x boost
  • Multi-Chunk Expansion: Returns up to 3 sequential chunks from highly relevant documents
  • Document Grouping: Results grouped by source document for better context

Semantic Chunking

Unlike naive character-splitting, this uses smart boundaries:

  1. Markdown headers (##, ###) - keeps sections together
  2. Paragraph breaks (\n\n) - maintains topical coherence
  3. Sentences - fallback for unstructured text

Settings:

  • Chunk size: 2048 characters (whole sections, not fragments)
  • Overlap: 200 characters (prevents context loss at boundaries)

Web Documentation Scraper

Built-in tool to download and convert web docs:

# Download single page
python download_docs.py https://wiki.example.com/Guide

# Crawl multiple pages (with smart duplicate detection)
python download_docs.py https://wiki.example.com/Main --crawl --max 50

# Force re-download (skip existing detection)
python download_docs.py https://wiki.example.com/Main --crawl --force

# Filter by language
python download_docs.py https://wiki.example.com/Main --crawl --languages en,de

Features:

  • Auto-detects and skips existing pages (saves time & bandwidth)
  • Respects same-domain and link patterns
  • Polite crawling with configurable delays
  • Converts HTML to clean Markdown with metadata
  • Preserves document structure (headers, lists, tables)

🔒 Security & Access Control

Built-in features:

  • 4-tier access hierarchy: PUBLIC → INTERNAL → RESTRICTED → CONFIDENTIAL
  • Query-time filtering based on user permissions
  • Full audit logging (every query tracked)
  • Local-only processing (no data leaves your machine)

Audit logs (data/audit.log):

{
  "timestamp": "2025-10-21T14:30:00Z",
  "event_type": "search",
  "user_id": "mcp_client",
  "query": "enemy types",
  "results_count": 5,
  "search_time_ms": 143
}

⚙️ Configuration

Edit schemas/config.py or set environment variables:

# Search weights
SEMANTIC_WEIGHT = 0.7  # Meaning-based search
KEYWORD_WEIGHT = 0.3   # Exact term matching

# Chunking
CHUNK_SIZE = 1280      # Larger chunks for complete sections
CHUNK_OVERLAP = 128    # Overlap for context continuity

# Model
EMBEDDING_MODEL = "all-MiniLM-L6-v2"  # Fast, accurate, small

🔧 Advanced Usage

REST API (for testing/debugging)

# Start FastAPI server
python main.py

# Search via REST
curl -X POST http://localhost:8000/api/v1/search_docs \
  -H "Content-Type: application/json" \
  -d '{"query": "getting started", "top_k": 5}'

# Interactive API docs
open http://localhost:8000/docs

Programmatic Usage

from indexing.embedder import Embedder
from indexing.vector_store import VectorStore
from indexing.hybrid_search import HybridSearchEngine

# Initialize
embedder = Embedder(model_name="all-MiniLM-L6-v2")
vector_store = VectorStore(embedding_dim=384)
search_engine = HybridSearchEngine(vector_store, keyword_searcher, embedder)

# Search
results = search_engine.search("how to configure auth", top_k=5)
for chunk, score in results:
    print(f"{score:.3f}: {chunk.text[:100]}...")

Technical Highlights

  1. Hybrid search outperforms pure semantic or keyword alone
  2. Smart chunking preserves document structure
  3. Title boosting dramatically improves ranking quality
  4. Multi-chunk expansion provides complete context
  5. Zero cloud dependencies - privacy-first architecture
  6. MCP native - works with any compatible AI assistant

🙏 Acknowledgments

Built with:

  • FastAPI - Modern Python web framework
  • Sentence Transformers - Semantic embeddings
  • FAISS - Vector similarity search
  • BM25 (rank-bm25) - Keyword ranking
  • MCP - Model Context Protocol by Anthropic

Status: ✅ Demo Ready | Version: 1.0.0 | Updated: October 2025

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选