docsray-mcp
An MCP server that provides AI assistants with advanced document perception capabilities including text extraction, structure analysis, and deep content understanding through multiple tools and providers.
README
🔍 Docsray MCP Server
Docsray is a powerful Model Context Protocol (MCP) server that gives AI assistants like Claude advanced document perception capabilities. Extract text, navigate pages, analyze structure, and understand any document with ease.
✅ Status: Published to PyPI and TestPyPI - Working in Cursor, Claude Desktop, and other MCP clients
✨ Features
🎯 Five Powerful Tools
docsray_peek- Quick document overview with format detection and provider capabilitiesdocsray_map- Generate comprehensive document structure maps with cachingdocsray_xray- AI-powered deep analysis extracting entities, relationships, and insightsdocsray_extract- Extract content in multiple formats (markdown, text, JSON, tables)docsray_seek- Navigate to specific pages, sections, or search for content
🔌 Multi-Provider Architecture
-
PyMuPDF4LLM - Lightning-fast PDF processing (✅ Implemented)
- Fast markdown extraction
- Basic table detection
- Multi-page support
- Always enabled as fallback
-
LlamaParse - Deep document understanding with LLMs (✅ Implemented)
- AI-powered entity extraction
- Custom analysis instructions
- Comprehensive caching in .docsray directories
- Rich format preservation (markdown, images, tables)
-
PyTesseract - OCR for scanned documents (🔄 Planned)
-
Mistral OCR - AI-powered OCR and analysis (🔄 Planned)
🚀 Key Benefits
- Universal Input Support - Local files (./path, ../path, /absolute) and URLs (https://)
- Intelligent Provider Selection - Automatically chooses the best tool for each task
- Smart Caching - LlamaParse results cached in .docsray directories for instant access
- Dynamic Discovery - Tools report actual capabilities based on what's enabled
- Production Ready - Comprehensive error handling, logging, and 56 tests
- Self-Documenting - Built-in resources for discovery by MCP clients
📦 Installation
Quick Start with uvx (Recommended)
# Run directly without installation
uvx docsray-mcp start
# Or install globally
uv tool install docsray-mcp
# Then run with:
docsray start
# or
docsray-mcp start
Alternative: Install with pip
# Basic installation (PyMuPDF4LLM only)
pip install docsray-mcp
# With LlamaParse for AI analysis
pip install "docsray-mcp[ai]"
# Development installation
pip install -e ".[dev]"
🚀 Quick Start
1. Set up API Keys (Optional but Recommended)
Create a .env file in your project:
# For AI-powered analysis with LlamaParse
LLAMAPARSE_API_KEY=llx-your-key-here
# Or use environment variables
export LLAMAPARSE_API_KEY=llx-your-key-here
Get your free LlamaParse API key at cloud.llamaindex.ai
2. Configure with Your MCP Client
For Cursor
Add to your Cursor settings:
{
"mcpServers": {
"docsray": {
"command": "uvx",
"args": ["docsray-mcp"],
"env": {
"LLAMAPARSE_API_KEY": "llx-your-key-here"
}
}
}
}
For Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"docsray": {
"command": "uvx",
"args": ["docsray-mcp"],
"env": {
"LLAMAPARSE_API_KEY": "llx-your-key-here"
}
}
}
}
📚 Usage Examples
Basic Document Overview
Peek at ./document.pdf to see its structure and available formats
Extract Entities from Contracts
Xray ./contract.pdf and extract all parties, dates, payment terms, and obligations
Navigate Documents
Map the complete structure of ./manual.pdf including all sections and subsections
Extract Specific Content
Extract pages 10-20 from ./report.pdf as markdown
Analyze Web Documents
Analyze https://arxiv.org/pdf/2301.00234.pdf for methodology and key findings
Compare Providers
Extract text from document.pdf with provider pymupdf4llm (fast)
Xray document.pdf with provider llama-parse (AI analysis)
🛠️ Advanced Configuration
Environment Variables
# Provider Configuration
DOCSRAY_PYMUPDF4LLM_ENABLED=true # Always true by default
DOCSRAY_LLAMAPARSE_ENABLED=true
LLAMAPARSE_API_KEY=llx-your-key
# Performance Tuning
DOCSRAY_CACHE_ENABLED=true
DOCSRAY_CACHE_TTL=3600
DOCSRAY_MAX_CONCURRENT_REQUESTS=5
DOCSRAY_TIMEOUT_SECONDS=30
# Logging
DOCSRAY_LOG_LEVEL=INFO
Provider Capabilities
PyMuPDF4LLM (Always Available)
- ✅ Fast text extraction
- ✅ Markdown formatting
- ✅ Basic table detection
- ✅ Multi-page support
- ❌ No AI analysis
- ❌ No OCR
LlamaParse (When API Key Configured)
- ✅ AI-powered analysis
- ✅ Entity extraction
- ✅ Custom instructions
- ✅ Table extraction
- ✅ Image extraction
- ✅ Layout preservation
- ✅ Relationship mapping
- ✅ Result caching
🧪 Testing
# Run all tests
pytest tests/
# Run only unit tests (no API calls)
pytest tests/unit/
# Run integration tests
pytest tests/integration/
# Run with coverage
pytest tests/ --cov=src/docsray --cov-report=html
Current test coverage: 52 tests passing with comprehensive coverage across all components
📖 API Reference
Tool: docsray_peek
Get quick document overview and metadata.
{
"document_url": "path/to/document.pdf",
"depth": "structure", # metadata | structure | preview
"provider": "auto" # auto | pymupdf4llm | llama-parse
}
Tool: docsray_map
Generate comprehensive document structure map.
{
"document_url": "path/to/document.pdf",
"include_content": false,
"analysis_depth": "deep", # basic | deep | comprehensive
"provider": "auto"
}
Tool: docsray_xray
Deep AI-powered document analysis.
{
"document_url": "path/to/document.pdf",
"analysis_type": ["entities", "key-points"],
"custom_instructions": "Extract all dates and amounts",
"provider": "llama-parse"
}
Tool: docsray_extract
Extract content in various formats.
{
"document_url": "path/to/document.pdf",
"extraction_targets": ["text", "tables"],
"output_format": "markdown", # markdown | text | json
"pages": [1, 2, 3], # Optional: specific pages
"provider": "auto"
}
Tool: docsray_seek
Navigate to specific document locations.
{
"document_url": "path/to/document.pdf",
"target": {"page": 5}, # or {"section": "Introduction"} or {"query": "search text"}
"extract_content": true,
"provider": "auto"
}
🏗️ Architecture
docsray-mcp/
├── src/docsray/
│ ├── server.py # FastMCP server with discovery resources
│ ├── providers/ # Provider implementations
│ │ ├── base.py # Provider interface
│ │ ├── pymupdf4llm.py # Fast PDF extraction
│ │ └── llamaparse.py # AI-powered analysis
│ ├── tools/ # MCP tool implementations
│ │ ├── peek.py # Document overview
│ │ ├── map.py # Structure mapping
│ │ ├── xray.py # Deep analysis
│ │ ├── extract.py # Content extraction
│ │ └── seek.py # Navigation
│ └── utils/ # Utilities
│ ├── cache.py # Document caching
│ └── llamaparse_cache.py # LlamaParse .docsray cache
├── tests/
│ ├── unit/ # Fast isolated tests
│ ├── integration/ # Component interaction tests
│ └── manual/ # Debugging scripts
└── PROMPTS.md # Example prompts for all use cases
🤝 Contributing
We welcome contributions! See CONTRIBUTING.md for guidelines.
Development Setup
# Clone the repository
git clone https://github.com/docsray/docsray-mcp.git
cd docsray-mcp
# Install in development mode
pip install -e ".[dev]"
# Run tests
pytest tests/
# Run linting
ruff check src/
📄 License
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
🙏 Acknowledgments
- Built on FastMCP framework
- Document processing powered by PyMuPDF4LLM
- AI analysis powered by LlamaParse
- Inspired by the Model Context Protocol specification
📬 Support
Made with ❤️ for the MCP ecosystem
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。