DocMistral MCP Server

DocMistral MCP Server

Converts documents and images to Markdown using Mistral AI's OCR, enabling AI-powered document processing via MCP-compatible clients like Claude Desktop.

Category
访问服务器

README

DocMistral MCP Server

Model Context Protocol Python License CI Contributions npm version

A powerful MCP (Model Context Protocol) server that converts documents and images to Markdown using Mistral AI's advanced OCR and document processing capabilities. Perfect for integrating document processing into Claude Desktop and other MCP-compatible clients.

🚀 Features

MCP Server Capabilities

  • 🔗 MCP Compatible: Works with Claude Desktop, Continue, and other MCP clients
  • 📦 One-Command Install: npx @trsdn/mistraldocai-mcp-server
  • 🔄 Automatic Setup: Manages Python environment and dependencies
  • 🌍 Cross-Platform: Windows, macOS, and Linux support

Document Processing

  • 📄 Documents: PDF, PPTX, DOCX via Mistral's OCR API
  • 🖼️ Images: PNG, JPG, JPEG, GIF, BMP, AVIF support
  • 🧠 AI-Powered: Advanced document understanding with complex layouts
  • ✍️ OCR Support: Scanned documents and handwritten text
  • ⚡ Fast Processing: Up to 2,000 pages per minute
  • 💰 Cost-Effective: $0.001 per page ($1 per 1,000 pages)

🚀 Quick Start

Step 1: Install the MCP Server

# Install and test with one command
npx @trsdn/mistraldocai-mcp-server --test

Step 2: Get API Key

Get your Mistral API key from console.mistral.ai

Step 3: Configure Your MCP Client

For Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "mistraldocai": {
      "command": "npx",
      "args": ["@trsdn/mistraldocai-mcp-server"],
      "env": {
        "MISTRAL_API_KEY": "your_mistral_api_key_here"
      }
    }
  }
}

For Other MCP Clients

Use the command: npx @trsdn/mistraldocai-mcp-server with environment variable MISTRAL_API_KEY

Step 4: Start Using!

The server provides 2 tools:

  • process_document - Convert documents/images to Markdown
  • get_supported_formats - List supported file formats

Manual Installation (Python Tool)

For direct Python usage:

  1. Clone this repository:
git clone <repository-url>
cd DocMistral
  1. Create a virtual environment:
python3 -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
  1. Install dependencies:
pip install -r requirements.txt

Configuration

API Key Setup

  1. Get a Mistral API key from console.mistral.ai

  2. Create a .env file in the project directory:

cp .env.example .env
  1. Edit .env and add your API key:
MISTRAL_API_KEY=your_api_key_here

Alternatively, you can set it as an environment variable:

export MISTRAL_API_KEY=your_api_key_here

Usage

MCP Server Usage

The MCP server provides two tools for document processing:

1. Process Single Document

Convert a document or image file to Markdown:

{
  "name": "process_document",
  "arguments": {
    "file_path": "/path/to/document.pdf"
  }
}

Or with base64 content (useful for MCP clients):

{
  "name": "process_document",
  "arguments": {
    "base64_content": "base64_encoded_file_content",
    "file_name": "document.pdf"
  }
}

2. Get Supported Formats

Get information about supported file formats:

{
  "name": "get_supported_formats",
  "arguments": {}
}

Python Tool Usage

For direct command-line usage:

# Process all files in the input directory
python docmistral.py

# Convert a single file
python docmistral.py --file document.pdf

Custom Directories

Specify custom input and output directories:

python docmistral.py --input /path/to/docs --output /path/to/markdown

Command Line Options

  • --input, -i: Input directory (default: input)
  • --output, -o: Output directory (default: output)
  • --mistral-api-key, -k: Mistral AI API key (required)
  • --file, -f: Convert a single file instead of a directory

Directory Structure

DocMistral/
├── docmistral.py       # Main script
├── requirements.txt    # Python dependencies
├── .env.example        # Environment variables template
├── README.md          # This file
├── input/             # Default input directory
│   └── .gitkeep      # Ensures directory is tracked
└── output/            # Default output directory
    └── .gitkeep      # Ensures directory is tracked

Requirements

  • Python 3.8+
  • See requirements.txt for Python package dependencies

Supported Formats

  • Documents: PDF, PPTX, DOCX (via OCR API)
  • Images: PNG, JPG, JPEG, GIF, BMP, AVIF (via OCR API)
  • File size limit: 50 MB
  • Page limit: 1,000 pages per document

How it Works

  • Uses Mistral's dedicated OCR API (client.ocr.process) for all supported formats
  • Advanced document understanding handles complex layouts, tables, and equations
  • Processes up to 2000 pages per minute
  • Pricing: $0.001 per page ($1 per 1,000 pages)

🔧 MCP Tools Reference

process_document

Converts documents and images to Markdown format.

Parameters:

  • file_path (string): Path to the document/image file
  • OR base64_content (string) + file_name (string): Base64 content with filename
  • mime_type (string, optional): MIME type of the file

Example Usage:

{
  "name": "process_document",
  "arguments": {
    "file_path": "/path/to/document.pdf"
  }
}

With Base64 Content:

{
  "name": "process_document",
  "arguments": {
    "base64_content": "base64_encoded_file_content",
    "file_name": "document.pdf"
  }
}

get_supported_formats

Lists all supported file formats and their limitations.

Parameters: None

Example Usage:

{
  "name": "get_supported_formats",
  "arguments": {}
}

📋 Supported Formats

Format Extensions Processing Method Notes
Documents .pdf, .pptx, .docx Mistral OCR API Up to 1,000 pages
Images .png, .jpg, .jpeg, .gif, .bmp, .avif Mistral OCR API Up to 50 MB

Limitations:

  • Maximum file size: 50 MB
  • Maximum pages: 1,000 per document
  • Processing speed: Up to 2,000 pages/minute
  • Cost: $0.001 per page

🎯 Use Cases

  • Research: Convert academic papers and reports to Markdown
  • Documentation: Process technical manuals and guides
  • Data Extraction: Extract text from scanned documents
  • Content Migration: Convert legacy documents to modern formats
  • OCR Processing: Digitize handwritten notes and forms

🔌 MCP Compatibility

This server is fully compatible with the Model Context Protocol (MCP) specification and works with:

  • Claude Desktop - Anthropic's desktop application
  • Continue - VS Code extension
  • Zed - Code editor with MCP support
  • Custom MCP clients - Any application implementing the MCP protocol

MCP Registry

This server is available in the MCP ecosystem:

  • Package: @trsdn/mistraldocai-mcp-server
  • Command: npx @trsdn/mistraldocai-mcp-server
  • Protocol Version: MCP 1.0
  • Transport: stdio

🏷️ Tags & Discovery

Find this MCP server using these tags:

  • mcp-server - MCP compatible server
  • mistral - Uses Mistral AI
  • ocr - Optical Character Recognition
  • document-processing - Document conversion
  • pdf-to-markdown - PDF conversion
  • image-to-text - Image text extraction
  • ai-powered - AI-enhanced processing

📦 Installation Methods

NPX (Recommended)

npx @trsdn/mistraldocai-mcp-server

Global Installation

npm install -g @trsdn/mistraldocai-mcp-server
mistraldocai-mcp

Local Development

git clone https://github.com/yourusername/MistralDocAI-mcp.git
cd MistralDocAI-mcp
npm install && npm run build
npm start

🛠️ Development

Building from Source

# Clone the repository
git clone <repository-url>
cd MistralDocAI-mcp

# Install npm dependencies
npm install

# Build TypeScript
npm run build

# Test the build
npm test

Publishing

npm run build
npm publish

Notes

  • The tool preserves the directory structure when converting files
  • All documents are processed through Mistral AI for consistency
  • Output files are saved with the .md extension
  • Supports fallback processing for edge cases
  • API key is required for all operations
  • The MCP server automatically manages Python virtual environments
  • Cross-platform support (Windows, macOS, Linux)

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选