Utility MCP Server

Utility MCP Server

An MCP server that provides AI-powered document processing and search capabilities, including PDF summarization, text extraction, metadata retrieval, and web search via Google Custom Search.

Category
访问服务器

README

MCPs

Think of MCP like the 𝐔𝐒𝐁-𝐂 𝐩𝐨𝐫𝐭 𝐨𝐧 𝐚 𝐥𝐚𝐩𝐭𝐨𝐩. Any device that wants to connect, whether it’s an external hard drive, monitor, or power supply, must follow the USB-C standard. Similarly, MCP is a standardized protocol developed by Anthropic that defines how 𝐋𝐋𝐌𝐬 𝐜𝐨𝐧𝐧𝐞𝐜𝐭 𝐭𝐨 𝐜𝐨𝐧𝐭𝐞𝐱𝐭𝐬 𝐚𝐧𝐝 𝐝𝐚𝐭𝐚 𝐬𝐨𝐮𝐫𝐜𝐞𝐬.

The laptop = MCP host The USB-C port = MCP itself The monitor, hard drive, power supply = MCP servers

For more understanding of MCPs, visit https://medium.com/@BH_Chinmay/basics-of-mcps-why-and-what-9579c21caac4

Utility MCP Server

A Model Context Protocol (MCP) server that provides AI-powered document processing and search capabilities.

Overview

This project implements an MCP server with tools for:

  • PDF Document Processing: Extract text, metadata, and statistics from PDF files
  • Document Summarization: Use Azure OpenAI to generate intelligent summaries of documents and pages
  • Web Search: Perform Google searches and retrieve top results

Architecture

MCPs/
├── mcp-servers/              # MCP server entrypoint and tools
│   ├── server.py             # Server initialization and tool registration
│   ├── mcp_instance.py       # FastMCP instance with logging configuration
│   └── tools/                # Tool implementations
│       ├── document_summerizer.py  # PDF summarization tools
│       ├── google_search.py        # Web search tool
│       └── pdf_reader.py           # Basic PDF reading tool
├── utils/                    # Shared utility modules
│   ├── llm_utils.py          # Azure OpenAI integration
│   └── pdf_utils.py          # PDF processing utilities
├── requirements.txt          # Python dependencies
├── .env                      # Environment variables (secrets)
└── README.md                 # This file

Tools

Document Summarization Tools (tools/document_summerizer.py)

summarize_document(pdf_path)

Analyzes a complete PDF document and generates an intelligent summary.

  • Extracts all text from the PDF
  • Splits content into manageable chunks
  • Summarizes each chunk using Azure OpenAI
  • Generates a final summary from partial summaries
  • Returns: File path, page count, chunk count, and final summary

summarize_page(pdf_path, page_number)

Generates a summary for a specific page in a PDF.

  • Extracts text from the specified page
  • Processes through Azure OpenAI
  • Returns: Page number and page summary

extract_document_text(pdf_path)

Extracts all text content from a PDF without summarization.

  • Returns: Full text content

document_statistics(pdf_path)

Calculates text statistics for a document.

  • Returns: Character count, word count, and line count

document_metadata(pdf_path)

Retrieves metadata from a PDF document.

  • Returns: Page count, title, author, creator, and producer

Google Search Tool (tools/google_search.py)

google_search(query, num_results)

Performs a Google search and returns top results.

  • Uses Google Custom Search API
  • Parameters:
    • query: Search query string
    • num_results: Number of results to return (default: 5)
  • Returns: List of results with title, link, and snippet

PDF Reader Tool (tools/pdf_reader.py)

read_pdf(pdf_path)

Basic PDF text extraction tool.

  • Reads and returns all text from a PDF
  • Returns: Extracted text content

Utilities

LLM Utilities (utils/llm_utils.py)

Handles Azure OpenAI integration:

  • _get_client(): Initializes Azure OpenAI client with environment configuration
  • llm_summary(text): Sends text to Azure OpenAI for summarization

PDF Utilities (utils/pdf_utils.py)

Core PDF processing functions:

  • extract_text(pdf_path): Extracts all text from a PDF
  • chunk_text(text, chunk_size): Splits text into chunks
  • get_metadata(pdf_path): Extracts PDF metadata
  • get_page_text(pdf_path, page_number): Extracts text from a specific page

Configuration

Environment Variables (.env)

The application requires the following environment variables:

Google Search Configuration:

GOOGLE_API_KEY=<your-google-api-key>
GOOGLE_SEARCH_ENGINE_ID=<your-search-engine-id>

Azure OpenAI Configuration:

AZURE_OPENAI_ENDPOINT=<your-azure-endpoint>
AZURE_OPENAI_API_KEY=<your-azure-api-key>
AZURE_OPENAI_API_VERSION=<api-version>
AZURE_OPENAI_DEPLOYMENT=<deployment-name>

Important: Never commit .env with actual credentials to version control.

Setup and Installation

Prerequisites

  • Python 3.10+
  • Virtual environment (recommended)

Installation

  1. Create a virtual environment:
python -m venv .venv
  1. Activate the virtual environment:
# Windows
.venv\Scripts\activate

# macOS/Linux
source .venv/bin/activate
  1. Install dependencies:
pip install -r requirements.txt
  1. Create and configure .env:
cp .env.example .env
# Edit .env and add your API credentials

Running the Server

Development Mode

cd mcp-servers
mcp dev server.py

Production Mode

cd mcp-servers
python server.py

Logging

The application uses Python's built-in logging module with the following configuration:

  • Level: INFO (use DEBUG for detailed output)
  • Format: YYYY-MM-DD HH:MM:SS - logger_name - LEVEL - message
  • Security: All credentials and API keys are masked in logs

Log Levels

  • DEBUG: Detailed operation information (page extraction, chunk processing)
  • INFO: Normal operation events (tool invocations, completion status)
  • WARNING: Warning messages (invalid page numbers, missing configuration)
  • ERROR: Error events with full exception tracebacks

Enabling Debug Logging

To see more detailed logs during development:

import logging
logging.getLogger().setLevel(logging.DEBUG)

Dependencies

  • mcp (1.28.1+): Model Context Protocol framework
  • openai (1.0.0+): Azure OpenAI client
  • python-dotenv: Environment variable management
  • httpx (0.27.0+): Async HTTP client for Google Search API
  • PyMuPDF (1.24.0+): PDF text extraction

See requirements.txt for complete list.

Error Handling

All tools include comprehensive error handling:

  • File existence validation before processing
  • Exception logging with full tracebacks
  • User-friendly error messages in responses
  • No sensitive data logged in error messages

Security Considerations

  1. Credentials: Store all API keys and endpoints in .env file
  2. Logging: Credentials are never logged or printed
  3. Environment: Use separate .env files for different environments (dev, staging, production)
  4. Access: Restrict access to .env file permissions (never commit to version control)

Development Notes

Adding New Tools

  1. Create a new file in tools/ directory
  2. Import and register with @mcp.tool() decorator
  3. Add comprehensive logging with logger.info() and logger.error()
  4. Document the tool in this README

Code Style

  • Use descriptive variable names
  • Include docstrings for all functions
  • Log important operations and errors
  • Never log sensitive information (API keys, authentication tokens)

Troubleshooting

Import Errors

If you encounter ModuleNotFoundError:

  1. Ensure virtual environment is activated
  2. Run pip install -r requirements.txt
  3. Check that Python path includes both mcp-servers/ and project root directories

Missing Environment Variables

If you see "Missing Azure OpenAI environment variables":

  1. Verify .env file exists in the project root
  2. Check that all required variables are set (not empty)
  3. Restart the server after updating .env

PDF Processing Issues

  • Ensure PDF file exists and is readable
  • Check that PyMuPDF (fitz) is properly installed
  • Verify sufficient disk space for large PDF files

Support

For issues or questions, please refer to the logging output for detailed error information.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选