pdf-automation

pdf-automation

A local MCP server that drives PDFium and pypdf to perform comprehensive PDF operations including inspection, assembly, page editing, watermarking, rendering, extraction, form filling, encryption, compression, attachments, bookmarks, and metadata management.

Category
访问服务器

README

pdf-engine-mcp

한국어 안내 → README.ko.md

A local MCP server that drives two native PDF enginesPDFium (Chromium's PDF renderer, via pypdfium2) and pypdf (+cryptography) — to cover the full everyday PDF workflow: inspect, merge/split, page surgery, watermark, render pages to images, extract text/images, fill AcroForm fields, encrypt/decrypt, compress, embed/extract attachments, rewrite bookmarks and metadata.

Design philosophy: same as its siblings (word / excel / ppt / hwp) — expose engine calls, not hand-rolled file poking. Unlike the Office siblings there is no desktop app to automate, so this server is stateless: no COM session, no worker thread, fully cross-platform. Every tool call opens the file, works, writes to out_path, and closes. Originals are never modified.

Requirements

  • Python 3.10+ — verified on 3.12 (Windows 11; no OS-specific dependency)
  • Claude Code or any MCP client
  • No Adobe Acrobat, no MS Office, no Ghostscript needed

Install

git clone https://github.com/Feynman520/d01-p05-pdf-engine-mcp.git
cd d01-p05-pdf-engine-mcp
py -3.12 -m venv .venv          # or: python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -r requirements.txt

Register with Claude Code

Run this in the cloned folder (uses absolute paths, so it works from anywhere afterwards):

claude mcp add pdf-automation --scope user -- "$PWD\.venv\Scripts\python.exe" "$PWD\server.py"

--scope user makes it available in every project. Use --scope project to limit it to one project.

Verify

$py = ".\.venv\Scripts\python.exe"; $env:PYTHONUTF8 = "1"
& $py tests\smoke_engine.py   # runs all 17 tool paths on self-generated fixture PDFs
& $py tests\server_tools.py   # MCP tool registration

Tools (17 core + 1 diagnostic)

Group Tool Input → Output Engine
Inspect pdf_info path, password? → pages, page sizes, metadata, encryption, form fields, attachments, bookmark tree pypdf
Assemble pdf_merge inputs[{path,pages?,password?}], out_path → merged file pypdf
pdf_split src_path, pages?/out_path or out_dir, every → extracted / chunked files pypdf
pdf_pages op: rotate|delete|reorder (+pages/degrees/order) pypdf
pdf_watermark text (built-in Helvetica, Latin) or stamp_path (any PDF), mode: overlay|background pypdf
Render pdf_render_images src_path, out_dir, pages?, dpi, png|jpg → page images PDFium
pdf_extract_text src_path, pages? → per-page text (honest empty result for scans) PDFium
pdf_extract_images src_path, out_dir, pages? → embedded image originals pypdf
Forms pdf_form_fields path → AcroForm field names/types/values pypdf
pdf_fill_form fields{name:value}, flatten? → filled (optionally locked) form pypdf
Security pdf_encrypt user_password, owner_password?, AES-256, allow_printing?, allow_copying? pypdf
pdf_decrypt password → unencrypted copy (for files whose password you know) pypdf
Optimize pdf_compress stream compression + duplicate removal (lossless), image_quality? (lossy) pypdf
Attach pdf_attach_files / pdf_extract_attachments embed files into / extract from the PDF pypdf
Structure pdf_bookmarks nested [{title,page,children?}] → rewritten outline pypdf
pdf_set_metadata title/author/subject/keywords/creator/producer pypdf
pdf_health → engine versions (stateless, instant) both

Page specs are 1-based strings: "3", "1-3,5", "4-", "-2". Paths should be absolute. PDF→Word conversion is intentionally not here — the word sibling's word_convert owns it.

Architecture notes

  • Stateless by design: PDF has no resident desktop app, so there is no session to manage — each call is open → work → write out_path → close. Blocking work is delegated to a thread (anyio.to_thread) to keep the event loop responsive.
  • Two engines, one rule: PDFium does what pure Python cannot (rasterize, layout-aware text); pypdf does document surgery. PyMuPDF was deliberately avoided (AGPL vs this repo's MIT).
  • Text watermark uses the built-in Helvetica font (Latin-1 only) drawn by a tiny built-in raw-PDF generator (engine/rawpdf.py) — for CJK watermarks pass a stamp PDF via stamp_path.
  • Honest extraction: scanned PDFs return an empty text result with a note pointing to pdf_render_images + OCR, never hallucinated text.
  • Originals preserved: results are always written to a new out_path/out_dir.

Limitations

  • Text watermark supports Latin scripts only (use stamp_path for CJK).
  • pdf_decrypt requires the correct password — this is a convenience tool, not a cracker.
  • XFA forms (legacy Adobe LiveCycle) are not supported; AcroForm only.

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选