omapdf-mcp
Gives AI agents tools to read PDFs, list form fields, apply operations, highlight, add notes, fill fields, place signatures, list signatures, and flatten PDFs.
README
omapdf
Preview.app for Linux — but your agent can drive it.

omapdf is an agent-native PDF tool: a fast GTK4 editor for reading, annotating, and signing; a scriptable CLI; an MCP server for AI agents; and an Omarchy integration that ties them all to the operating system. One op engine underneath — anything a human can do by hand, an agent can do by instruction, and vice versa.
omapdf edit lease.pdf # the editor
omapdf sign lease.pdf --page 4 --at 120,540 --date -o signed.pdf
# …or just tell your agent: "fill out this lease, highlight anything
# unusual, and get it ready for my signature"
Why
The "someone emailed me a PDF, I need to highlight two things, sign it, and send it back" workflow is macOS Preview's killer feature — and Linux has no lightweight equivalent. And nobody anywhere treats AI agents as first-class PDF users. omapdf does both, with one architecture:
One operations API, every client is thin
Every action — highlight, comment, fill a field, stamp a signature, draw ink — is a small JSON operation. A document edit is a list of them. The editor, the CLI, and the MCP server all funnel through the same engine:
human agent
│ │
┌─────┴─────┐ ┌────────┴────────┐
│ editor │ │ MCP server / │
│ (GTK4) │ │ Claude skill │
└─────┬─────┘ └────────┬────────┘
│ ops (JSON) │
└──────────────┬─────────────────┘
┌───────┴───────┐
│ op engine │ validate → resolve → apply → report
│ (PyMuPDF) │
└───────┬───────┘
document.pdf
Every applied op echoes back its resolved geometry, so either side can show the other exactly what changed and where. Everything written is a standard PDF annotation — Acrobat, Preview, and Evince users see your notes and highlights as normal comments.
The editor (omapdf edit)
A native GTK4 editor, Preview-fast, designed for Omarchy but plain-GTK portable:
- Tools (hand-drawn vector icon set): select/drag, pen with a tap-again color palette, highlighter, text, sticky notes, signature placement, green-check and red-cross stamps
- Ghost model: everything you place is a draggable, nudgeable pending item until Save bakes it through the op engine — and undo crosses the save boundary: Ctrl+Z after saving reverts the file and resurrects the saved items as editable ghosts
- Reading comforts: thumbnail sidebar (F9), fit-width zoom that tracks the live viewport, zoom presets + Ctrl+scroll, full-document search (Ctrl+F) with match cycling, page navigation by Up/Down, PgUp/PgDn, Home/End, Ctrl+G go-to-page
- Comments open on click: click any saved annotation to read its text — including comments left by agents or by other people's PDF apps
- Agent proposals as ghosts:
omapdf edit doc.pdf --ops proposal.jsonloads an agent's dry-run ops as selected, draggable overlays — nudge, then Save. Agent proposes, human confirms. - Ask your agent (✦): type a question, it opens your OS default agent
(
omarchy agent prompt) with the document attached. The editor watches the file and reloads itself when the agent saves changes. - Share (native, no fake share sheet): email attach, LocalSend, copy the file (or a zip of it) straight onto the clipboard, show in folder — with an optional flatten-copy-first toggle
- Save celebration included. You'll see.
The CLI
omapdf read doc.pdf # structured JSON: text+bboxes, fields, annots
omapdf read doc.pdf --text-only # just the words
omapdf fields form.pdf # fillable fields with names and rects
omapdf snapshot doc.pdf --page 2 --grid 50 # page PNG with a labeled
# coordinate grid — agents read placement
# coordinates straight off the image
omapdf annotate doc.pdf --page 2 --match "termination clause"
omapdf note doc.pdf --page 2 --at 400,300 --text "negotiate this"
omapdf fill form.pdf --field tenant_name "Peter Bergin" --field rent "1800"
omapdf apply doc.pdf --ops edits.json # atomic batch of ops
omapdf sig draw # draw your signature once (GTK window)
omapdf sig add ~/sig.png --name work # …or import an image
omapdf sign doc.pdf --page 4 --at 120,540 --date -o signed.pdf
omapdf flatten doc.pdf -o final.pdf # bake everything in for any viewer
omapdf open doc.pdf # opens the omapdf editor
Every edit command takes -o (default: in place), --dry-run, and
--json (each op returns its resolved geometry). Coordinates are PDF
points, origin top-left, 1-based pages — identical to what read reports.
The ops vocabulary is specified in docs/ops.md.
Agents
claude mcp add omapdf -- omapdf-mcp
MCP tools: read_pdf, list_form_fields, apply_ops, highlight,
add_note, fill_field, place_signature (dry-run by default —
confirm-before-ink), list_signatures, flatten_pdf.
The Claude Code skill in skill/ is a full PDF-assistant
playbook: recipes for review-and-highlight, form filling, signing,
redlining, checklists, extraction with page citations — built around a
precision ladder: text anchors → grid snapshot (look at the page) →
dry-run → verify the written result visually → hand a draggable ghost to
the human when taste matters.
Install
git clone https://github.com/pbergin11/omapdf && cd omapdf
python -m venv --system-site-packages .venv # system gi for the GTK editor
.venv/bin/pip install -e '.[mcp]'
ln -s "$PWD/.venv/bin/omapdf" ~/.local/bin/omapdf
ln -s "$PWD/.venv/bin/omapdf-mcp" ~/.local/bin/omapdf-mcp
ln -s "$PWD/bin/omapdf-pick" ~/.local/bin/omapdf-pick
Requires Python ≥ 3.11, PyMuPDF, and (for the editor and sig draw)
PyGObject + GTK4 from your distro. Arch packaging in
packaging/PKGBUILD.
Make omapdf your system PDF handler:
cp share/omapdf.desktop ~/.local/share/applications/
xdg-mime default omapdf.desktop application/pdf
Omarchy integration
- Top-bar widget (
shell-plugin/): a PDF pill — click for a recent-PDFs menu, pick one, it opens in the editor (ln -s .../shell-plugin/omapdf.bar ~/.config/omarchy/plugins/omapdf.bar, thenomarchy bar put omapdf.bar --section right) - Ask-agent uses
omarchy agent prompt, so it launches whatever default agent the user picked (omarchy default agent), in Omarchy's native agent window; Voxtype dictation works in the ask box like any text field - Packaged in the spirit of omasnap: a focused, single-purpose native tool
Status & roadmap
Working today: op engine, CLI, MCP server + skill, signature store and
drawing window, the full editor, bar widget, agent ask/watch loop, tests.
See docs/roadmap.md for what's next — headlines:
signature-line auto-detection (sign --auto), annotation deletion/editing
of saved items, comments summary page, cryptographic (PAdES) signing via
pyHanko.
License
AGPL-3.0-or-later (matching our PyMuPDF dependency). See LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。