DeckProbe MCP Server
Enables agents to inspect PDF, Microsoft Office, and Apple iWork documents without rendering them, providing page/slide counts, metadata, security signals, structure, and integrity as deterministic JSON via typed tools and batch operations.
README
<div align="center">
DeckProbe MCP Server
Let an agent ask what's inside a PDF, Office, or iWork file — without opening it.
Install · Tools · Configuration · Security · How it works · DeckProbe
</div>
An MCP server that exposes
DeckProbe — ffprobe for documents — as
four typed tools. Ask for page counts, slide counts, metadata, encryption and
macro signals, structure, or integrity, and get back bounded, deterministic JSON
with confidence, evidence, and measured I/O cost.
Nothing is rendered, no macro runs, no external reference is followed, and no network connection is opened. It is safe to point at untrusted files.
// probe { "path": "deck.pptx", "targets": ["slide_count"], "view": "values" }
{
"schema_version": 2,
"status": "ok",
"driver": { "id": "powerpoint", "profile": "pptx" },
"values": { "powerpoint.slide_count": 31 },
"view": "values"
}
Install
Nothing to install ahead of time — npx fetches the server and the engine
together.
Claude Code
claude mcp add deckprobe -- npx -y @deckflow/deckprobe-mcp
Claude Desktop, Cursor, VS Code, Zed, and anything else reading mcpServers
{
"mcpServers": {
"deckprobe": {
"command": "npx",
"args": ["-y", "@deckflow/deckprobe-mcp"]
}
}
}
For a pinned install, npm install -g @deckflow/deckprobe-mcp and use
deckprobe-mcp as the command.
Requires Node.js 20 or newer. The engine binary arrives as a per-platform
optional dependency for macOS, Linux (glibc and musl), and Windows on x86-64 and
ARM64; anywhere else the server falls back to the same engine compiled to
WebAssembly, so npx works wherever Node does.
Tools
| Tool | Use it for |
|---|---|
probe |
Everything about one document |
probe_batch |
Inventory or triage many documents in one call |
list_formats |
Which formats are supported, and where support stops |
list_targets |
The exact target names a format offers |
There is also one resource, deckprobe://schema, carrying the report JSON
Schema bundled with the running engine.
probe
{
"path": "reports/q3.pptx",
"targets": ["@summary", "@security"], // presets, short names, or canonical names
"level": "metadata", // header | metadata | deep
"min_confidence": "high", // low | medium | high | exact
"target_confidence": { "slide_count": "exact" },
"view": "report", // report | values
"budget": { "max_physical_bytes": 8388608, "timeout_ms": 1000 }
}
targets accepts short names (slide_count), canonical names
(powerpoint.slide_count), and presets:
| Preset | Expands to |
|---|---|
@header |
Container identity only — format, size, extension match, encryption flag |
@summary |
Identity, common metadata, and primary structure |
@security |
Encryption, macros, signatures, external references, active content |
@structure |
Format-owned counts, names, and dimensions |
@assets |
Images, media, previews, fonts, embedded objects |
@quality |
Integrity, repair, extension match, conformance |
@format |
Every format-specific target at the active level |
@all |
Everything available at the active level |
@summary deliberately omits statistics that need a full-file read. A PDF's
page_count is the notable case — ask for it explicitly.
probe_batch
{ "paths": ["a.pdf", "b.pptx", "c.xlsx"], "targets": ["@security"] }
One engine process handles the whole batch. Results come back in input order,
each with its own report or its own error, so one bad file never spoils the run.
Defaults to the compact values view. Literal paths only — expand globs
yourself.
list_formats and list_targets
list_targets takes a format (pdf, docx, xlsx, pptx, doc, xls,
ppt, key, numbers, pages) and returns each target's aliases,
description, value type, minimum level, cost class, and selector membership.
Pass detail: "full" for the engine's complete report, including per-target
JSON Schema fragments and expanded selector lists.
Both are cached for the lifetime of the server process.
Reading a report
The tool result is the engine's own schema-v2 envelope, unmodified. Two things are worth knowing before consuming it:
status: "partial"is not a failure. It means at least one requested target could not be resolved at the requested confidence. It is named inexecution.unresolved_targets, and every other result still stands.confidence_scoreis a fixed constant per label (0.4,0.7,0.95,1.0), not a calibrated probability.0.95does not mean the value is right 95% of the time.
Only results with status resolved or estimated carry a value. unknown is
common and usually means the document simply does not record that fact.
A failing call returns isError with the engine's error envelope plus one line
saying what to do about it. Failures the server itself raises before the engine
runs — a missing path, a directory, a path outside the allow-list, an exceeded
deadline — use the same envelope shape with an MCP_-prefixed code and
origin: "mcp-server".
Configuration
Every setting is an environment variable, set in your client's MCP config. All are optional.
| Variable | Default | Meaning |
|---|---|---|
DECKPROBE_MCP_BIN |
– | Engine binary to use instead of the bundled one |
DECKPROBE_MCP_ROOTS |
unrestricted | Allowed directories, separated like PATH |
DECKPROBE_MCP_TIMEOUT_MS |
30000 |
Hard per-call deadline on an engine process |
DECKPROBE_MCP_MAX_CONCURRENCY |
4 |
Concurrent engine processes |
DECKPROBE_MCP_MAX_BATCH |
64 |
Paths accepted by one probe_batch call |
{
"deckprobe": {
"command": "npx",
"args": ["-y", "@deckflow/deckprobe-mcp"],
"env": { "DECKPROBE_MCP_ROOTS": "/Users/me/Documents:/Users/me/Downloads" }
}
}
Security
DeckProbe is built for untrusted input: bounded parsing, no renderer, no macro interpreter, no external-reference resolution, and no network access. This server adds two things on top.
- Process isolation and a hard deadline. Each probe runs in its own
short-lived process, killed if it outruns
DECKPROBE_MCP_TIMEOUT_MS. - An optional read allow-list.
DECKPROBE_MCP_ROOTSpins the reachable tree; paths are symlink-resolved before the check, so a link cannot step around it. The default is unrestricted, matching the CLI the user could run themselves — set it for shared or automated deployments.
Reports describe a document (metadata, counts, signals) rather than reproducing its contents. Note that report values such as a document title are still attacker-controlled strings: the server passes them through as JSON data and never interpolates them into instructions, and a consumer should treat them the same way.
Report a vulnerability privately as described in SECURITY.md.
How it works
MCP client
│ JSON-RPC over stdio
▼
deckprobe-mcp ── validates arguments, resolves the path, maps the result
│ argv + stdout (one process per probe, or one --jsonl process per batch)
▼
DeckProbe engine ── plans the cheapest paths that answer the request
The server spawns the native DeckProbe CLI rather than calling the WebAssembly build. The CLI reads only the byte ranges a probe plan needs, where the WebAssembly path holds the whole file in memory, and a separate OS process both isolates untrusted parsing and can be killed outright. The engine is chosen in this order:
DECKPROBE_MCP_BIN- the binary that ships with this package's
@deckflow/deckprobedependency deckprobeonPATH- the bundled WebAssembly engine
The resolved engine is logged to stderr at startup. stdout belongs to the MCP transport and carries nothing else.
MCP server or agent skill?
DeckProbe also ships an Agent Skill that teaches a shell-capable agent to use the CLI directly. Both teach the same vocabulary. Use the skill when the agent has a shell and you want the CLI's full surface; use this server when it does not, or when you want typed arguments validated before the engine ever runs.
Development
npm install
npm test # typecheck, lint, build, and the full suite
npm run test:watch
Contributions are welcome — see CONTRIBUTING.md. The design rationale, including the alternatives that were rejected, is in docs/rfc.md.
License
MIT. See LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。