docx_mcp_server_ts
Enables programmatic processing of DOCX/OOXML files through MCP, covering text, tables, images, headers/footers, structured data tags, comments, and track changes.
README
DOCX MCP Server
A comprehensive TypeScript-based MCP (Model Context Protocol) server for universal DOCX processing with full OOXML support. Process Word documents programmatically with support for text, tables, images, headers/footers, SDTs, comments, and more.
Features
- Complete OOXML Access: Read/write DOCX parts at ZIP level with full namespace support
- Text Operations: Extract, find, and replace text with minimal diff preservation
- Table Management: Insert/delete rows, modify cells, merge/split operations
- Image Handling: Add inline/positioned images with EMU-based sizing
- Structured Data Tags (SDT): Access content controls by tag or alias
- Headers/Footers: List and modify section headers and footers
- Track Changes: Accept/reject revisions, handle insertions/deletions
- Comments: Manage document comments
- Metadata: Read/write core and app properties
- LRU Caching: Efficient memory management with part caching
- Lossless XML: Preserves document structure with fast-xml-parser
Installation
npm install
npm run build
Quick Start
Start the Server
npm start
The server will listen on stdin/stdout for MCP protocol messages.
Installation & Configuration
Claude Code CLI
claude mcp install docx \
--command node \
--args /full/path/to/docx_mcp_server_ts/dist/index.js \
--env LOG_LEVEL=INFO
~/.claude.json (для Claude Code)
Отредактировать ~/.claude.json добавить в раздел "projects":
{
"projects": {
"/full/path/to/docx_mcp_server_ts": {
"mcpServers": {
"docx": {
"command": "node",
"args": ["/full/path/to/docx_mcp_server_ts/dist/index.js"],
"env": {
"LOG_LEVEL": "INFO"
}
}
}
}
}
}
Пример для Linux/WSL:
{
"projects": {
"/mnt/c/Users/pavelk/Desktop/Projects/MCP-servers/docx_mcp_server_ts": {
"mcpServers": {
"docx": {
"command": "node",
"args": ["/mnt/c/Users/pavelk/Desktop/Projects/MCP-servers/docx_mcp_server_ts/dist/index.js"],
"env": {
"LOG_LEVEL": "INFO"
}
}
}
}
}
}
MCP Tools
Document Management
docx.open
Open a DOCX document from file or base64 buffer.
Input:
{
"path": "/path/to/document.docx",
"bufferBase64": "..." // OR provide base64 data
}
Output:
{
"docId": "uuid-string",
"parts": ["word/document.xml", ...],
"props": { "core": {}, "app": {} }
}
docx.close
Close a document and release resources.
Input: { "docId": "uuid" }
docx.save
Save document to file or return as base64.
Input:
{
"docId": "uuid",
"path": "/output/path.docx", // optional
"returnBase64": true // optional
}
docx.list_parts
List all parts in document.
docx.part_read / docx.part_write
Read/write individual XML parts for low-level access.
Text Operations
docx.get_text
Extract all text from document.
Input: { "docId": "uuid", "scope": "document|headers|footers|all" }
docx.replace_text
Replace text preserving run structure.
Input:
{
"docId": "uuid",
"match": "search text",
"replace": "replacement",
"mode": "literal|regex",
"where": "document|headers|footers|all"
}
Output: { "replaced": 5 }
docx.find
Find text with context.
Output:
{
"hits": [
{
"text": "found text",
"context": "...found text...",
"offset": 150
}
]
}
Table Operations
docx.tables_list
List all tables with dimensions.
Output:
{
"tables": [
{
"tableXPath": "//w:tbl[1]",
"rows": 5,
"colsApprox": 3
}
]
}
docx.table_edit
Perform table operations.
Input:
{
"docId": "uuid",
"tableXPath": "//w:tbl[1]",
"op": {
"kind": "setCellText",
"row": 0,
"col": 0,
"text": "new value"
}
}
Supported operations:
{ "kind": "setCellText", "row": number, "col": number, "text": string }{ "kind": "insertRow", "at": number }{ "kind": "deleteRow", "at": number }{ "kind": "insertCol", "at": number }{ "kind": "deleteCol", "at": number }
Structured Data Tags (SDT)
docx.sdt_get
Get content control content.
Input: { "docId": "uuid", "tagOrAlias": "control_tag" }
Output:
{
"xml": "<w:p>...</w:p>",
"textPreview": "Control content..."
}
docx.sdt_put
Update content control.
Input:
{
"docId": "uuid",
"tagOrAlias": "control_tag",
"xmlFragment": "<w:p>...</w:p>"
}
Image Operations
docx.images_list
List all images with metadata.
Output:
{
"images": [
{
"rId": "rId4",
"path": "word/media/image1.png",
"sizeEMU": { "cx": 914400, "cy": 914400 }
}
]
}
docx.image_add
Insert image inline or anchored.
Input:
{
"docId": "uuid",
"target": {
"afterParagraphXPath": "//w:p[1]",
"sdtTagOrAlias": "imageControl" // OR use SDT
},
"image": {
"path": "/local/image.png",
"base64": "...", // OR base64 data
"filename": "image.png",
"contentType": "image/png"
},
"placement": {
"kind": "inline" // OR { "kind": "anchor", "xEMU": 0, "yEMU": 0 }
},
"size": {
"widthMM": 50,
"heightMM": 50
},
"altText": "Description"
}
docx.image_update_position
Update anchored image position/size.
Advanced Operations
docx.styles_get / docx.styles_set
Read/write styles.xml
docx.numbering_get / docx.numbering_set
Read/write numbering.xml
docx.headers_footers_list
List headers and footers with section info.
docx.headers_footers_get / docx.headers_footers_set
Read/write specific header or footer.
docx.comments_list / docx.comments_add / docx.comments_delete
Manage document comments.
docx.changes_accept_all
Accept all tracked changes (remove w:del, unwrap w:ins).
Output: { "removedDel": 3, "flattenedIns": 5 }
docx.metadata_get / docx.metadata_set
Read/write document properties (core.xml, app.xml).
Size Conversions
The server handles EMU (English Metric Unit) conversions internally:
- 1 inch = 914,400 EMU
- 1 mm ≈ 36,000 EMU
- 1 point ≈ 12,700 EMU
Examples
Extract and Replace Text
// Open document
const openResult = await client.call('docx.open', {
path: '/tmp/document.docx'
});
const docId = openResult.docId;
// Get text
const textResult = await client.call('docx.get_text', { docId });
console.log(textResult.text);
// Replace text
await client.call('docx.replace_text', {
docId,
match: 'old text',
replace: 'new text',
mode: 'literal'
});
// Save
await client.call('docx.save', {
docId,
path: '/tmp/document-modified.docx'
});
// Close
await client.call('docx.close', { docId });
Modify Table
// List tables
const tablesResult = await client.call('docx.tables_list', { docId });
const tableXPath = tablesResult.tables[0].tableXPath;
// Update cell
await client.call('docx.table_edit', {
docId,
tableXPath,
op: {
kind: 'setCellText',
row: 0,
col: 0,
text: 'Updated Value'
}
});
// Insert row
await client.call('docx.table_edit', {
docId,
tableXPath,
op: {
kind: 'insertRow',
at: 1
}
});
Add Image
const fs = require('fs').promises;
const imageBuffer = await fs.readFile('/path/to/image.png');
const base64 = imageBuffer.toString('base64');
await client.call('docx.image_add', {
docId,
target: {
afterParagraphXPath: '//w:p[1]'
},
image: {
base64,
filename: 'image.png',
contentType: 'image/png'
},
placement: {
kind: 'inline'
},
size: {
widthMM: 100,
heightMM: 75
},
altText: 'My image'
});
Architecture
src/
├── index.ts # MCP server entry point
├── logger.ts # Logging utility
├── errors.ts # Error types and codes
├── ooxml/
│ ├── namespaces.ts # OOXML constants and namespaces
│ ├── emu.ts # Unit conversion utilities
│ ├── dom.ts # XML DOM utilities (xmldom + fontoxpath)
│ ├── xmlParser.ts # FXP parser with order preservation
│ ├── parts.ts # ZIP part reading/writing
│ ├── rels.ts # Relationship management
│ ├── text.ts # Text operations with diff-match-patch
│ ├── tables.ts # Table manipulation
│ ├── sdt.ts # Structured Data Tags
│ ├── drawings.ts # Image handling
│ ├── headersFooters.ts # Header/footer operations
│ ├── comments.ts # Comment management
│ ├── changes.ts # Track changes handling
│ ├── styles.ts # Styles XML access
│ └── numbering.ts # Numbering XML access
├── store/
│ ├── types.ts # Store type definitions
│ └── docStore.ts # Document store with LRU cache
└── mcp/
└── tools.ts # MCP tool implementations
Performance
- Memory: LRU cache limits per-document parts to 50 cached items
- Total Size: Supports documents up to 100MB in-memory
- Partial Access: Only requested parts are parsed from ZIP
- Minimal Diffs: Text replacements preserve run structure when possible
Limitations
- Page layout calculations are not performed (Word's rendering engine needed)
- Advanced DrawingML transformations are read-only
- VBA macros and embedded OLE objects not supported
- Extremely large documents (>500MB) may require streaming
Development
# Install dependencies
npm install
# Type check
npm run type-check
# Build
npm run build
# Run dev server
npm run dev
# Debug with inspector
npm run dev:debug
Logging
Control log level via environment variable:
LOG_LEVEL=DEBUG npm start # Verbose
LOG_LEVEL=INFO npm start # Default
LOG_LEVEL=WARN npm start # Warnings only
LOG_LEVEL=ERROR npm start # Errors only
Protocol Support
- Transport: stdio
- Protocol: MCP (Model Context Protocol)
- Handler: @modelcontextprotocol/sdk
License
MIT
Resources
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。