Grasp MCP Server
A Windows desktop control MCP server that gives Claude eyes and hands—screen capture with coordinate-grid overlays, pixel-perfect DPI-correct mouse/keyboard input via a Rust backend, and direct PowerShell command execution, enabling full local desktop automation through natural language.
README
<div align="center">
🖐️ Grasp
Give Claude eyes and hands — full desktop control over MCP.
Grasp is a Model Context Protocol server that lets Claude see your screen and drive your mouse, keyboard and shell. Install it once and Claude Code — or any MCP host — can operate your Windows PC the way a person would: look, point, click, type, run commands.
</div>
What "eyes and hands" means. Eyes — Grasp captures the screen, resizes it to a vision-friendly resolution and draws a coordinate grid on top, so Claude can read exact positions instead of guessing. Hands — Grasp injects real, DPI-correct mouse and keyboard input at the OS level (a Rust backend), and runs shell commands directly. Together they close the loop: Claude looks, decides, acts, looks again.
Why Grasp
Most "computer use" setups are either a cloud VM you don't control, or a browser-only automation that can't touch the rest of your machine. Grasp runs locally, drives the real desktop, and plugs into the tools you already use through a single mcp add command.
- Pixel-perfect on scaled displays. The Rust input backend respects Windows display scaling (125% / 150% / 200%) — clicks land exactly where Claude looked, not two centimetres off.
- Vision tuned for accuracy. Screenshots are cropped/resized to ~1280 px with a 100-px coordinate grid overlay, the single biggest lever on click precision.
- Shell without the screenshot tax.
run_commandreturns real text output — no opening a terminal, typing into it, and screenshotting the result. - Safe by construction. Input is released on startup/shutdown so a stuck modifier can never freeze your keyboard; failing and elevated commands are reported honestly, never as silent successes.
Tools
| Tool | What it does | |
|---|---|---|
| 👁️ | screenshot |
Capture the screen (full or active window) as an image with a coordinate grid. |
| 👁️ | get_active_window |
Title, process and bounds of the focused window. |
| 👁️ | list_windows |
All open top-level windows. |
| 👁️ | get_screen_info |
Monitor geometry, resolution and DPI scale. |
| ✋ | click |
Glide the cursor to a point and click (left / right / middle). |
| ✋ | double_click |
Double-click at a point. |
| ✋ | move_mouse |
Move the cursor (to reveal hover-only menus). |
| ✋ | drag |
Press, glide with the button held, release — selections, sliders, drag-and-drop. |
| ✋ | type_text |
Type literal text at the current focus. |
| ✋ | press_key |
Keys and combos: enter, ctrl+c, alt+f4, win+r, ctrl+shift+esc… |
| ✋ | scroll |
Scroll up/down at a point. |
| ⚙️ | run_command |
Run a shell command (PowerShell), optionally elevated via UAC; returns text output. |
| ⚙️ | wait |
Pause N ms to let an app settle before the next screenshot. |
How the coordinate model works
- Claude calls
screenshot. Grasp captures the screen, resizes the longest edge to 1280 px, and overlays a grid labelled every 100 px. - Claude reads a coordinate straight off the grid (e.g. "the button is near
x≈540, y≈300"). - Claude calls
clickwith those numbers. Grasp maps them back through the resize + crop + DPI transform to the exact physical pixel and clicks there.
You never do the math — pass coordinates in the same space you see them. (Advanced: pass coord_space: "screen" to use raw physical pixels instead.)
Requirements
- Windows 10/11, x64. The mouse/keyboard backend ships as a prebuilt native binary for
win32-x64. Screen capture, window enumeration and image processing are cross-platform, but Grasp is Windows-first today. - Node.js ≥ 18.
- An MCP host — Claude Code, Claude Desktop, or anything that speaks MCP over stdio.
Install
git clone https://github.com/Jamshed7470/grasp-mcp.git
cd grasp-mcp
npm install
npm run doctor # verify Grasp can see and control this machine
npm run doctor should end with ✅ Grasp is ready — Claude has eyes and hands.
Add to Claude Code
From the repo directory:
claude mcp add grasp -- node "%CD%\src\index.js"
…or with an absolute path from anywhere:
claude mcp add grasp -- node "C:\path\to\grasp-mcp\src\index.js"
Then just ask Claude, e.g. "take a screenshot and open Settings for me."
Add to Claude Desktop
Edit claude_desktop_config.json
(%APPDATA%\Claude\claude_desktop_config.json) and add:
{
"mcpServers": {
"grasp": {
"command": "node",
"args": ["C:\\path\\to\\grasp-mcp\\src\\index.js"]
}
}
}
Restart Claude Desktop. Grasp's tools appear under the 🔨 menu.
Usage examples
Once installed, drive it in plain language — Claude picks the tools:
- "Open Notepad, type today's date, and save it to the Desktop."
- "Find the Wi-Fi icon in the tray and tell me which network I'm on." (screenshot → read)
- "Update all my winget packages." (
run_commandwithelevated: true) - "Scroll the page down and click the first search result."
Configuration
| Environment variable | Purpose |
|---|---|
GRASP_ENIGO_PATH |
Absolute path to the node-enigo .node binary, if it isn't in the bundled native/ folder. |
Per-tool options (image size, grid on/off, JPEG quality, command timeout, working directory, elevation) are passed as tool arguments — see each tool's description in the MCP schema.
⚠️ Security
Grasp gives an AI model real control of your computer — the same reach you have. Treat it accordingly.
- You are always in the loop. MCP hosts ask before each tool call by default. Keep it that way for anything destructive.
- Close sensitive windows (password managers, private messages) before running an agent — whatever is on screen is what Claude sees, and screenshots may briefly hold that content.
- Elevated commands need your hand.
elevated: truetriggers a Windows UAC prompt you must approve; UAC runs on a secure desktop that cannot be automated. If you dismiss it, Grasp reports the command as not run — never a silent success. - Never commit screenshots. This repo's
.gitignoreblocks*.png/*.jpgfor exactly this reason. - Run Grasp only on machines and tasks you're comfortable handing to an assistant.
Troubleshooting
| Symptom | Fix |
|---|---|
enigo NOT loaded — hands disabled |
The native binary wasn't found. Confirm native/node-enigo-win32-x64.node exists, or set GRASP_ENIGO_PATH. You're likely not on Windows x64. |
node-screenshots NOT loaded |
Run npm install (native module needs its prebuild). |
| Clicks land in the wrong place | Take a fresh screenshot right before clicking — coordinates are relative to the most recent capture. |
run_command can't find cmd/ping/net |
Grasp already restores PATH/PATHEXT for hosts that strip the environment; update to the latest version if you see this. |
Run npm run doctor any time to re-check every backend.
How it's built
src/
index.js MCP server — registers the 13 tools over stdio
native.js screen capture + mouse/keyboard (node-enigo, node-screenshots, get-windows)
image-processor.js sharp pipeline: crop → resize → coordinate-grid overlay
shell.js PowerShell runner with UTF-8 output + UAC elevation
doctor.js standalone backend health check
native/
node-enigo-win32-x64.node prebuilt Rust input backend
test/
mcp-test.js end-to-end test over real MCP stdio (read-only + safe tools)
live-notepad.js opt-in live test that types into a throwaway file and reads it back
Backends: node-enigo (Rust, mouse/keyboard) · node-screenshots (Rust, capture) · get-windows · sharp · @modelcontextprotocol/sdk.
Development
npm run doctor # health check, no server
node test/mcp-test.js # 11 end-to-end assertions over MCP stdio
node test/live-notepad.js # opt-in: drives the real desktop, cleans up after itself
Contributions welcome — especially a macOS/Linux input backend to make Grasp truly cross-platform.
License
MIT. Grasp bundles a prebuilt node-enigo binary and depends on node-screenshots, both built on the Rust enigo / xcap ecosystem.
<div align="center"> <sub>Grasp — <i>to grasp</i> is both to see and to hold. Built for Claude.</sub> </div>
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。