jetkvm-mcp

jetkvm-mcp

MCP server that enables AI to see and control a physical computer via a JetKVM device for screen viewing, mouse/keyboard input, media mounting, and power management.

Category
访问服务器

README

jetkvm-mcp

Give an AI eyes and hands on a physical computer.

jetkvm-mcp is an MCP server that turns a JetKVM — a small open-source KVM-over-IP device — into a machine that Claude (or any MCP client) can see and operate directly: watch the screen, type, click, mount boot media, and control power. Because the JetKVM sits on the HDMI and USB ports, the AI drives the computer below the OS — BIOS screens, bootloaders, installers, headless boxes with no network, machines that are wedged. No agent, no SSH, nothing installed on the target.

you:    "Screenshot the machine. It's stuck — what's wrong?"
claude: → screenshot → "It's sitting at a GRUB rescue prompt. The root partition
         UUID changed. Want me to boot it manually?" → type_text → enter → fixed

Works against stock JetKVM firmware — no modifications to the device.

The two planes

Plane Tools Nature
Screen control (eyes + hands) screenshot, click, double_click, move_mouse, type_text, press_key, scroll vision loop — the AI looks, then acts
Device control mount_media_url, mount_media_storage, upload_media, upload_and_mount, unmount_media, list_storage, delete_storage_file, storage_space, virtual_media_state, power, power_state, dc_power, wake_host, wol, usb_emulation, video_state, reboot_device deterministic RPC

Full parameter reference: docs/tools.md.

How it works

One WebRTC peer connection to the device drives everything:

Claude ──MCP/stdio──▶ server.py (this repo, runs on your workstation)
                        │
                        └──WebRTC over LAN──▶ JetKVM ──HDMI-in / USB-HID-out──▶ target machine
                             ├─ H.264 video track ─▶ decoded locally (PyAV) ─▶ JPEG screenshots
                             └─ "rpc" data channel ─▶ keyboard / mouse / media / power JSON-RPC
  • The device already streams its HDMI capture as an H.264 video track — the client decodes it locally and hands the AI JPEG snapshots on demand. (JetKVM has no snapshot endpoint; it doesn't need one.)
  • A reliable rpc data channel carries every JSON-RPC method the device's own web UI uses: keyboardReport, absMouseReport (absolute 0–32767, drift-free), mountWithHTTP, setATXPowerAction, and friends.
  • The server connects lazily on the first tool call and keeps the one connection alive.

Deep dive — handshake, codec negotiation, the keyframe/PLI story, coordinate mapping: docs/architecture.md.

Requirements

  • A JetKVM attached to the target machine, reachable on your network
  • Python 3.11+ on the machine that runs Claude
  • aiortc/av wheels bundle FFmpeg on macOS/Linux; if a build from source is triggered, install FFmpeg dev libraries first (brew install ffmpeg / apt install libavdevice-dev)

Quickstart

git clone https://github.com/shvartzj1/jetkvm-mcp.git
cd jetkvm-mcp
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env        # set JETKVM_URL (+ JETKVM_PASSWORD if your device has one)

Prove the pipeline before wiring it into anything — this connects, holds the stream open, and saves four screenshots:

set -a; source .env; set +a
python smoke_test.py

Expected output — sustained ~60 fps, snapshots in single-digit milliseconds after the first:

connected. video_state: {'ready': True, 'width': 1280, 'height': 1024, 'fps': 60}
  snapshot 0: 1280x1024   179578 bytes  (grab 3129 ms)  frames_seen=1
  snapshot 1: 1280x1024   178280 bytes  (grab   11 ms)  frames_seen=128
  ...

Wire it into Claude

Claude Code (one command, available in every session):

claude mcp add jetkvm --scope user \
  --env JETKVM_URL=http://192.168.1.50 \
  --env JETKVM_VERIFY_TLS=false \
  -- /abs/path/jetkvm-mcp/.venv/bin/python /abs/path/jetkvm-mcp/server.py

Claude Desktop (claude_desktop_config.json):

{
  "mcpServers": {
    "jetkvm": {
      "command": "/abs/path/jetkvm-mcp/.venv/bin/python",
      "args": ["/abs/path/jetkvm-mcp/server.py"],
      "env": {
        "JETKVM_URL": "http://192.168.1.50",
        "JETKVM_PASSWORD": "",
        "JETKVM_VERIFY_TLS": "false"
      }
    }
  }
}

Then just talk to it: "Screenshot the machine, open a terminal, and check disk usage." The AI calls screenshot → reasons → click / type_text → repeats.

The killer workflow: hands-free bare-metal provisioning

Device control and screen control compose into something no in-OS agent can do — installing an operating system on an empty machine:

mount_media_url("https://mirror.lan/rocky-9.iso", "CDROM")   # host the ISO yourself
power("reset")                                               # reboot into the installer
# screenshot → click → type_text … the AI walks through the installer by sight
unmount_media()

The device's own storage partition is tiny, so mount_media_url (the device streams the image over HTTP with range requests) is the right path for full-size ISOs; upload_and_mount is for small recovery images.

Gotchas (read this before filing a bug)

  • First screenshot takes ~3 s; the rest are instant. The device only emits an H.264 keyframe when asked via RTCP PLI. Browsers request keyframes automatically; aiortc does not — so this client sends PLI on connect and whenever frames go stale (_request_keyframe in client.py). Without that, decode fails on every packet forever (avcodec_send_packet: Invalid data). If you're building your own client: this is the trap.
  • Keyboard layout: type_text maps ASCII → USB HID usage codes assuming the US layout on the target OS. On other layouts, shifted symbols swap (on a UK target, " arrives as @). Letters, digits, and / - . ; = are layout-stable; prefer them in critical commands.
  • Coordinates: click/move_mouse take pixel coordinates on the most recent screenshot; the client maps them to the HID absolute range using the live frame dimensions, so there is no drift.
  • getVideoState may report streaming: 0 even while frames flow at 60 fps — cosmetic quirk, ignore it.
  • TLS: stock firmware serves plain HTTP on the LAN. The device supports optional TLS (Settings → Advanced) — enable it and set JETKVM_URL=https://…, plus JETKVM_VERIFY_TLS=true if the cert is trusted. WebRTC media/control is DTLS/SRTP-encrypted peer-to-peer regardless of how the signaling travelled.
  • power needs the ATX extension board wired to the motherboard header; without it the tool is a no-op (power_state reads power: false).

Safety

This lets a language model drive a real computer with real consequences. Recommendations:

  • Point it at a test box or lab machine first, not your production NAS.
  • The destructive tools are power, reboot_device, dc_power, mount_*, delete_storage_file, and any press_key of a reboot chord — consider requiring per-call confirmation for them in your MCP client's permission settings.
  • Set a device password (and TLS) if the JetKVM is reachable by anyone but you.

Development

jetkvm/client.py   WebRTC + JSON-RPC client (connect, snapshot, HID input, uploads)
jetkvm/keymap.py   ASCII / key-combo → USB HID usage codes
server.py          FastMCP server exposing the 24 tools
smoke_test.py      live end-to-end check against a real device
docs/              architecture + tool reference

Validated end-to-end against a JetKVM v2 on firmware/app 0.5.8 (Jul 2026): sustained 60 fps decode, keyboard input, HTTP CDROM mount/unmount, ATX/DC state reads — including driving it from a live Claude session. The RPC surface is verified against the jetkvm/kvm source.

Ideas / roadmap

  • Keyboard layout profiles for type_text (US hardcoded today)
  • Gate destructive tools behind an env flag
  • Native getSnapshot RPC upstream in the firmware would remove the H.264 decode dependency entirely (see jetkvm/kvm#1459)

License

MIT. Not affiliated with JetKVM/Improve Robotics — this is an independent client of the device's public API.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选