o2-mcp
Enables interaction with the HMS O2 cluster via SSH, allowing job submission, remote command execution, file transfers, and disk management while minimizing Duo authentication prompts.
README
o2-mcp
Generic, project-agnostic access to the HMS O2 cluster, exposed both as a Python
library (o2mcp) and as an MCP server (o2-mcp) so an agent can submit Slurm work,
run remote commands, monitor jobs, move files, and keep disk tidy — without triggering a
Duo push on every action.
Extracted from clock-oscillation-analysis so the cluster tooling is shared
infrastructure (used by multiple analysis projects) rather than living inside one of them.
Project-specific layers (e.g. run-organization for a particular pipeline) build on this
package rather than living in it.
Duo model (read this first)
HMS O2 uses Duo autopush: every new SSH connection fires a Duo push, even key-only / BatchMode. The tools are built so this costs exactly one push per session:
- Call
o2_start_masteronce (one push you approve) to open a persistent SSH ControlMaster (stays up ~8h viaControlPersist). - Every other tool reuses that master and costs no additional push.
Never open the master in a loop or run these tools on a short timer — a periodic reconnect
is what causes "a Duo call every minute". The .agent_locks/O2_DISABLED lock file is a hard
stop honored by every operation.
Install
# The core (config/connection/sync/slurm/async_transfer/keepalive/workspace) is pure-stdlib
# and runs on Python 3.9. The MCP server needs the mcp SDK (Python >= 3.10):
pip install -e ".[o2]" # on a 3.10+ env
MCP server config
{
"mcpServers": {
"o2": {
"type": "stdio",
"command": "/path/to/venv/bin/o2-mcp",
"env": {
"O2_SSH_HOST_ALIAS": "o2",
"O2_SSH_TRANSFER_ALIAS": "o2-transfer",
"O2_SSH_LOCK_FILE": "/path/to/.agent_locks/O2_DISABLED"
}
}
}
}
Requires Host o2 (and optionally Host o2-transfer) blocks in ~/.ssh/config with
ControlMaster auto + a ControlPath socket.
Tools
| Tool | Purpose | Hint |
|---|---|---|
o2_status |
Lock state, ControlMaster state, hostname; whoami; date probe |
read-only |
o2_start_master |
Open the persistent SSH master (needs allow_new_login) |
write |
o2_run |
Run an arbitrary command on a login node | write |
o2_submit_job |
sbatch a script (existing path or staged script_text); returns the job id |
write |
o2_squeue |
squeue -u <user> as structured rows |
read-only |
o2_job_status |
sacct -j <id> accounting (state, elapsed, exit code, MaxRSS) |
read-only |
o2_tail_log |
Tail a remote log file | read-only |
o2_cancel_job |
scancel <id> |
destructive |
o2_push / o2_pull |
rsync up/down (reuses the master; use_transfer_node for big moves) |
write |
o2_push_async / o2_pull_async |
Non-blocking rsync: launch detached, return a transfer_id immediately |
write |
o2_transfer_status |
Progress/state of async transfers (running/done/failed/crashed); omit id to list all |
read-only |
o2_transfer_cancel |
SIGTERM a running async transfer's process group | destructive |
o2_disk_report |
Per-tier usage + hygiene flags (regenerable/redundant/misplaced) | read-only |
o2_workspace_gc |
Prune regenerable + redundant disk (detached, dry-run default) | destructive |
o2_place |
Resolve the canonical output path for a kind (+project) per tier | read-only |
Non-blocking transfers
o2_push_async / o2_pull_async launch a detached rsync and return a transfer_id right
away, so the agent can keep working and poll o2_transfer_status instead of blocking a tool
call for a multi-GB transfer. The transfer keeps running between tool calls and survives an
MCP-server restart (a wrapper records rsync's exit code to disk); re-running the same command
resumes it (rsync --partial). Remote paths are escaped so spaces transfer intact while
~/$VAR/${VAR} still expand. State lives under ~/.cache/clock_o2_mcp/transfers
(O2_ASYNC_STATE_DIR to override).
Safety contract
- All SSH uses
BatchMode=yes(public key only) — a dead master or missing key fails fast instead of triggering an interactive MFA prompt. - Remote commands run only through an already-established ControlMaster; opening a new login
requires explicit opt-in (
allow_new_login). - The
O2_DISABLEDlock hard-stops every operation. - Destructive/transfer-node operations default to dry-run where applicable and verify before freeing scratch.
Development
pip install -e ".[dev,o2]"
ruff check src tests && black --check src tests && pytest -m "not o2" -q
The core stays import-light (stdlib only); mcp/pydantic/anyio are needed only by the
server. Tests inject the subprocess seam, so they run fully offline (no cluster, no network).
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。