slurm-mcp
Enables agents to query Slurm scheduler state safely through a read-only allowlist, with progressive disclosure to minimize context usage.
README
slurm-mcp
A read-only MCP server exposing Slurm scheduler state to agents. The allowlist is enforced in code, not requested in a prompt, and the tool surface uses progressive disclosure so a Slurm-shaped question does not cost a Slurm-shaped context window.
v0.1.0. The guard, the topic surface, and the stdio server are built and tested. Runs against a real cluster or against recorded fixtures with no Slurm installed.
Why the guard is in code
A system prompt saying "only use read-only commands" is a request, not a
control. It fails open. A jailbreak, a confused tool call, or an ordinary
hallucination is enough to reach scontrol update on a production controller.
Anything that could drain a node must be impossible to express, not merely
discouraged. So the allowlist lives in guard.py and
runs on every invocation:
- only eight read binaries may run at all;
scontrolis permitted forshowand refused forupdate,reconfigure,shutdown,reboot,requeue,hold,power, and fifteen more;- shell metacharacters in arguments are refused, and commands execute with
shell=Falseanyway — defence in depth, not the only barrier.
The threat model is not a malicious user. It is an agent that has read a confusing log line at 3am and is about to do something decisive.
The tests drive this with 51 real mutating and injection attempts rather than asserting on prompt text, because a prompt-level promise cannot be tested:
make test
Progressive disclosure
The obvious design exposes one tool per binary — sinfo, squeue, sacct,
sdiag, sprio, sshare, scontrol, sacctmgr — each carrying a schema for
its flags. Slurm's flag surface is enormous and most of it is irrelevant to any
given question, but all of it sits in context on every turn.
This server exposes three tools:
| tool | when |
|---|---|
slurm_overview |
the snapshot most sessions open with — nodes, queue, diagnostics in one call |
slurm_query |
one topic plus optional filters. Topics are a closed vocabulary, not a command line |
slurm_describe |
column meanings and filters for one topic, fetched only when needed |
Measured, and reproducible with make footprint:
resident, three tools 1088 chars
detail, fetched on request 3029 chars
flat one-tool-per-binary 4441 chars (4.1x resident)
That is a context-cost measurement, not a quality claim. It says the detail is not resident until asked for; it does not say the agent will diagnose anything better. Nothing here has been scored on slurm-rca-bench.
Topics: queue, nodes, accounting, priority, fairshare,
diagnostics, config.
Quickstart
No cluster required.
make install
make surface # the read-only allowlist, and what is denied
make demo # a full overview against recorded fixtures
make footprint # reproduce the context measurement above
make check # ruff, ruff format, mypy --strict, pytest
Against a real cluster, drop --fixtures:
slurm-mcp overview
slurm-mcp query queue --filter user=alice
slurm-mcp describe accounting
As an MCP server
pip install -e ".[server]"
slurm-mcp serve # stdio; add --fixtures to serve recorded data
Point any MCP client at that command. Every response is labelled [live] or
[fixture] so recorded data can never be mistaken for a cluster read.
Limitations
- Read-only is enforced for this server's own surface. It does not sandbox the host, and it says nothing about what other tools an agent has been given.
- The fixtures are illustrative, not a benchmark. They are a plausible cluster shape used to make the surface runnable; no result should be quoted from them.
- No scoring. The claim here is about the tool surface, not diagnostic
accuracy. Measuring whether progressive disclosure changes an agent's
diagnosis is
cluster-sre-agent's job, and it has not been done. - Single Slurm dialect. Output formats are tested against Slurm 25.x column layouts. Older versions may differ.
sacctcan block. A degraded accounting path makes it hang rather than error; calls time out at 20s and say so, because an agent waiting forever is worse than an agent told the read failed.
Related
- cluster-sre-agent — the agent that consumes a surface like this one. Its internal read-only tool layer is where this design came from; this repo is the standalone server.
- cluster-ops-skills — the runbooks an agent follows once it can read the cluster.
- slurm-rca-bench — where a claim about diagnostic quality would have to be proven.
License
MIT.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。