slurm-mcp

slurm-mcp

Enables agents to query Slurm scheduler state safely through a read-only allowlist, with progressive disclosure to minimize context usage.

Category
访问服务器

README

slurm-mcp

A read-only MCP server exposing Slurm scheduler state to agents. The allowlist is enforced in code, not requested in a prompt, and the tool surface uses progressive disclosure so a Slurm-shaped question does not cost a Slurm-shaped context window.

v0.1.0. The guard, the topic surface, and the stdio server are built and tested. Runs against a real cluster or against recorded fixtures with no Slurm installed.


Why the guard is in code

A system prompt saying "only use read-only commands" is a request, not a control. It fails open. A jailbreak, a confused tool call, or an ordinary hallucination is enough to reach scontrol update on a production controller.

Anything that could drain a node must be impossible to express, not merely discouraged. So the allowlist lives in guard.py and runs on every invocation:

  • only eight read binaries may run at all;
  • scontrol is permitted for show and refused for update, reconfigure, shutdown, reboot, requeue, hold, power, and fifteen more;
  • shell metacharacters in arguments are refused, and commands execute with shell=False anyway — defence in depth, not the only barrier.

The threat model is not a malicious user. It is an agent that has read a confusing log line at 3am and is about to do something decisive.

The tests drive this with 51 real mutating and injection attempts rather than asserting on prompt text, because a prompt-level promise cannot be tested:

make test

Progressive disclosure

The obvious design exposes one tool per binary — sinfo, squeue, sacct, sdiag, sprio, sshare, scontrol, sacctmgr — each carrying a schema for its flags. Slurm's flag surface is enormous and most of it is irrelevant to any given question, but all of it sits in context on every turn.

This server exposes three tools:

tool when
slurm_overview the snapshot most sessions open with — nodes, queue, diagnostics in one call
slurm_query one topic plus optional filters. Topics are a closed vocabulary, not a command line
slurm_describe column meanings and filters for one topic, fetched only when needed

Measured, and reproducible with make footprint:

  resident, three tools            1088 chars
  detail, fetched on request       3029 chars
  flat one-tool-per-binary         4441 chars  (4.1x resident)

That is a context-cost measurement, not a quality claim. It says the detail is not resident until asked for; it does not say the agent will diagnose anything better. Nothing here has been scored on slurm-rca-bench.

Topics: queue, nodes, accounting, priority, fairshare, diagnostics, config.

Quickstart

No cluster required.

make install
make surface      # the read-only allowlist, and what is denied
make demo         # a full overview against recorded fixtures
make footprint    # reproduce the context measurement above
make check        # ruff, ruff format, mypy --strict, pytest

Against a real cluster, drop --fixtures:

slurm-mcp overview
slurm-mcp query queue --filter user=alice
slurm-mcp describe accounting

As an MCP server

pip install -e ".[server]"
slurm-mcp serve              # stdio; add --fixtures to serve recorded data

Point any MCP client at that command. Every response is labelled [live] or [fixture] so recorded data can never be mistaken for a cluster read.

Limitations

  • Read-only is enforced for this server's own surface. It does not sandbox the host, and it says nothing about what other tools an agent has been given.
  • The fixtures are illustrative, not a benchmark. They are a plausible cluster shape used to make the surface runnable; no result should be quoted from them.
  • No scoring. The claim here is about the tool surface, not diagnostic accuracy. Measuring whether progressive disclosure changes an agent's diagnosis is cluster-sre-agent's job, and it has not been done.
  • Single Slurm dialect. Output formats are tested against Slurm 25.x column layouts. Older versions may differ.
  • sacct can block. A degraded accounting path makes it hang rather than error; calls time out at 20s and say so, because an agent waiting forever is worse than an agent told the read failed.

Related

  • cluster-sre-agent — the agent that consumes a surface like this one. Its internal read-only tool layer is where this design came from; this repo is the standalone server.
  • cluster-ops-skills — the runbooks an agent follows once it can read the cluster.
  • slurm-rca-bench — where a claim about diagnostic quality would have to be proven.

License

MIT.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选