KubeWhisper
An AI-DevOps MCP server that gives LLMs read-only-by-default access to Kubernetes clusters, Prometheus metrics, and GitHub Actions, enabling natural language queries about infrastructure status and safe write operations with previews.
README
🗣️ KubeWhisper
An AI-DevOps MCP server — talk to your infrastructure in plain English, safely.
KubeWhisper is a Model Context Protocol (MCP) server that gives an LLM like Claude read-only-by-default access to a live environment, so you can ask:
"Which pods are crash-looping in
staging?" "What's the p95 latency right now?" "Did the last GitHub Actions run onmainpass?" "Scale thewebdeployment to 5 — but show me the preview first."
…and the model answers grounded in your actual cluster, metrics, and CI — not generic guesses.
MCP crossed ~97M monthly SDK downloads in March 2026 and is now adopted by Anthropic, OpenAI, Google, Microsoft and Amazon. Plenty of people use MCP servers; far fewer have built one for real DevOps work. This is a small, honest, safety-first example of how.
✨ Tools
| Tool | Mode | What it does |
|---|---|---|
list_clusters |
🟢 read-only | list configured clusters + their routing hints |
resolve_cluster |
🟢 read-only | decide which cluster an alert/error belongs to |
k8s_get_resources |
🟢 read-only | pods / deployments / services / nodes / events, with status |
prometheus_query |
🟢 read-only | run an instant PromQL query |
github_actions_runs |
🟢 read-only | list recent workflow runs for a repo |
scale_deployment |
🔴 write | scale a Deployment — disabled by default, previews first, every call audited |
send_slack_message |
🟡 outbound | post an incident/remediation note to Slack |
🌐 Multi-cluster + error-context routing
Drop a clusters.json next to the server (or point KUBEWHISPER_CLUSTERS at it).
Each cluster maps to a kubeconfig context, a Prometheus URL, and routing hints:
{
"default_cluster": "prod-mumbai",
"clusters": [
{ "name": "prod-mumbai", "kube_context": "gke_proj_asia-south1_prod",
"prometheus_url": "http://prometheus.prod-mumbai:9090",
"match": { "keywords": ["prod","mumbai","live","p1"], "namespaces": ["streaming","ingest"] } },
{ "name": "staging", "kube_context": "gke_proj_asia-south1_staging",
"prometheus_url": "http://prometheus.staging:9090",
"match": { "keywords": ["staging","qa"], "namespaces": ["streaming-staging"] } }
]
}
Now the agent can route by the error context: given an alert like
"P1: transcode pods OOMing on the Mumbai live stream", it calls resolve_cluster,
matches mumbai / live / transcode → prod-mumbai, and targets that cluster's
kube context + Prometheus automatically. Every cluster-scoped tool also accepts an
explicit cluster argument. With no clusters.json, it falls back to your current
kube context. See clusters.example.json.
💬 Slack bot
Two ways to use Slack:
- Outbound notes from any MCP client — the
send_slack_messagetool (setSLACK_BOT_TOKENorSLACK_WEBHOOK_URL). - A full Slack bot (
slack_bridge.py) — @mention it in a channel ("@KubeWhisper which pods are failing in prod?") and it runs a Claude tool-use loop over these tools, routes to the right cluster, and replies in-thread.
pip install -r requirements-slack.txt
export ANTHROPIC_API_KEY=... # console.anthropic.com
export SLACK_BOT_TOKEN=xoxb-... # scopes: app_mentions:read, chat:write
export SLACK_APP_TOKEN=xapp-... # Socket Mode
python slack_bridge.py
The bot is read-only by default — scale_deployment stays preview-only unless an
operator sets KUBEWHISPER_ALLOW_WRITES=true on the bot process.
🔐 Safety model (the whole point)
- Read-only by default. The single write tool is off unless
KUBEWHISPER_ALLOW_WRITES=true. - Preview, then confirm. A write with
confirm=Falseonly describes the change. - Audit trail. Every write intent — refused, previewed, or applied — is appended to a JSONL log.
- No destructive verbs.
delete/drain/cordonare deliberately not implemented in v0.
See docs/architecture.md for the diagram.
🚀 Quickstart
# 1. Install
python -m venv .venv && source .venv/bin/activate # (Windows: .venv\Scripts\activate)
pip install -r requirements.txt
# 2. Point at a cluster + Prometheus (a local `kind` cluster is perfect)
export PROMETHEUS_URL=http://localhost:9090
# KUBEWHISPER_ALLOW_WRITES stays unset → read-only
# 3. Run it
python -m kubewhisper.server
Use it from Claude Desktop
Copy the block in config.example.json into your
claude_desktop_config.json, fix the cwd path, restart Claude Desktop, and ask it about your cluster.
🧪 Tests
pip install pytest
pytest -q # guardrail tests prove writes are refused unless explicitly enabled
CI runs lint + import check + guardrail tests on every push (see .github/workflows/ci.yml).
🗺️ Roadmap
- v0.1 — read-only tools + guarded scale + audit (this release)
- v0.2 — Terraform
plan(read-only) tool + deploy "what-changed" diff - v0.3 — eval suite proving destructive-op refusal + OPA policy layer over writes
- v0.4 — one-command
docker run, publish to an MCP registry
⚠️ Disclaimer
A learning / portfolio project. Run it against clusters you own. Keep it in read-only mode unless you fully understand the write path. Built independently, on personal infrastructure.
📄 License
MIT — see LICENSE.
Built by Abdulhussain Kanchwala · Portfolio · LinkedIn · GitHub
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。