opsagent

opsagent

A read-only MCP server that exposes Kubernetes cluster telemetry tools (pods, events, logs, metrics, ArgoCD syncs) with automatic redaction for incident triage and hypothesis ranking.

Category
访问服务器

README

ai-automation-lab

CI

Python n8n MCP pytest ArgoCD

An incident triage agent that runs against my own K3s cluster. When Alertmanager fires, it investigates using read-only access to the cluster's telemetry and returns a ranked hypothesis a human can act on. Then it records whether it was right.

That last part is the point. Piping an alert into a language model is a weekend project. Measuring whether the output was correct, bounding what it costs, and proving it cannot touch anything it should not, is the actual work.

This is the third lab in a series: devops-homelab-k3s-hybrid-cloud is the platform it watches, and qa-engineering-lab is the test suite that found six real defects in that platform.

Architecture

flowchart TB
    subgraph cluster["K3s cluster"]
        AM["Alertmanager"] -->|webhook| N8N["n8n<br/>workflows deployed from git"]
        N8N -->|"POST /investigations"| AGENT["opsagent<br/>FastAPI + agent loop"]

        AGENT -->|"read-only ServiceAccount"| TOOLS["tool layer"]
        TOOLS --> K8S["Kubernetes API<br/>pods, events, deploys"]
        TOOLS --> LOKI["Loki<br/>container logs"]
        TOOLS --> PROM["Prometheus<br/>PromQL"]
        TOOLS --> ARGO["ArgoCD<br/>sync history"]

        TOOLS -->|redaction| AGENT
        AGENT --> PG[("PostgreSQL<br/>investigations, cost, verdicts")]
    end

    AGENT -->|"redacted prompt"| LLM["LLM provider<br/>mock by default"]
    N8N --> TG["Telegram"]
    N8N --> GH["GitHub issue<br/>new alert class only"]
    HUMAN["me"] -->|"actual root cause"| PG
    PG --> EVAL["accuracy report"]

Two properties are structural rather than conventional. Tool output passes through redaction before it reaches the model, so nothing unredacted can leave the cluster even if the agent misbehaves. And the model never executes anything: it reads, it reasons, it proposes. Remediation is out of scope for v1.

Status

Built phase by phase, and this table is the honest state of it.

Phase What it delivers State
0 Repository skeleton, tooling, CI Done
1 n8n as a GitOps workload, workflow export/import CLI Tooling done, deploy pending
2 Cluster tool layer over MCP, redaction Done
3 The agent: provider abstraction, guardrails, persistence Planned
4 Alertmanager to Telegram, resolution capture Planned
5 Metrics, Grafana dashboard, report page, runbook Planned
6 Fault injection and the accuracy evaluation Planned
7 Daily digest, manifest review bot, CVE triage Planned

The full breakdown, including the definition of done for each phase and the parts of the brief I argued against, is in plan.md.

Running it

Nothing here needs an API key, a database or cluster access. The default provider is a deterministic mock, which is also what CI uses.

uv sync
uv run pytest
uv run opsagent show-config
environment=local
log_level=INFO
log_json=None

Quality gates, the same four CI runs:

uv run ruff check .
uv run mypy
uv run pytest
uv run opsagent n8n validate

Workflow sync needs a running instance and an API key, so it is the one thing that does not work from a clean clone:

opsagent n8n export    # instance to git, produces a reviewable diff
opsagent n8n diff      # compare, exits non-zero on drift, used as a CI gate
opsagent n8n import    # git to instance, reconciles activation state
opsagent n8n validate  # offline checks, no API key needed

Driving the tools by hand

The tool layer is an MCP server before it is an agent's dependency, so the tools can be used from an editor session against the real cluster. Register it:

{
  "mcpServers": {
    "opsagent": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/ai-automation-lab", "python", "-m", "opsagent.mcp"],
      "env": { "OPSAGENT_LOKI_URL": "http://localhost:3100" }
    }
  }
}

It reads whatever kubeconfig context is active, so point it at a read-only one. The six tools are get_pod_status, get_events, query_logs, query_metrics, get_recent_deploys and get_runbook. Every result carries how many values were redacted and whether it was truncated, so a caller can never mistake a partial answer for a complete one.

Design decisions worth defending

The repository runs with zero API keys and zero spend. A reviewer who clones this gets a working system, not a README describing one. That forced the provider abstraction to exist from the start rather than being retrofitted.

Redaction sits at the tool boundary, not before the prompt. Putting it in the agent means every future caller of the tool layer has to remember to redact. Putting it in the tools means it is impossible to forget, and it is the most heavily tested code in the repository.

Redaction preserves identity instead of erasing it. The same address always becomes the same <ip-1>, so the model can still reason that the pod at <ip-1> cannot reach <ip-2> and correlate that across a log excerpt and an event. Masking everything to one <redacted> would destroy exactly the structure a root cause is made of.

The agent's ServiceAccount cannot read secrets and cannot write anything. The fault injection harness in phase 6 needs write access to break things on purpose, so it carries its own separate credential. The agent never gets one.

Log lines are untrusted input. Anyone who can write to a log I read can write instructions to my agent. That is in the threat model, and phase 6 measures what actually happens rather than assuming the prompt held.

Documentation

Document What it covers
plan.md Phases, definitions of done, data model, open questions
docs/adr/ Decisions and the alternatives rejected
docs/assumptions.md Everything assumed rather than verified
docs/threat-model.md Trust boundaries, RBAC scope, prompt injection
docs/cost-model.md Token and cost accounting per investigation
docs/eval-report.md Accuracy numbers, including the misses

Author

Kostiantyn Osmakov cv.batpepe.online | @batpepe

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选