Splunk Automated Triage MCP Server

Splunk Automated Triage MCP Server

Provides tools to search Splunk logs, inject test data, and send triage emails, enabling AI-driven incident investigation.

Category
访问服务器

README

Agentic Automated Triage for Splunk

CI

This is an agentic incident-triage pipeline for Splunk: a saved-search alert fires a webhook → an orchestrator starts a Claude tool-use conversation → Claude autonomously drives a chain of SPL investigation queries through three MCP tools (search → populate-if-empty → re-search → email) → an AI-written triage report — severity, root-cause hypothesis, breakdowns, and a clickable Splunk deep link — lands in your inbox.

This isn't a static dashboard or a fixed alert template. The agent decides, on each run, what to search, whether the result set needs enriching, and how to characterise the incident — the same three tools are exposed both to the orchestrator's tool-use loop and as a standalone MCP server, so any MCP-compatible client can drive the same investigation.

 Splunk saved-search alert  (index=triage_demo, error_count > 0, every 1 min)
        │  webhook action (HTTP POST JSON)
        ▼
 ┌───────────────────────────┐        ┌──────────────────────────────┐
 │ webhook listener (FastAPI)│  call  │ orchestrator (Anthropic loop) │
 │   POST /webhook  :5001    │ ─────► │   model = claude-sonnet-4-6   │
 └───────────────────────────┘        └───────────────┬──────────────┘
                                                       │ tool-use
            ┌──────────────────────────────────────────┼─────────────────────────┐
            ▼                                           ▼                          ▼
   search_splunk_logs                      populate_splunk_test_data        send_email
   (Splunk REST :8089)                     (Splunk HEC :8088)               (SMTP/mailpit :1025)
            └────────────── same triage.tools module also exposed by the MCP server ┘
                                  (triage.mcp_server, MCP/SSE :8050)

The three tools live in one module (triage/tools.py). That module is exposed two ways: as MCP tools by triage/mcp_server.py (the "one MCP server"), and called directly by the orchestrator's tool-use loop. One implementation, two surfaces.


Components

Path Role
triage/tools.py The 3 tools: search_splunk_logs, populate_splunk_test_data, send_email (shared)
triage/splunk_client.py Splunk REST search + HEC inject (Bearer-token auth only)
triage/deeplink.py Builds the clickable Splunk Web search URL for the email
triage/mcp_server.py FastMCP server exposing the 3 tools (MCP/SSE on :8050)
triage/orchestrator.py Anthropic tool-use loop — the agent brain
triage/webhook.py FastAPI listener: /webhook, /test-triage, /health
scripts/setup_splunk.py Mints tokens, creates index + HEC + the webhook alert
scripts/trigger_alert.py Injects events to fire the alert, or --manual posts a synthetic alert
scripts/verify_email.py Polls mailpit and prints the delivered email
tests/test_e2e.py Drives the whole chain and asserts the email arrived
docker-compose.yml Brings up the MCP server + webhook/orchestrator

Prerequisites (this dev box)

These already run as standalone dev containers (Docker Desktop auto-starts them):

Container Ports Used for
splunk-dev (Splunk Enterprise) 8000 web, 8089 REST, 8088 HEC searches + data injection
mailpit (test SMTP) 1025 SMTP, 8025 web UI receiving the triage email

Check: docker ps should show both. Python 3.12 on the host is only needed for the scripts/ helpers (py on this machine — the bare python alias is the broken MS Store stub).


Quick start

cd agentic-automated-triage-for-splunk

# 1. Prepare Splunk: mint tokens, create index + HEC + the webhook alert.
#    Writes SPLUNK_API_TOKEN + SPLUNK_HEC_TOKEN into .env (created from .env.example).
py scripts\setup_splunk.py

# 2. (OPTIONAL) Add your Claude API key to .env  (>>> SUBSTITUTE <<<)
#    ANTHROPIC_API_KEY=sk-ant-...
#    Leave it blank to run the deterministic "scripted" mode (see Run modes below) —
#    that is how the boss demo is driven, no key required.

# 3. Bring up the pipeline (MCP server + webhook/orchestrator).
docker compose up --build -d

# 4a. Immediate end-to-end run (no waiting for Splunk's scheduler):
py scripts\trigger_alert.py --manual

# 4b. ...or the real path: inject errors and let the scheduled alert fire (~1 min):
py scripts\trigger_alert.py --count 30

# 5. Verify the email arrived.
py scripts\verify_email.py --subject "[Triage]"
#    ...or just open the mailbox: http://localhost:8025

Automated check of the whole chain:

py tests\test_e2e.py

How the agent behaves

On each alert the orchestrator (Claude) is instructed to:

  1. search_splunk_logs for the alert's index over the last 15 minutes.
  2. If that returns no/insufficient data → populate_splunk_test_data (realistic sample events via HEC, stamped now) → search_splunk_logs again to confirm.
  3. Summarise: counts, top error codes / affected services, a P1–P4 severity, next actions.
  4. send_email once — subject starts [Triage], body includes the alert name, findings, the exact SPL, the severity, and the Splunk deep link (in links).

The /test-triage endpoint seeds an empty result set on purpose, so it always exercises the populate-then-re-search branch.

Run modes (with or without an API key)

run_triage() reports its mode explicitly:

  • agentic — ANTHROPIC_API_KEY is set. Claude drives the tools in a real tool-use loop and chooses the sequence itself.
  • scripted — no key. A deterministic stand-in performs the identical documented procedure (search → populate-if-empty → re-search → stats breakdowns → email) with no model call, so the full pipeline — and the rich HTML report — is demonstrable offline. This is the mode the boss demo runs in. Flip to agentic any time by adding the key and docker compose up -d --force-recreate.

Either way the email is built by triage/report.py from real Splunk stats results, so the breakdowns (top error codes, affected services, regions, latency, severity) are genuine aggregates of the indexed events — not hard-coded.

Resetting the demo data

Injected events stay inside the 15-minute search window for ~15 min, so repeated fires within that window stack up. For a pristine single-incident screenshot, clear the index first (admin Bearer token, config-safe — no index/HEC teardown):

# deletes all events in triage_demo; the next fire re-populates a clean, skewed batch
curl.exe -sk -H "Authorization: Bearer $env:SPLUNK_API_TOKEN" `
  https://127.0.0.1:8089/services/search/jobs `
  --data-urlencode "search=search index=triage_demo | delete" `
  -d exec_mode=oneshot -d output_mode=json -d earliest_time=-24h -d latest_time=now

Verifying the email step

  • Web UI: open http://localhost:8025 — the triage email appears at the top.
  • CLI: py scripts\verify_email.py --subject "[Triage]" prints From/To/Subject/body and exits 0 on success.
  • API: curl http://localhost:8025/api/v1/messages returns the JSON message list.

>>> SUBSTITUTE for your environment <<<

Everything is env-driven via .env (copied from .env.example). Flagged values:

Variable Default (this dev box) Substitute when…
ANTHROPIC_API_KEY (empty) always — your Claude key sk-ant-...
SPLUNK_PASSWORD changeme-dev-1 your splunk-dev admin password differs
SPLUNK_API_TOKEN / SPLUNK_HEC_TOKEN (minted) auto-filled by setup_splunk.py; replace if pointing at a different Splunk
SPLUNK_HOST host.docker.internal running processes on the host → 127.0.0.1; real Splunk Cloud → its hostname
SPLUNK_WEB_BASE / SPLUNK_WEB_LOCALE http://localhost:8000 / en-US Splunk Cloud → stack URL + en-GB
SPLUNK_VERIFY_SSL false production → true (or a CA-bundle path)
SMTP_HOST / SMTP_PORT host.docker.internal / 1025 a real mail relay
SMTP_TO / SMTP_FROM *.local.test real recipient/sender

Splunk Cloud note: the previous Victoria trial (prd-p-6oxft) was decommissioned, so this pipeline targets the local splunk-dev container. To repoint at Splunk Cloud, set SPLUNK_HOST to the stack host, supply an ACS/HEC token, and create the alert via the ACS API instead of setup_splunk.py's REST call. All runtime auth stays Bearer-token.


Auth model

Runtime auth is Bearer tokens only — no admin:password, no Basic header, no -u (matches the repo-wide rule). setup_splunk.py performs a single bootstrap form-login (/services/auth/login, a session key — not a Basic header) purely to mint the JWT auth token and HEC token; every subsequent call uses those tokens.

Guardrails

The agent composes its own SPL at runtime, so it is never trusted to behave — it is constrained:

  • Read-only SPL guard (triage/spl_guard.py) — every agent-issued query passes through a deny-by-default gate at the single Splunk choke point (splunk_client.run_search) before it reaches the REST API. Commands that mutate state or exfiltrate data (delete, collect, outputlookup, sendemail, script, …) are blocked as pipeline commands — the literal word "delete" appearing in log text still searches fine. Unit-tested offline in tests/test_spl_guard.py (runs in CI).
  • Bounded action surface — the agent has exactly three tools (search, populate test data, send email). It cannot touch Splunk config, users, or apps; the only outbound side effect is the triage email, and in the lab that lands in mailpit, not a real mailbox.
  • Secrets stay in the environment — API keys and tokens come from .env (gitignored, .env.example is placeholders-only); nothing is hardcoded and nothing is echoed into the report.
  • Deterministic fallback — with no ANTHROPIC_API_KEY set, the pipeline runs a scripted mode that exercises the identical tool chain, so the guardrails are testable without a live model.

The MCP server on its own

triage/mcp_server.py is a standalone MCP server you can point any MCP client at:

# SSE on :8050 (default)
py -m triage.mcp_server
# or stdio
$env:MCP_TRANSPORT="stdio"; py -m triage.mcp_server

It exposes exactly search_splunk_logs, populate_splunk_test_data, send_email.


Verification

See VERIFICATION.md for a full end-to-end verification record — the exact tool-call sequence observed, the rendered report contents, and the reproduction steps used to confirm the pipeline works as described.

License

MIT — see LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选