Enterprise Infrastructure & Metrics MCP Server
Provides read-only access to host system metrics (CPU, memory, disk), Docker container health/logs, and sandboxed log file analysis via MCP tools, enabling AI agents to monitor enterprise infrastructure safely.
README
Enterprise Infrastructure & Metrics MCP Server
A production-grade Model Context Protocol (MCP) server, written in native async Python, that gives an LLM agent (Claude Desktop, Claude Code, or any MCP-compatible host) safe, read-only introspection into:
- Host system resources — live CPU (per-core), memory, and disk metrics via
psutil - Docker containers — list, inspect health/state, and tail logs via the Docker SDK
- Application/system logs — sandboxed tail-and-filter of log files, with strict path-traversal protection
Built as a portfolio project demonstrating enterprise MCP server engineering: strict Pydantic v2 contracts, structured (never-raise) error handling, async-safe wrapping of blocking I/O, and an explicit, auditable security boundary.
Architecture
┌──────────────────────────┐ stdio (JSON-RPC 2.0) ┌───────────────────────────────────────────┐
│ MCP Host │ <---------------------------------> │ enterprise-mcp-server (this project) │
│ (Claude Desktop / │ subprocess, stdin/stdout │ │
│ Claude Code / other) │ │ ┌─────────────────────────────────────┐ │
└──────────────────────────┘ │ │ server.py (FastMCP app) │ │
│ │ - registers 3 tools │ │
│ │ - stdio transport │ │
│ │ - logging -> stderr ONLY │ │
│ └───────────────┬─────────────────────┘ │
│ │ validated Pydantic │
│ ▼ input models │
│ ┌─────────────────────────────────────┐ │
│ │ tools/ │ │
│ │ ├─ system_metrics.py (psutil) │ │
│ │ ├─ docker_manager.py (docker SDK) │ │
│ │ └─ log_analyzer.py (sandboxed FS)│ │
│ └───────────────┬─────────────────────┘ │
│ │ asyncio.to_thread │
│ │ (never blocks loop) │
│ ┌────────────────▼─────────────────────┐│
│ │ utils/ ││
│ │ ├─ security.py (path sanitization) ││
│ │ ├─ errors.py (structured JSON) ││
│ │ └─ retry.py (jittered backoff) ││
│ └────────────────────────────────────────│
└───────────────┬───────────────┬───────────┘
│ │
┌───────────────▼───┐ ┌────────▼──────────┐
│ Host OS │ │ Docker daemon │
│ /proc, psutil │ │ /var/run/ │
│ sandboxed log dir │ │ docker.sock │
└────────────────────┘ └────────────────────┘
Request lifecycle: MCP host → JSON-RPC tools/call over stdin → FastMCP
parses & validates arguments against the tool's Pydantic input model →
tool function executes, wrapping every blocking call (psutil, Docker SDK,
file I/O) in asyncio.to_thread → result serialized to JSON (success
payload or structured ToolError — the function never raises past this
boundary) → written to stdout as the JSON-RPC response.
The three tools
| Tool | Purpose | Mutates host state? |
|---|---|---|
get_system_metrics |
CPU (aggregate + per-core), memory (RAM + swap), disk (usage + I/O counters) | No |
manage_docker_containers |
List containers, inspect health/config, tail container logs | No — read-only by design |
analyze_local_logs |
Tail + filter a log file inside a sandboxed root directory | No — read-only, sandboxed |
Security boundaries
This project treats the LLM as an untrusted caller operating a read-only monitoring surface, not an operator with host control. Three concrete boundaries enforce that:
-
manage_docker_containersexposes no lifecycle verbs. The Docker SDK and daemon support starting, stopping, restarting, executing commands in, and removing containers. None of that is wired up. Onlylist_containers,inspect_container, andget_container_logsexist as actions — an LLM cannot use this server to take down a container or run arbitrary commands inside one, even if prompted to. -
analyze_local_logsis sandboxed to a single, explicit root directory (MCP_LOG_ROOT_DIR, default/var/log), enforced inutils/security.py. Every requested path is:- stripped of leading
/and..segments (blocks absolute-path override), - joined onto the resolved root,
- resolved again with
Path.resolve()(collapses remaining..segments and follows symlinks, closing the symlink-escape vector), - and finally checked with
Path.is_relative_to()against the resolved root before any file is opened.
A request for
../../etc/shadow,/etc/shadow, or a symlink inside the sandbox that points outside it is rejected with a structuredPATH_TRAVERSAL_BLOCKEDerror — never a Python traceback, and never a silent read. - stripped of leading
-
Secrets are never echoed back.
inspect_containerreturnsenv_var_count(an integer), not the environment variables themselves, since container env vars routinely contain credentials and API keys. -
Every response size is bounded.
MCP_MAX_LOG_LINEShard-caps log tails server-side regardless of what a caller requests, and Docker log tails are capped at 500 lines — both protect the LLM's context window and prevent a single tool call from returning gigabytes of data.
Reliability & engineering standards
- Never-raise tool boundary: every tool function wraps its entire body
in
try/exceptand returns a structured JSONToolError(utils/errors.py) on failure — malformed input, a missing container, a down Docker daemon, or a permissions error all produce a well-formed, LLM-parseable payload instead of crashing the server process. - Async-safe by construction:
psutil, thedockerSDK, and file I/O are all synchronous/blocking under the hood. Every call site wraps them inasyncio.to_threadso a slow disk read or a stalled Docker socket cannot stall the event loop and starve other concurrent tool calls. - Jittered exponential backoff (
utils/retry.py) around Docker daemon calls, since a momentarily busy socket is a transient condition worth retrying — capped atMCP_TOOL_RETRY_ATTEMPTSattempts. - Strict Pydantic v2 contracts (
schemas.py) for every tool's input and output. FastMCP derives the JSON Schema exposed to the LLM host directly from the input models, so the tool's documented contract and its runtime validation can never drift apart. - stdout is sacred: the stdio transport uses stdout exclusively for
JSON-RPC frames. All logging is configured to write to stderr
(
server.py) — a strayprint()or misconfigured logger on stdout would silently corrupt the protocol stream for every connected host.
Project layout
enterprise-mcp-server/
├── pyproject.toml
├── README.md
├── .env.example
├── src/
│ └── enterprise_mcp_server/
│ ├── __init__.py
│ ├── server.py # FastMCP app, tool registration, stdio entrypoint
│ ├── config.py # Env-driven settings + security boundary (log root)
│ ├── schemas.py # Pydantic v2 input/output contracts for all tools
│ ├── tools/
│ │ ├── system_metrics.py
│ │ ├── docker_manager.py
│ │ └── log_analyzer.py
│ └── utils/
│ ├── security.py # Path-traversal sanitization
│ ├── errors.py # Structured ToolError contract
│ └── retry.py # Jittered exponential backoff
└── tests/
├── test_security.py
└── test_system_metrics.py
Setup
Prerequisites
- Python 3.11+
- Docker (optional — only required for
manage_docker_containers; the other two tools work without it)
Install
git clone https://github.com/<your-username>/enterprise-mcp-server.git
cd enterprise-mcp-server
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
Configure (optional)
Copy .env.example to .env and adjust as needed, or export directly:
export MCP_LOG_ROOT_DIR=/var/log # sandbox root for analyze_local_logs
export MCP_MAX_LOG_LINES=1000 # hard ceiling on lines returned
export MCP_DOCKER_TIMEOUT_SECONDS=10 # Docker daemon socket timeout
export MCP_TOOL_RETRY_ATTEMPTS=3 # retry attempts for transient failures
Run standalone (for smoke-testing)
enterprise-mcp-server
# or
python -m enterprise_mcp_server.server
The process will sit waiting for JSON-RPC frames on stdin — this is expected; it's designed to be launched by an MCP host, not run interactively. Use the MCP Inspector (below) for interactive testing.
Test with MCP Inspector
npx @modelcontextprotocol/inspector enterprise-mcp-server
This opens a browser UI where you can call each tool directly and inspect
the JSON Schema FastMCP generated from schemas.py.
Run the test suite
pytest -v
Claude Desktop configuration
Add the following to your Claude Desktop MCP config file
(~/Library/Application Support/Claude/claude_desktop_config.json on
macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):
{
"mcpServers": {
"enterprise-infra-metrics": {
"command": "/absolute/path/to/enterprise-mcp-server/.venv/bin/enterprise-mcp-server",
"args": [],
"env": {
"MCP_LOG_ROOT_DIR": "/var/log",
"MCP_MAX_LOG_LINES": "1000",
"MCP_DOCKER_TIMEOUT_SECONDS": "10"
}
}
}
}
Note: Claude Desktop launches this as a subprocess with a minimal environment, so
commandmust be the absolute path to the virtualenv's console script (not a bareenterprise-mcp-server, which relies onPATHbeing inherited — it usually isn't).
Restart Claude Desktop, and the hammer icon in the composer should show
get_system_metrics, manage_docker_containers, and analyze_local_logs
as available tools.
Example interactions
"Is my machine under memory pressure right now?" → calls
get_system_metrics(scope="memory")
"Are any of my Docker containers unhealthy?" → calls
manage_docker_containers(action="list_containers"), theninspect_containeron anything with a non-healthystatus
"Check
nginx/access.logfor the last hour's errors" → callsanalyze_local_logs(relative_log_path="nginx/access.log", severity_filter="error_and_above")
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。