es-error-lens

es-error-lens

Enables LLM agents to search and analyze Elasticsearch logs for errors, detect recurring patterns, analyze error-rate trends, and retrieve full trace context through MCP tools.

Category
访问服务器

README

es-error-lens

CI License: MIT Python 3.10+ Offline tests

An MCP server that gives LLM agents eyes on your Elasticsearch logs. Point it at a cluster holding ECS-format logs and any MCP client (Claude Desktop, Claude Code, or your own agent) can search errors, detect recurring patterns, analyze error-rate trends, and pull full trace context — the queries an engineer runs by hand during triage, exposed as tools.

Tool What it answers
search_errors "What's failing right now?" — filtered by window, service, level
get_error_patterns "Is this systemic or a one-off?" — message aggregation with occurrence counts
analyze_error_trend "Is it getting worse?" — time-series histogram, peak detection, trend verdict
get_error_context "What led up to this?" — every log sharing the error's trace ID, in order
compare_errors "Are these two failures related?" — attribute + fuzzy-message comparison
health_check "Can I even reach the cluster?"

Elasticsearch is spoken to over its plain REST API via httpx — no Elasticsearch client dependency, and the HTTP layer is transport-injectable, so the entire server runs offline against the bundled FakeElasticsearch for tests and the demo.

Demo (offline, no cluster, no API keys)

git clone https://github.com/TerraCo89/es-error-lens && cd es-error-lens
python -m venv .venv && source .venv/bin/activate   # .venv\Scripts\activate on Windows
pip install -e ".[dev]"

pytest              # 22 tests, all offline, < 1 second
es-error-lens-demo  # synthetic incident walk-through

The demo mounts FakeElasticsearch with a deterministic 60-document ECS corpus simulating a small incident — a steady payment-provider failure, an escalating search-timeout problem, and background noise — then triages it:

[get_error_patterns] 2 recurring patterns in 57 errors:
   42x  [TimeoutError] Search query timed out after 30s  (search-api)
   13x  [UpstreamHTTPError] Payment provider returned 502 Bad Gateway  (checkout-api)
  error types: TimeoutError=42, UpstreamHTTPError=13, TemplateError=1, ConnectionResetError=1

[analyze_error_trend] search-api: 42 errors, trend=INCREASING
  peak: 8 errors at 2026-07-10T06:00:00Z

[get_error_context] trace-checkout-7f3a -> [UpstreamHTTPError] Payment provider returned 502 Bad Gateway
  3 related logs on the same trace:
    info  POST /checkout started
    info  Cart validated: 3 items
    warn  Payment provider latency 4100ms exceeds SLO

Demo result: OK

Using it against a real cluster

export ELASTICSEARCH_URL=http://localhost:9200   # default
export ES_INDEX_PATTERN=logs-*                   # default

es-error-lens                                    # stdio, for Claude Desktop / Claude Code
es-error-lens --transport streamable-http --port 8080   # HTTP clients

Claude Desktop / Claude Code config:

{
  "mcpServers": {
    "es-error-lens": {
      "command": "es-error-lens",
      "env": { "ELASTICSEARCH_URL": "http://localhost:9200" }
    }
  }
}

Expected document shape — standard ECS fields, the ones Filebeat and the ECS logging libraries emit by default: @timestamp, log.level, message, service.name, error.type, error.stack_trace, trace.id, host.name. Anything missing degrades gracefully (fields come back null).

Testing without a cluster

FakeElasticsearch (also exported) implements just enough of the _search API for this server: bool queries with term/range clauses, sorting, terms aggregations with top_hits sub-aggregations, and date_histogram with fixed intervals. Mount it in your own tests:

from es_error_lens import FakeElasticsearch, set_transport, search_errors

set_transport(FakeElasticsearch(my_ecs_docs).transport())
result = await search_errors(time_range="1h", service="checkout-api")

The test suite runs entirely through it — deterministic, offline, fast — and a hygiene test enforces that all bundled demo data stays synthetic (hosts under .example, no real domains or addresses).

Provenance

Extracted from a personal observability platform where it fronts the Elasticsearch instance that aggregates structured logs from a fleet of side-project services, letting coding agents triage production errors during development sessions. All demo and test data in this repository is synthetic.

License

MIT © 2026 Kris Cernjavic

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选