Agent Reliability MCP server

Agent Reliability MCP server

Enables AI agents to query a curated, cited knowledge graph on testing, benchmarking, and auditing autonomous agents, returning claims with sources, confidence values, and evidence tiers through eight read-only tools over a remote streamable-HTTP endpoint with no authentication required.

Category
访问服务器

README

Agent Reliability — MCP server

Testing, benchmarking and auditing autonomous AI agents — methods, harnesses, evidence

A remote MCP server over a curated knowledge graph. Every claim it returns is bound to a registered source: the tools hand back claims with their citations and a confidence value, so an agent can show its work instead of asserting.

Nothing to install. It is a hosted streamable-HTTP endpoint:

https://agentreliability.dev/mcp

Add it to a client

Claude Code

claude mcp add --transport http agent-reliability https://agentreliability.dev/mcp

Claude Desktop / any client reading mcpServers

{
  "mcpServers": {
    "agent-reliability": {
      "type": "streamable-http",
      "url": "https://agentreliability.dev/mcp"
    }
  }
}

No API key, no account, no auth. Read-only.

Check it answers, without any client at all:

curl -s https://agentreliability.dev/mcp \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'mcp-protocol-version: 2025-06-18' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

Tools

Eight, each with an outputSchema, each returning structuredContent.

tool arguments what it does
get_overview — Corpus overview: what this instance knows, counts by type, published tags, freshness. Start here when you land and do not yet know whether this corpus can answer your question.
search query, limit? Full-text search over the knowledge graph. Accent- and apostrophe-insensitive, so query in the user's own words; every hit carries its relevance score and the fields it matched.
answer question Answer a question from the corpus. Returns the matched object's claims with sources and confidence — never an unsourced answer.
get_entity id Fetch one knowledge object by id, with its claims and the sources each claim cites.
get_topic tag List the knowledge objects carrying a tag (topics are content-backed tags).
get_related id Graph neighbours of an object: outgoing and incoming relations, each with its relation type.
get_sources object_id? The whole source registry, or just the sources cited by one object. Use it to judge the corpus before trusting it.
get_latest limit? Most recently verified knowledge objects — a freshness signal.

The intended path is get_overview → search or answer → get_entity → get_related. get_overview exists because an agent that has just arrived needs to know whether this corpus can help before it spends a call guessing.

What is in the corpus

knowledge objects 38
registered sources 31
published topics 93
type objects
entity 25
guide 9
comparison 2
faq 1
glossary 1

Subject matter: evals and benchmarks (GAIA, AgentBench, Inspect), LLM-as-judge and its failure modes, Goodhart and benchmark contamination, fault injection and chaos testing, approval gates and autonomy levels, grounding and faithfulness.

Questions it is built to answer

  • How do I tell a real eval from a benchmark my agent has memorised?
  • What does calibration mean for an LLM judge, and how is it measured?
  • Which failure modes does fault injection actually catch?

What an answer actually looks like

A real call against the live endpoint — answer with "how do I tell a real eval from benchmark contamination" — returns this structuredContent, trimmed:

{
  "answered": true,
  "entity": {
    "id": "agent-reliability-glossary",
    "name": "Agent reliability glossary",
    "evidence_tier": "secondary",
    "confidence": 0.85,
    "last_verified": "2026-08-08",
    "canonical_url": "https://agentreliability.dev/k/agent-reliability-glossary"
  },
  "claims": [
    {
      "text": "An eval is a structured, repeatable test that measures an LLM or LLM-based system against a defined dimension; frameworks package evals as registries of reusable templates.",
      "sources": [{ "title": "openai/evals — framework for evaluating LLMs and LLM systems" }]
    }
  ]
}

Note what travels with the answer: the evidence tier, a confidence, the date it was last verified, and the source behind the claim — not as prose an agent has to parse, but as fields it can act on. An agent can decline to use a weak claim, or cite the primary source directly.

When the corpus cannot answer, answered is false. It does not improvise, and the miss is recorded so the gap can be filled.

Machine-readable surfaces

The MCP endpoint is one of several. The same corpus is served as plain files an agent can read directly:

surface what it is
/llms.txt the index, as text/plain
/llms-full.txt the whole corpus in one file
/ai-index.json every surface this instance publishes, with its content type
/api/index.json one JSON document per knowledge object
/api/sources.json the source registry, in full
/.well-known/mcp/server.json this server's manifest

Each knowledge object has a human page and a machine twin at the same id, with a canonical URL that agrees across all of them.

Behaviour worth knowing before you integrate

  • POST only. Every other method answers 405 with an Allow: POST, OPTIONS header.
  • Rate limit: 120 requests per minute per client, counted in a shared store, published on every response as RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset (all three exposed via CORS). It fails open: if the store is unreachable the request is served.
  • Malformed input gets a spec-correct JSON-RPC error — -32700 for unparseable bodies, -32602 for an unknown tool — never an HTML error page.
  • Request bodies are capped and validated before transport.

Privacy

No accounts, no cookies, no ads. Usage is measured in aggregate with daily-rotating hashed identifiers and a 200-day retention; raw IPs are never stored. Full policy: PRIVACY.md.

Provenance and licence

Knowledge content is CC-BY-4.0: use it, cite it. The source registry is public precisely so a claim can be checked rather than trusted — get_sources returns what any given claim rests on.

Claims carry an evidence tier and a last_verified date. Where the evidence is weaker, the object says so rather than rounding up.

How it is built

Compiled and served by Citarium, an open-source framework for turning a knowledge graph into a website, an API, an MCP server and agent-readable files from a single source — under external evaluation, with the guardians and the falsification record in the open.

This repository is the server's public face: its manifest and its documentation. The corpus itself lives at agentreliability.dev.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选