ARWP Resolver MCP

ARWP Resolver MCP

Enables MCP clients to discover the concrete interfaces a website actually exposes—llms.txt, OpenAPI, Agent Skills, WebMCP, MCP, A2A—and select the right one for a given intent without site-specific code.

Category
访问服务器

README

Agent-Ready Web Profile

ARWP validation Reference verification

Resolve how a website can actually be used by agents.

Modern websites may expose HTML, Markdown negotiation, HTTP Link discovery, llms.txt, datasets, retrieval indexes, OpenAPI, Agent Skills, MCP, A2A, OAuth resource metadata, agents.json and other surfaces. A client should not need site-specific code — or guess which manifest is authoritative — to understand them.

ARWP now has two complementary parts:

  1. ARWP Profile — an experimental publisher-maintained service map at /ai/site-profile.json.
  2. ARWP Resolver — an interoperability engine that reads ARWP plus existing upstream/community/web discovery, preserves evidence and conflicts, and selects an interface for a concrete intent.

The profile is useful. It is not required to use the Resolver and it is not intended to replace upstream standards.

Public project site: https://dkharlanau.github.io/agent-ready-web-profile/

Profile contract: experimental v0.1. Released validator/Action: v0.1.0. The current main toolchain is version 0.2.0; npm publication remains an external release gate and must not be described as complete until it succeeds.

The problem

A site can legitimately publish several independent discovery surfaces:

                         WEBSITE
                            |
      +---------------------+----------------------+
      |          |          |         |            |
    HTTP       ARWP      agents.*   API/A2A    Agent Skills
 Link/HTML    profile                metadata
      |                     |                      |
      +-------------- MCP / OAuth / web ---------+
                            |
                       ARWP RESOLVER
                            |
                 evidence-backed service map
                            |
               +------------+-------------+
               |            |             |
             read         search         tools
             data       structured       agent

The Resolver does not ask every ecosystem to converge on one file. It answers:

What does this website actually expose, where did each claim come from, do the claims conflict, and which interface should a client use for this task?

Resolve, explain and plan

node bin/arwp.mjs resolve https://example.com
node bin/arwp.mjs explain https://example.com
node bin/arwp.mjs plan https://example.com --intent=search

Machine-readable output is available with --json.

Supported planning intents:

  • read
  • search
  • structured
  • tools
  • agent

Planning is deterministic. It preserves source authority and fallbacks rather than hiding decisions behind a readiness score.

For read, richer publisher surfaces such as llms.txt or Markdown can be preferred, while ordinary canonical HTML remains an honest low-priority fallback. A plain website is therefore not treated as unusable merely because it has no AI-specific metadata.

See docs/RESOLVER.md.

Discovery surfaces

The Resolver currently normalizes evidence from:

  • canonical HTML and bounded ordinary web discovery;
  • HTTP Link relations including api-catalog, service-desc, service-doc, Markdown alternate and ARWP describedby;
  • Accept: text/markdown content-negotiation observation;
  • valid ARWP profiles;
  • /agents.txt and /agents.json as a community convention;
  • RFC 9727 API Catalog;
  • RFC 9728 root Protected Resource Metadata;
  • A2A /.well-known/agent-card.json;
  • Agent Skills /.well-known/agent-skills/index.json;
  • experimental MCP AI Catalog / Server Card discovery.

Experimental/community sources remain explicitly labeled. Static metadata is never silently upgraded into runtime conformance.

Source authority stays visible

Authority Example
ietf-standard RFC 8288 / RFC 9727 / RFC 9728
upstream-standard A2A Agent Card
upstream-convention Agent Skills discovery
community-convention agents.txt / agents.json
experimental-upstream current MCP Server Card / AI Catalog work
project-profile ARWP publisher profile
observed-web directly observed HTML/HTTP evidence

Authority is not authorization or a security trust rank.

Batch resolution

Inventory/research workflows can resolve several sites without site-specific wrappers:

node bin/arwp.mjs resolve-many targets.txt
node bin/arwp.mjs resolve-many targets.json --concurrency=4 --json

The library primitive is bounded to 100 targets and max concurrency 10. Same-origin work is serialized inside a batch and failures are isolated per site.

Snapshots and drift

Create compact operational state:

node bin/arwp.mjs snapshot https://example.com --output=example.snapshot.json

Compare two observations:

node bin/arwp.mjs drift before.snapshot.json after.snapshot.json --json

Snapshots keep identity, discovery sources, normalized interfaces, conflicts and deterministic intent plans. They do not copy canonical datasets.

Drift distinguishes added/removed/changed sources, interfaces, conflicts, identity and routing-plan changes. Observation time alone is not drift.

Resolver monitoring

A small monitor runtime builds on the same snapshots:

cp monitor/example.config.json arwp-resolver-monitor.json
npm run monitor:resolver

A monitor may fail only on selected operational classes:

  • identity
  • source-removed
  • interface-removed
  • conflict-added
  • plan-changed
  • resolution-failed
  • any

templates/github-actions/resolver-monitor.yml provides a scheduled workflow with cached operational snapshots and always-uploaded drift reports.

Runtime evidence is opt-in

A normal resolve remains static/bounded discovery and does not open MCP sessions or perform cryptographic trust checks.

The Resolver MCP exposes explicit verification tools when deeper evidence is wanted.

MCP runtime reconciliation

verify_mcp_runtime:

  • sends real modern server/discover where supported;
  • falls back to legacy initialize + notifications/initialized lifecycle;
  • records negotiated/self-reported server metadata;
  • reports authorization-required separately from runtime failure;
  • blocks cross-origin runtime redirects;
  • never invokes MCP tools;
  • never sends credentials discovered from metadata automatically;
  • surfaces static/runtime identity mismatches as conflicts.

A2A signature verification

verify_a2a_signatures:

  • validates the current v1 Agent Card required shape;
  • treats unsigned cards as unsigned, not invalid;
  • retrieves explicitly declared public HTTPS JWKS under the same network bounds;
  • verifies RS256 and ES256 signatures by kid;
  • distinguishes signature-verified, signature-invalid, key-unavailable, unsupported-algorithm, invalid-card and not-assessed.

Internal RSA/EC fixtures and tampering detection pass CI. Broad cross-SDK interoperability is still an explicit external gate because current A2A implementations have had canonicalization/default-field inconsistencies. A cryptographically valid signature also does not by itself make a signer trustworthy.

Resolver as MCP

npm run resolver:mcp

Current tools:

  • resolve_site
  • resolve_sites
  • search_resolved_sites
  • explain_site
  • plan_site_interface
  • verify_mcp_runtime
  • verify_a2a_signatures

The prepared Official MCP Registry artifact launches this Resolver and does not require an ARWP profile.

The older ARWP-profile gateway remains available separately:

ARWP_PROFILE=https://example.com/ai/site-profile.json npm run mcp:start

Remote Streamable HTTP profile gateway:

ARWP_PROFILE=https://example.com/ai/site-profile.json \
ARWP_HTTP_ALLOWED_HOSTS=mcp.example.com \
npm run mcp:http

See docs/GATEWAY.md.

Resolver-backed federation

The original directory federation remains available:

node bin/arwp.mjs directory
node bin/arwp.mjs federated-search "outside view"

The newer Resolver MCP search_resolved_sites starts from canonical site URLs. It does not require ARWP profiles.

Generic federation deliberately executes only resolved static JSON/JSONL/NDJSON retrieval indexes. It does not invent OpenAPI, MCP or A2A calls when operation semantics are unknown. Each result preserves source site, discovery source/authority and selected interface.

Public ARWP reference directory:

https://dkharlanau.github.io/agent-ready-web-profile/directory.json

See docs/DIRECTORY.md.

ARWP Profile

Publishers that want one explicit service map can expose:

/ai/site-profile.json

Minimal profile:

{
  "$schema": "https://raw.githubusercontent.com/dkharlanau/agent-ready-web-profile/v0.1.0/schema/site-profile.schema.json",
  "profileVersion": "0.1",
  "id": "example-knowledge-site",
  "name": "Example Knowledge Site",
  "canonicalUrl": "https://example.com/",
  "description": "A reviewed public knowledge library.",
  "web": {
    "sitemap": "https://example.com/sitemap.xml",
    "llms": "https://example.com/llms.txt"
  }
}

Optional HTML advertisement:

<link rel="describedby" type="application/json" href="/ai/site-profile.json">

This is an ARWP convention, not a registered .well-known location. See SPEC.md.

Adopt ARWP from an existing website

node bin/arwp.mjs scan https://example.com
node bin/arwp.mjs init https://example.com
node bin/arwp.mjs validate ai/site-profile.json
node bin/arwp.mjs verify https://example.com/ai/site-profile.json
node bin/arwp.mjs health https://example.com

scan observes bounded public evidence. init generates a conservative profile and does not invent unverified MCP, WebMCP, Skills or A2A capabilities.

Reusable Action:

- name: Validate Agent-Ready Web Profile
  uses: dkharlanau/agent-ready-web-profile@v0.1.0
  with:
    profile: ai/site-profile.json

templates/github-actions/propose-arwp-profile.yml provides an opt-in profile-update PR workflow.

Bounded hosted discovery service

The server runtime exposes only fixed operations:

GET  /health
POST /scan
POST /resolve
POST /explain
POST /plan

It includes HTTPS-only target rules, DNS/private-network rejection, redirect revalidation, request/response bounds, explicit browser Origin allow-listing and shared rate limiting. It is not an arbitrary URL proxy.

ARWP_SCANNER_ALLOWED_ORIGINS=https://dkharlanau.github.io \
npm run scanner:http

A container artifact is in scanner-service/. Public hosting remains an external deployment gate.

Real reference suite

Five owner-controlled public knowledge-site architectures publish ARWP profiles and are live-verified:

  • Dzmitryi Kharlanau — SAP Knowledge;
  • Brali Practical Knowledge Library;
  • Cognitive Biases Knowledge Library;
  • CBT Cards;
  • Metkagram.

They are implementation/regression evidence, not independent adoption evidence.

Benchmark before marketing claims

Synthetic regression:

npm run benchmark:resolver

Independent-corpus runner:

npm run benchmark:external -- --output=benchmark-results/external.json

The external runner has a strict reviewed fixture schema. Aggregate results count only ownership=independent. Ground truth is manually reviewed public evidence and cannot be generated from Resolver output itself.

Subset strategy comparisons are selection-only projections over the same observed resolution. Request/byte/time metrics are attributed only to the actual Resolver network run.

The first pilot corpus contains 10 independent documentation sites and deliberately includes ordinary HTML controls and path-scoped discovery that the Resolver may miss. The target after reviewing the pilot is 20–50 sites.

No benchmark result is evidence of token savings, search ranking, adoption or answer quality. Raw negative results must remain visible. See docs/BENCHMARK.md.

What ARWP deliberately does not replace

Do not create ARWP-native replacements for:

  • RFC 8288 Web Linking;
  • RFC 9727 API Catalog;
  • RFC 9728 Protected Resource Metadata;
  • A2A Agent Cards;
  • MCP runtime discovery / Server Cards;
  • Agent Skills;
  • crawler AI-use preferences;
  • payment/commerce protocols.

Project rule:

UPSTREAM EXISTS
      ↓
resolve / verify / normalize it

UPSTREAM DOES NOT EXIST
      ↓
collect a concrete interoperability failure

ONLY THEN
      ↓
consider an ARWP-specific extension

Security boundaries

  • public HTTPS targets only;
  • private/reserved/link-local targets rejected;
  • redirect destinations revalidated;
  • bounded requests and responses;
  • no URL credentials;
  • generic federation does not invent operations;
  • metadata never grants permission;
  • static reachability never proves runtime conformance;
  • runtime probes are opt-in and do not invoke MCP tools;
  • signature verification does not establish signer trust;
  • conflicts remain visible instead of being hidden by a score.

Development direction

The North Star is:

How many external sites can ARWP correctly resolve and route without site-specific integration code?

Immediate work is increasingly external/evidence-driven:

  1. preserve and review the first independent benchmark pilot;
  2. fix systematic discovery gaps only after the baseline is recorded;
  3. expand the corpus to 20–50 sites;
  4. publish/install the 0.2.x Resolver package and MCP Registry artifact;
  5. deploy the bounded public HTTPS discovery service;
  6. obtain three independent adopters/consumers;
  7. prove A2A signature interoperability against independent upstream implementations;
  8. decide from evidence whether the ARWP Profile contract needs another version at all.

See ROADMAP.md.

Repository map

License

Apache License 2.0. See LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选