seo-geo-mcp-server

seo-geo-mcp-server

An MCP server that lets an AI agent audit a page for SEO and GEO — on-page tags, structured data, robots.txt, sitemaps, hreflang, and whether ChatGPT, Claude, Perplexity and Gemini can actually crawl and cite you. No API keys required.

Category
访问服务器

README

seo-geo-mcp-server

An MCP server that lets an AI agent audit a page for SEO and GEO (Generative Engine Optimization) — on-page tags, structured data, robots.txt, sitemaps, hreflang, and whether ChatGPT, Claude, Perplexity and Gemini can actually crawl and cite you. No API keys required.

MCP TypeScript License: MIT

Ask Claude "how is this page doing, and will AI assistants cite it?" and it runs a full audit and hands you a graded report with prioritised fixes — instead of you pasting a URL into six different web tools.

> Audit https://example.com/guide and tell me if ChatGPT can cite it

  seo_audit(url="https://example.com/guide", include_geo=true)

  Overall: A (92/100) · indexable: yes
    Meta tags & social preview   95 (A)
    Heading structure           100 (A)
    Structured data              75 (C)

  GEO readiness: B (85/100)
    ✅ AI crawler access        25/25
    ✅ Server-rendered content  20/20
    ❌ Authorship & entity       2/10
  1. Add author and Organization markup with `sameAs` links to official profiles.

Why this exists

Two gaps, one server.

The SEO gap: the on-page checkers are all web UIs. None of them let an agent run the audit, read the result and fix the code in the same loop.

The GEO gap: "generative engine optimization" tooling is mostly rank-tracking dashboards behind a subscription. The mechanics that actually decide whether an AI assistant can cite you are cheap to check and almost never checked:

  • Can the AI search crawlers reach you? Blocking GPTBot stops training. Blocking OAI-SearchBot stops you being cited. Most sites that meant to do the first have accidentally done the second. This server separates them.
  • Does your content exist without JavaScript? Googlebot renders JS. GPTBot, ClaudeBot, PerplexityBot and CCBot largely do not. A client-rendered page can rank fine in Google and be invisible to every AI assistant.

It is the agent-facing companion to the tools at ortamarco.me, and shares its core (SSRF-guarded fetching, public-resolver DNS, host validation) with domain-security-mcp-server.

Tools

Audits

Tool What it does
seo_audit One fetch → seven weighted sections (meta, headings, content, schema, images, links, crawlability) → 0–100 score, A–F grade, prioritised fixes. include_geo=true adds the GEO dimension
geo_audit AI answer-engine readiness: crawler access (25), server-rendered content (20), structured data (15), extractable structure (15), authorship (10), freshness (8), depth (7)

GEO

Tool What it does
ai_crawler_access Resolves ~35 AI crawler tokens against robots.txt. Separates training from citation bots, flags blocks that vendors document as unenforceable, handles the Applebot→Googlebot fallback
render_check Whether content survives without JavaScript — detects unhydrated SPA shells that AI crawlers cannot read
llms_txt_check Detects and validates /llms.txt against the llmstxt.org proposal — and reports its real adoption status rather than implying it earns visibility

On-page

Tool What it does
meta_tags_check Title, description, canonical, robots (meta and X-Robots-Tag), lang, charset, viewport
social_preview_check Open Graph + Twitter Card, and verifies the og:image actually loads
heading_structure Full h1–h6 outline, multiple h1s, skipped levels, question-shaped headings
structured_data_check JSON-LD/microdata/RDFa extraction, parse errors, and Google rich-result requirements for 17 schema types
content_analysis Word count, Flesch reading ease, thin-content detection, text-to-HTML ratio, term density (EN + ES stopwords)
image_seo_check Missing alt text, missing dimensions (layout shift), lazy loading, WebP/AVIF adoption

Technical

Tool What it does
robots_txt_check RFC 9309 parse; flags wildcard Disallow: / and 5xx responses (which Google reads as disallow-all)
sitemap_check Discovery via robots.txt → conventional paths; index following, gzip, 50k/50MiB limits, lastmod validity
canonical_host_check All four http/https × apex/www variants — do they converge on one canonical URL, and via 301 or 302?
redirect_trace Hop-by-hop chain with loop and temporary-redirect detection
link_audit Internal/external split, rel attributes, generic anchor text, optional broken-link sampling
hreflang_check BCP-47 validity, self-reference, x-default, duplicates — plus optional reciprocity verification

Every tool is read-only, declares an outputSchema and returns structuredContent (validated by the SDK) alongside human-readable Markdown (default) or JSON (response_format="json"), plus actionable error messages.

Honesty notes

This server deliberately refuses to overstate two things that most GEO content gets wrong. Both are surfaced in tool output, not buried here:

  • llms.txt is not an adopted standard. It is a community proposal from September 2024. No major AI vendor has documented that its crawlers read it from third-party sites, and Google has publicly said it does not. The tool reports presence and validates shape — and geo_audit deliberately does not score it. (llms-full.txt is a docs-tooling convention, not part of the proposal.)
  • Some robots.txt blocks are advisory. Perplexity-User, ChatGPT-User and meta-externalfetcher are documented by their own vendors as ignoring or possibly ignoring robots.txt. Reporting those as cleanly "blocked" would be misleading, so they are listed separately as unenforceable.

Crawler tokens carry a provenance field distinguishing first-party vendor documentation from community aggregators, and vendors that publish no token at all (xAI/Grok, Microsoft Copilot) are named explicitly — because a missing rule cannot be read as either allowed or blocked.

Install

git clone https://github.com/OrtaMarco/seo-geo-mcp-server.git
cd seo-geo-mcp-server
npm install
npm run build

Use it with Claude Code

claude mcp add seo-geo -- node /absolute/path/to/seo-geo-mcp-server/dist/index.js

Use it with Claude Desktop

Add to claude_desktop_config.json (see examples/):

{
  "mcpServers": {
    "seo-geo": {
      "command": "node",
      "args": ["/absolute/path/to/seo-geo-mcp-server/dist/index.js"]
    }
  }
}

Restart Claude Desktop, then ask: "Audit the SEO and GEO of example.com."

Self-host (HTTP transport)

The same server speaks stateless Streamable HTTP for remote/multi-client use — handy behind a reverse proxy such as Coolify or Traefik.

TRANSPORT=http PORT=3000 npm start
# POST JSON-RPC to http://localhost:3000/mcp   ·   health at /healthz

Or with Docker:

docker build -t seo-geo-mcp .
docker run -p 3000:3000 -e TRANSPORT=http seo-geo-mcp

Set ALLOWED_ORIGINS=https://your.app to enable Origin-based DNS-rebinding protection (leave empty when a trusted proxy already restricts access).

Develop

npm run dev      # tsx watch (stdio)
npm test         # 32 deterministic unit tests (robots matcher, SPA detection, JSON-LD…)
npm run smoke    # call all 17 tools over MCP and validate structuredContent vs outputSchema
npm run inspect  # open the MCP Inspector against the built server
npm run build    # type-check + emit dist/

evals/ holds a 10-question LLM evaluation set (stable, verifiable) and instructions for running it — see evals/README.md.

How it works

src/
├── index.ts        # transport selection (stdio | http)
├── server.ts       # registers every tool on one McpServer
├── schemas.ts      # Zod outputSchema for each tool
├── core/           # pure logic, no MCP coupling — reusable & testable
│   ├── fetch.ts        # SSRF-safe fetch: per-hop guard, byte caps, manual redirects
│   ├── page.ts         # HTML loading + the shared parsed-document model
│   ├── meta.ts         # title/description/canonical/robots, Open Graph, hreflang
│   ├── content.ts      # headings, readability, word counts, image SEO
│   ├── structured-data.ts # JSON-LD/microdata + Google rich-result requirements
│   ├── robots.ts       # RFC 9309 parser and rule matcher
│   ├── ai-crawlers.ts  # the AI crawler registry (token, purpose, compliance, provenance)
│   ├── sitemap.ts      # discovery, index following, gzip, protocol limits
│   ├── links.ts        # link classification + broken-link sampling
│   ├── redirects.ts    # chain tracing + host canonicalisation
│   ├── geo.ts          # llms.txt, JS-rendering detection, GEO scoring
│   └── seo-audit.ts    # the composite audits (one fetch, every analyser)
└── tools/          # thin MCP wrappers (Zod schemas, descriptions, formatting)

The core/ layer is deliberately free of any MCP types, so the same logic can power both this server and a web UI.

Security: every user-supplied URL is validated and re-checked on each redirect hop against loopback, private, link-local and cloud-metadata ranges, so the server cannot be used to probe internal networks. Response bodies are read through a byte cap.

License

MIT © Marco Orta

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选