vision-relay-mcp
A tiny MCP server that lets text-only coding models analyze images via vision relay APIs, supporting single image analysis and side-by-side comparison.
README
Vision Relay MCP 🖼️
A tiny MCP server that lets text-only coding models analyze images via Anthropic/OpenAI-compatible vision relay APIs.
一个轻量级 MCP 服务器,让不支持图片的编程模型也能借助视觉中继 API 分析截图、对比界面、解读报错。
English
What It Does
Vision Relay MCP solves a common pain point:
Your main coding model (DeepSeek, etc.) Vision-capable model (Claude / GPT)
│ │
│ "What's in screenshot.png?" │
│ ─────────────► │
│ vision-relay MCP │
│ ◄───────────── │
│ "The screenshot shows a │
│ NullPointerException at line 42..." │
You use a cheap or text-only model for coding. When you need image analysis, this MCP forwards the image to a vision-capable model through your relay API and returns the result transparently.
Features
- 🔌 Dual protocol support — Anthropic-compatible (
/v1/messages) and OpenAI-compatible (/v1/chat/completions) - 🖼️ Two practical tools —
analyze_imagefor single images,compare_imagesfor side-by-side comparison - 🔒 API keys stay safe — All credentials read from environment variables, never in source code
- 🪶 Minimal footprint — Single 267-line file, zero dependencies beyond MCP SDK, extremely auditable
- 🎯 Smart invocation — Claude Code can auto-invoke it when images are present; you can also explicitly say "use vision-relay" for manual control
Requirements
- Node.js 20 or newer (Node.js 22+ recommended)
- Claude Code or another MCP-compatible client
- A relay API key for a vision-capable model
⚠️ Note: The model you set in
VISION_MODELmust actually support vision input. A text-only model will not work even if the API connection succeeds.
Quick Start
# 1. Clone and install
git clone https://github.com/zhoucoolboy/vision-relay-mcp.git
cd vision-relay-mcp
npm install
Then choose one of the two methods below to register the MCP server with Claude Code.
Add to Claude Code
Method A: CLI (recommended)
Run this in the project directory:
claude mcp add -s user vision-relay `
-e VISION_PROVIDER=anthropic `
-e VISION_BASE_URL=https://your-relay.example.com `
-e VISION_MODEL=claude-sonnet-4-6 `
-e VISION_API_KEY=your_api_key_here `
-- node "%CD%\index.js"
Method B: Edit .claude.json directly
Open your user-level Claude Code config file:
notepad "$env:USERPROFILE\.claude.json"
Add the vision-relay entry under mcpServers:
{
"mcpServers": {
"vision-relay": {
"command": "node",
"args": ["D:\\software\\Desktop\\vision-relay-mcp\\index.js"],
"env": {
"VISION_PROVIDER": "anthropic",
"VISION_BASE_URL": "https://your-relay.example.com",
"VISION_MODEL": "claude-sonnet-4-6",
"VISION_API_KEY": "your_api_key_here"
}
}
}
}
Replace the
argspath with the actual path toindex.json your machine.
Verify
claude mcp list
claude mcp get vision-relay
You should see vision-relay with Connected status.
Note: After changing environment variables or
.claude.json, restart Claude Code for the changes to take effect.
Usage
Put an image in your project directory and just ask Claude Code about it — the MCP may be auto-invoked:
What's in screenshot.png?
What does this error say? (with error.png in the project folder)
You can also explicitly call it:
Please call vision-relay to analyze screenshot.png.
Please use vision-relay to compare before.png and after.png.
Supported Vision Models
This project does not limit which model you can use. Any vision-capable model accessible through your relay API works — as long as it supports image input.
The model is entirely determined by what you set in VISION_MODEL. For example:
- Anthropic format —
claude-sonnet-4-6/claude-opus-4-6/ any other Claude model that accepts images - OpenAI format —
gpt-4o/gpt-4o-mini/gemini-2.5-pro/glm-4v/glm-4.5v/ any other vision model your relay provides
Just set VISION_MODEL to whatever your relay supports — check your relay provider's model list for details.
Environment Variables Reference
| Variable | Required | Description |
|---|---|---|
VISION_PROVIDER |
Yes | anthropic or openai — selects the API format |
VISION_BASE_URL |
Yes | Your relay API base URL |
VISION_MODEL |
Yes | Model name for vision analysis |
VISION_API_KEY |
Yes | Your relay API key |
VISION_MAX_TOKENS |
No | Max tokens for the vision response (default: 2000) |
How to set these values depends on which method you chose above:
- Method A (CLI) — values are stored via
claude mcp add -e. Update them withclaude mcp removethenclaude mcp addagain. - Method B (
.claude.json) — values are in theenvblock. Edit the JSON file directly.
For OpenAI-compatible relays: set
VISION_PROVIDER=openaiand make sureVISION_BASE_URLends with/v1or/v1/chat/completions.
Switching Models
Method A (CLI): remove and re-add:
claude mcp remove vision-relay -s user
claude mcp add -s user vision-relay `
-e VISION_PROVIDER=anthropic `
-e VISION_BASE_URL=https://your-relay.example.com `
-e VISION_MODEL=claude-opus-4-6 `
-e VISION_API_KEY=your_api_key_here `
-- node "%CD%\index.js"
Method B (.claude.json): edit the VISION_MODEL value in the env block directly, then restart Claude Code.
Troubleshooting
| Symptom | Likely Cause |
|---|---|
Missing environment variable(s): VISION_API_KEY |
Environment variables not set. Restart Claude Code after setting them. |
Vision relay failed (403) |
Relay restricts the endpoint. Try switching VISION_PROVIDER, or check your relay docs. |
Could not process image |
Image too large (>20MB), corrupt, or unsupported format. Use PNG/JPEG under 20MB. |
| Text works but images fail | Model doesn't support vision. Confirm with your relay provider. |
Connected but analysis fails |
MCP connection ≠ API works. Double-check API key, base URL, and model name. |
Security Notes
- ⚠️ Never commit real API keys to source control
- ⚠️ Images are sent to your configured relay API — treat screenshots as sensitive data
- ⚠️ Rotate any API key that has been shared in chat, logs, screenshots, or public repos
- ⚠️ Prefer environment variables or a secret manager for credentials
Architecture
Claude Code (MCP Client)
│ stdio
▼
index.js (MCP Server)
│ HTTP POST (Bearer Token)
▼
Your Relay API
/v1/messages or /v1/chat/completions
│
▼
Vision model returns analysis
In a Nutshell
Claude Code writes your code.
vision-relay handles the images.
The vision model reads the images.
The relay API forwards the requests.
Most of the time, just put an image in your project and ask Claude Code about it — it will auto-invoke vision-relay. Try it out!
Uninstall
claude mcp remove vision-relay -s user
To also remove environment variables:
[Environment]::SetEnvironmentVariable("VISION_PROVIDER", $null, "User")
[Environment]::SetEnvironmentVariable("VISION_BASE_URL", $null, "User")
[Environment]::SetEnvironmentVariable("VISION_MODEL", $null, "User")
[Environment]::SetEnvironmentVariable("VISION_API_KEY", $null, "User")
License
MIT — see LICENSE.
简体中文
它能做什么
Vision Relay MCP 解决了一个普遍痛点:
你用 DeepSeek 之类的模型写代码 有视觉能力的模型(Claude / GPT)
│ │
│ "screenshot.png 里有什么?" │
│ ─────────────► │
│ vision-relay MCP │
│ ◄───────────── │
│ "截图显示第 42 行有 │
│ NullPointerException..." │
主力编程模型不支持图片?这个 MCP 把图片转发给视觉模型分析,结果透明返回给 Claude Code。
特性
- 🔌 双协议支持 — 兼容 Anthropic 格式 (
/v1/messages) 和 OpenAI 格式 (/v1/chat/completions) - 🖼️ 两个实用工具 —
analyze_image分析单张图片,compare_images对比两张图片 - 🔒 密钥不进代码 — 全部走环境变量,仓库里不存任何密钥
- 🪶 极致简洁 — 单文件 267 行,仅依赖 MCP SDK,极端可审计
- 🎯 智能调用 — Claude Code 遇到图片时可自动调用,你也可以手动说"调用 vision-relay"来精准控制
你需要准备什么
1. Node.js(20 以上,推荐 22+)
2. Claude Code 或其他 MCP 客户端
3. 一个支持图片输入的中转站 API
4. 一个支持视觉的模型名
⚠️ 注意:普通文本模型不行。模型必须真的支持图片输入。
快速开始
# 1. 克隆并安装
git clone https://github.com/zhoucoolboy/vision-relay-mcp.git
cd vision-relay-mcp
npm install
然后从下面两种注册方式中选一种。
注册到 Claude Code
方式 A:CLI 命令(推荐)
在项目目录下执行:
claude mcp add -s user vision-relay `
-e VISION_PROVIDER=anthropic `
-e VISION_BASE_URL=https://你的中转站地址 `
-e VISION_MODEL=你的视觉模型名 `
-e VISION_API_KEY=你的API密钥 `
-- node "%CD%\index.js"
方式 B:直接编辑 .claude.json
打开用户级配置文件:
notepad "$env:USERPROFILE\.claude.json"
在 mcpServers 下添加 vision-relay:
{
"mcpServers": {
"vision-relay": {
"command": "node",
"args": ["D:\\software\\Desktop\\vision-relay-mcp\\index.js"],
"env": {
"VISION_PROVIDER": "anthropic",
"VISION_BASE_URL": "https://你的中转站地址",
"VISION_MODEL": "你的视觉模型名",
"VISION_API_KEY": "你的API密钥"
}
}
}
}
把
args里的路径换成你电脑上index.js的实际路径。
验证
claude mcp list
claude mcp get vision-relay
看到 vision-relay ... Connected 就说明成功了。
注意: 修改环境变量或
.claude.json之后,需要重启 Claude Code 才能生效。
使用方式
把图片放到项目目录里,直接用自然语言问 Claude Code 就行——MCP 可能会自动被调用:
screenshot.png 里有什么?
这个报错是什么意思?(项目里有 error.png)
你也可以显式调用:
请调用 vision-relay 分析 screenshot.png
请用 vision-relay 看一下 error.png,告诉我报错的原因
请调用 vision-relay 对比 before.png 和 after.png
支持的视觉模型
本项目不限制你用什么模型。只要是支持图片输入的视觉模型,且你的中转站能访问到,就可以用。
模型完全由你设置的 VISION_MODEL 决定。举个例子:
- Anthropic 格式 —
claude-sonnet-4-6/claude-opus-4-6/ 任何支持图片的 Claude 模型 - OpenAI 格式 —
gpt-4o/gpt-4o-mini/gemini-2.5-pro/glm-4v/glm-4.5v/ 你的中转站提供的任何其他视觉模型
把你中转站支持的视觉模型名填到 VISION_MODEL 就行,具体去你的中转站后台看支持列表。
环境变量参考
| 变量 | 必填 | 说明 |
|---|---|---|
VISION_PROVIDER |
是 | anthropic 或 openai,选择 API 格式 |
VISION_BASE_URL |
是 | 你的中转站地址 |
VISION_MODEL |
是 | 视觉模型名称 |
VISION_API_KEY |
是 | 中转站 API 密钥 |
VISION_MAX_TOKENS |
否 | 视觉模型最大返回 token 数(默认 2000) |
这些值的设置方式取决于你上面选的是哪种注册方式:
- 方式 A(CLI) — 参数通过
claude mcp add -e存储。想修改就先claude mcp remove再重新claude mcp add。 - 方式 B(
.claude.json) — 参数在env块里,直接编辑 JSON 文件即可。
OpenAI 格式中转站: 把
VISION_PROVIDER设为openai,VISION_BASE_URL通常以/v1或/v1/chat/completions结尾。
换模型
方式 A(CLI): 先删除再重新添加:
claude mcp remove vision-relay -s user
claude mcp add -s user vision-relay `
-e VISION_PROVIDER=anthropic `
-e VISION_BASE_URL=https://你的中转站地址 `
-e VISION_MODEL=claude-opus-4-6 `
-e VISION_API_KEY=你的API密钥 `
-- node "%CD%\index.js"
方式 B(.claude.json): 直接编辑 env 块里的 VISION_MODEL 值,保存后重启 Claude Code 即可。
换中转站
方式 A(CLI): 删掉重新添加,换掉 VISION_BASE_URL 或 VISION_API_KEY。
方式 B(.claude.json): 直接编辑 env 块里的对应字段,保存后重启 Claude Code。
常见问题
| 现象 | 可能原因 |
|---|---|
Missing environment variable(s): VISION_API_KEY |
环境变量没设。设完记得重启 Claude Code。 |
Vision relay failed (403) |
中转站限制端点或客户端类型。试试切换 VISION_PROVIDER。 |
Could not process image |
图片太大(>20MB)、损坏或格式不支持。用 PNG/JPEG 且小于 20MB。 |
| 文字能返回但图片分析失败 | 模型不支持视觉。跟中转站确认该模型是否支持图片输入。 |
| MCP 显示 Connected 但识图失败 | Connected 只代表 Claude Code 能启动 MCP。真正识图还需要 API key 正确、模型支持图片、账号有额度。 |
安全提醒
- ⚠️ 绝对不要把 API key 写进
index.js/README.md/.env.example或提交到 GitHub - ⚠️ 图片会发送到你配置的中转站,注意截图中的敏感信息
- ⚠️ 密钥一旦泄露(聊天记录、截图、公开仓库),请立即轮换
- ⚠️ 推荐放在 Windows 环境变量中
架构
Claude Code(MCP 客户端)
│ stdio
▼
index.js(MCP 服务器)
│ HTTP POST(Bearer Token)
▼
你的中继 API
/v1/messages 或 /v1/chat/completions
│
▼
视觉模型返回分析结果
一句话理解
Claude Code 负责写代码
vision-relay 负责接收图片
视觉模型负责看图
中转站负责转发 API 请求
大多数情况下,直接把图片放项目里问 Claude Code 就行,它会自动调用 vision-relay。欢迎验证!
卸载
claude mcp remove vision-relay -s user
如果要同时删除环境变量:
[Environment]::SetEnvironmentVariable("VISION_PROVIDER", $null, "User")
[Environment]::SetEnvironmentVariable("VISION_BASE_URL", $null, "User")
[Environment]::SetEnvironmentVariable("VISION_MODEL", $null, "User")
[Environment]::SetEnvironmentVariable("VISION_API_KEY", $null, "User")
许可证
MIT — 详见 LICENSE。
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。