mcp-eyes
A drop-in MCP server that pairs long-context reasoning LLMs with vision models in description-only mode, enabling any reasoning model to 'see' images without the vision model giving advice or solutions.
README
vision-extension
Drop-in vision capability pack for text-only reasoning LLMs. One repo containing an MCP server and a Claude Code skill, both engineered around a single contract: the vision model only describes — your reasoning model does the thinking.
English
What's in this repo
| Directory | What it is | Who installs it |
|---|---|---|
mcp-vision-extension/ |
The MCP server (Python package vision_extension). Pairs any text-only reasoning model with any vision model over OpenAI or Anthropic protocol. |
Required. Install once per machine. |
skills-vision-extension/ |
A Claude Code skill that knows the install playbook AND the day-to-day collaboration patterns between the reasoning model and the vision model. | Optional but strongly recommended for Claude Code users. Copy into ~/.claude/skills/. |
These two pieces are designed to work together. The MCP server gives your text model vision; the skill teaches your text model how to use that vision well.
Install everything in 2 commands
# 1. The MCP server
pip install "git+https://github.com/loudMore/vision-extension.git#subdirectory=mcp-vision-extension"
# 2. The Claude Code skill (optional)
git clone https://github.com/loudMore/vision-extension.git /tmp/vx
cp -r /tmp/vx/skills-vision-extension/vision-extension ~/.claude/skills/
Then point your MCP client at the new server. Detailed steps + provider presets in mcp-vision-extension/README.md.
Or just tell your agent
If you have Claude Code (or any MCP-aware agent), copy the skill once:
git clone https://github.com/loudMore/vision-extension.git
cp -r vision-extension/skills-vision-extension/vision-extension ~/.claude/skills/
Then say:
"Install vision-extension. Use the
<doubao | openai | qwen | gemini | ollama | …>provider. Here's my key:<KEY>."
The skill handles the rest. You don't write any JSON.
Why this exists
Long-context reasoning models (DeepSeek V4 Pro, GLM 5.2, Kimi K2, Qwen 3 Max, …) are extraordinary at code and analysis but cannot see images. Naively bolting on a vision API has two problems:
- No standard pipe — every IDE wires it differently.
- Vision models love to "help" — GPT-4o, Gemini, Doubao all reflexively produce advice, debugging hypotheses, and design opinions when you only wanted a description. The reasoning work gets fragmented.
vision-extension solves both:
- One MCP server, works with Claude Code, Cursor, Continue, Cline, Roo, or anything else that speaks MCP.
- Describe-only contract — the vision model is system-prompted into a pure visual scanner. No advice. No fixes. No opinions. Just verbatim transcription and structured description.
- One Claude Code skill that turns the install + daily-use rules into a single trigger phrase.
- Provider-agnostic — Anthropic protocol, OpenAI-compatible protocol. Switch with one env var.
License
MIT.
中文
仓库里有什么
| 目录 | 是什么 | 谁要装 |
|---|---|---|
mcp-vision-extension/ |
MCP server(Python 包 vision_extension),把任意纯文本推理模型和任意视觉模型用 OpenAI/Anthropic 协议接到一起 |
必装,每台机器装一次 |
skills-vision-extension/ |
Claude Code skill,把安装流程 + 主模型与视觉模型的日常协作规则打包好 | 强烈推荐,复制到 ~/.claude/skills/ 即可 |
两块组件协同设计。MCP server 给文本模型装上视觉;skill 教文本模型怎么用好这套视觉。
两条命令搞定
# 1. 装 MCP server
pip install "git+https://github.com/loudMore/vision-extension.git#subdirectory=mcp-vision-extension"
# 2. 装 Claude Code skill(可选)
git clone https://github.com/loudMore/vision-extension.git /tmp/vx
cp -r /tmp/vx/skills-vision-extension/vision-extension ~/.claude/skills/
然后让你的 MCP 客户端配新 server。详细步骤和 12 个 provider 预设见 mcp-vision-extension/README.md。
或者直接让你的 agent 装
装完 skill 之后,对你的 Claude Code(或任何支持 MCP 的 agent)说:
"装个 vision-extension。视觉模型用
<豆包 | openai | 通义 | 智谱 | ollama | …>,key 是<KEY>。"
skill 会按 7 步确定流程把剩下的全做完。你不用写任何 JSON。
为什么做这个
DeepSeek V4 Pro / GLM 5.2 / Kimi K2 / Qwen 3 Max 这类长上下文推理模型推理超强,但看不见图。直接接个视觉 API 拼起来有两个老问题:
- 没有统一通道 —— 每个 IDE 接法都不一样
- 视觉模型爱"帮忙" —— GPT-4o / Gemini / 豆包都会条件反射地给方案、提假设、写评价,把推理工作抢走一半,你只想要个描述
vision-extension 一并解决:
- 一个 MCP server,Claude Code / Cursor / Continue / Cline / Roo 通用
- describe-only 契约 —— 视觉模型被系统提示锁成纯扫描器,不给建议、不给方案、不给评价,只做逐字转录和结构化描述
- 一个 Claude Code skill 把安装流程和日常使用规则压成一句话触发
- 协议解耦 —— Anthropic 协议、OpenAI 协议都支持,一个环境变量切换
License
MIT。
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。