vision-mcp
Enables text-only models to understand images through a conversational MCP server, supporting multi-turn follow-ups, URL inputs, and OpenAI-compatible vision APIs.
README
vision-mcp
An MCP server that gives conversational image understanding to text-only models (DeepSeek, Claude Code, etc.).
The main model hands image paths and a question to a vision model, which returns a text description that the main model can reason over. A session mechanism enables "blind men and the elephant" style follow-ups: you can dig deeper into the same batch of images across multiple turns, and the vision model re-sees the full images and the conversation history on every turn.
Features
- Multi-turn conversational follow-up: follow-ups within a session automatically carry history context, supporting referential questions ("What does that sign say?")
- OpenAI-compatible vision API: any compatible endpoint works (default SiliconFlow; Qwen3.5-35B-A3B verified to accept images)
- Path deduplication: paths passed on session reuse are compared against existing ones; only new images are added
- URL support: pass http/https image URLs directly; they are forwarded to the vision API as-is (no local download)
- Image integrity validation: checks extension vs. actual format consistency; supports png/jpg/jpeg/webp/gif/bmp/tif/tiff
- Concurrency safe: operations on the same session are serialized; atomic writes; deletion is mutually exclusive with in-flight requests, so a deleted session can never be resurrected by a stale save
- Auto-expiry: sessions idle for 24 hours are cleaned up (configurable)
- Passthrough by default: images are sent as-is — no compression, no scaling, no re-encoding — unless compression is enabled (see
max_image_mb) - Staged compression (optional): when enabled, images over the threshold are compressed in stages, format-preserving where possible (JPEG/WebP lower quality first, then downscale; PNG keeps transparency by downscaling before falling back to JPEG)
Installation
git clone https://github.com/whyneedai/vision-mcp.git
cd vision-mcp
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
Configuration
The config file lives at ~/.config/vision-mcp/config.json:
{
"vision_model": {
"base_url": "https://api.siliconflow.cn/v1",
"api_key": "{env:SILICONFLOW_API_KEY}",
"model": "Qwen/Qwen3.5-35B-A3B",
"enable_thinking": false,
"max_tokens": 131072,
"temperature": 0.1
},
"max_history_rounds": 4,
"sessions_dir": "~/.local/share/vision-mcp/sessions",
"session_ttl_hours": 24,
"system_prompt": "optional, overrides the built-in vision system prompt"
}
| Field | Description |
|---|---|
max_history_rounds |
Number of recent Q&A rounds (1 round = one question + one answer) sent to the vision model as context. History beyond the window is kept on disk but not sent. Default 4; 0 sends no history. Capped by the provider's message limit (10) |
vision_model.base_url / api_key / model |
OpenAI-compatible endpoint; {env:XXX} references an environment variable |
vision_model.enable_thinking |
Disable thinking mode (otherwise the API returns an empty content) |
vision_model.max_image_mb |
Compression threshold in MB: images at or above this size are auto-compressed below it (staged, format-preserving). Unset / empty / 0 disables compression (default) |
sessions_dir |
Session storage directory (default ~/.local/share/vision-mcp/sessions) |
session_ttl_hours |
Session idle-expiry in hours (default 24) |
system_prompt |
Vision system prompt (default: strictly follow the question, no hallucination) |
Connecting to opencode
Add to the mcp section of your opencode config:
{
"mcp": {
"vision": {
"type": "local",
"command": ["/path/to/vision-mcp/.venv/bin/python", "/path/to/vision-mcp/server.py"],
"environment": {
"SILICONFLOW_API_KEY": "{env:SILICONFLOW_API_KEY}"
},
"enabled": true,
"timeout": 300000
}
}
}
Tools
ask_image
Ask a question about one or more images, with multi-turn follow-up support.
| Parameter | Required | Description |
|---|---|---|
question |
yes | The question |
session_id |
no | Existing session ID; omit to create a new session (image_path is then required) |
image_path |
no | List of image paths or http/https image URLs; may be omitted on session reuse (existing images are kept), new entries are deduplicated and appended. URLs are forwarded as-is to the vision API |
Returns {session_id, answer, image_paths}. Relative paths resolve against the opencode working directory.
end_session
Delete a session and all of its related files. Original images are never deleted.
Architecture
server.py MCP entry point: component wiring + tool registration
config.py Config loading ({env:XXX} resolution)
sessions.py Session storage: atomic writes, per-session locks, TTL cleanup
images.py Path/content validation, staged compression, data URL encoding
vision.py Vision client: OpenAI-compatible API
Tests
.venv/bin/python test/test_mcp_proto.py # MCP handshake and tool registration
.venv/bin/python test/test_e2e.py # end-to-end (requires a real API key)
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。