GLM Vision MCP
Enables text-only reasoning models to see images by wrapping vision-language models as MCP tools, supporting image description, OCR, chart analysis, and custom questioning within MCP-compatible IDEs.
README
GLM Vision MCP
<p align="center"> <strong>Give any text-only reasoning model a pair of eyes.</strong> </p>
<p align="center"> An MCP (Model Context Protocol) server that wraps vision-language models — like <strong>GLM-4.6V-Flash</strong> — as callable tools, so text-only LLMs (DeepSeek-R1, Qwen3-Thinking, etc.) can <em>see</em> images inside Cursor, Trae CN, Cline, Claude Desktop, and any MCP-compatible agent. </p>
How It Works
┌─────────────────────────────────────────────────────────┐
│ Your IDE / Agent │
│ │
│ Text-only model (e.g. DeepSeek-R1) │
│ │ │
│ │ "What's in this screenshot?" │
│ ▼ │
│ MCP Tool: see_image(path, question) │
│ │ │
│ ▼ (stdio / MCP protocol) │
│ ┌──────────────────────────────────────────────────┐ │
│ │ GLM Vision MCP Server │ │
│ │ │ │
│ │ Reads image → base64 → sends to vision model │ │
│ │ (GLM-4.6V-Flash / GPT-4o / any VLM) │ │
│ │ │ │
│ │ Returns: "A login error dialog showing..." │ │
│ └──────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ Text-only model continues reasoning with vision data │
└─────────────────────────────────────────────────────────┘
You configure 3 things: model provider, model ID, and API key. The server handles everything else — image reading, base64 encoding, API calls, error handling.
Quick Start
1. Install
# From PyPI (once published) or from source:
pip install -e .
Or use directly with uv / pipx without installing:
# Using uv (recommended for MCP)
uv run glm-vision-mcp
2. Configure
The server reads configuration from environment variables:
| Variable | Required | Default | Description |
|---|---|---|---|
VISION_API_KEY |
✅ Yes | — | Your model provider's API key |
VISION_MODEL_ID |
✅ Yes | glm-4.6v-flash |
The vision model to use |
VISION_MODEL_PROVIDER |
✅ Yes | zhipu |
Provider name (see table below) |
VISION_BASE_URL |
❌ No | provider default | Custom API base URL |
VISION_MAX_TOKENS |
❌ No | 2048 |
Max response tokens |
VISION_TEMPERATURE |
❌ No | 0.4 |
Sampling temperature |
3. Add to Your IDE
Cursor
Add to ~/.cursor/mcp.json (or .cursor/mcp.json in your project):
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
}
}
}
}
Trae CN
Add to Trae's MCP settings (设置 → MCP):
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
}
}
}
}
Cline (VS Code)
Add to ~/.cline/mcp_settings.json:
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
},
"disabled": false,
"autoApprove": []
}
}
}
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
}
}
}
}
Generic MCP Client (any MCP-compatible tool)
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-api-key"
}
}
}
}
Tip: If you installed via
uv, use"command": "uv"and"args": ["run", "glm-vision-mcp"]instead.
Supported Providers
| Provider | VISION_MODEL_PROVIDER |
Default Base URL | Example Models |
|---|---|---|---|
| Zhipu (智谱) | zhipu |
https://open.bigmodel.cn/api/paas/v4 |
glm-4.6v-flash, glm-4v-plus |
| OpenAI | openai |
https://api.openai.com/v1 |
gpt-4o, gpt-4o-mini |
| DeepSeek | deepseek |
https://api.deepseek.com/v1 |
deepseek-vl |
| Moonshot (Kimi) | moonshot |
https://api.moonshot.cn/v1 |
moonshot-v1-8k-vision |
| SiliconFlow | siliconflow |
https://api.siliconflow.cn/v1 |
Qwen/Qwen2-VL-72B |
| Custom | custom |
(you set VISION_BASE_URL) |
Any OpenAI-compatible VLM |
Using a Custom Endpoint
Set VISION_BASE_URL to point to your own server (vLLM, Ollama, LM Studio, etc.):
{
"env": {
"VISION_MODEL_PROVIDER": "custom",
"VISION_MODEL_ID": "your-model-name",
"VISION_API_KEY": "any-or-empty",
"VISION_BASE_URL": "http://localhost:8000/v1"
}
}
Tools
The server exposes 4 tools. Your IDE's agent will automatically call them when it needs vision:
see_image — Core Vision Tool
Ask any question about an image.
see_image(image, question="What is in this image?")
- image: File path, URL, or base64 string
- question: What you want to know (default: "What is in this image?")
describe_image — Image Description
Generate a text description/caption.
describe_image(image, detail_level="detailed")
- detail_level:
"brief"|"detailed"|"exhaustive"(default:"detailed")
extract_text — OCR
Extract all visible text from an image.
extract_text(image, language_hint="Chinese")
- language_hint: Optional — e.g.
"Chinese","English","mixed"
analyze_chart — Chart & Diagram Analysis
Analyze charts, graphs, architecture diagrams, or UI screenshots.
analyze_chart(image, question="")
- question: Optional specific question (default: general analysis)
Usage Example
Once configured, just talk to your IDE's agent normally:
You: "Look at the screenshot at
/tmp/error.png— what's wrong?"
The agent will:
- Call
see_image("/tmp/error.png", "What error is shown?") - The MCP server sends the image to GLM-4.6V-Flash
- GLM returns: "The dialog shows a 'Connection Refused' error..."
- Your text-only model uses that answer to help you
Getting an API Key
Zhipu (智谱) — Free Tier Available
- Visit https://open.bigmodel.cn
- Sign up / log in
- Go to API Keys → Create new key
- Copy the key (format:
xxxxxxxx.xxxxxxxx)
GLM-4.6V-Flash offers free quota — great for testing!
OpenAI
- Visit https://platform.openai.com/api-keys
- Create a new key
Development
Project Structure
glm-vision-mcp/
├── src/glm_vision_mcp/
│ ├── __init__.py
│ ├── __main__.py # python -m glm_vision_mcp
│ ├── server.py # MCP server + tool definitions
│ ├── config.py # Env-var config loader
│ ├── utils.py # Image encoding utilities
│ └── providers/
│ ├── base.py # VisionProvider (shared HTTP logic)
│ ├── zhipu.py # Zhipu GLM provider
│ ├── openai_compat.py # Generic OpenAI-compatible provider
│ └── registry.py # Provider name → class mapping
├── examples/
│ └── quickstart.py
├── pyproject.toml
├── requirements.txt
└── README.md
Running Tests
pip install -e ".[dev]"
pytest
Adding a New Provider
- Create
src/glm_vision_mcp/providers/my_provider.py:
from glm_vision_mcp.providers.base import VisionProvider
class MyProvider(VisionProvider):
def _build_headers(self):
# Custom auth if needed
return {"X-Api-Key": self.config.api_key}
- Register in
providers/registry.py:
_PROVIDERS["my_provider"] = MyProvider
FAQ
Q: Can I use this with a non-vision model? No — the configured model must support vision input (images). If you're unsure, GLM-4.6V-Flash is a good free option.
Q: Does it work with local images? Yes. Pass a file path and the server will read and base64-encode it automatically.
Q: How fast is it? Depends on the model provider. GLM-4.6V-Flash is very fast (typically 1-3 seconds per image).
Q: Can multiple images be analyzed at once?
Currently each tool call handles one image. For multi-image comparison, call see_image multiple times or extend the tools.
License
MIT © 2026 xiayuyang750
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。