GLM Vision MCP

GLM Vision MCP

Enables text-only reasoning models to see images by wrapping vision-language models as MCP tools, supporting image description, OCR, chart analysis, and custom questioning within MCP-compatible IDEs.

Category
访问服务器

README

GLM Vision MCP

<p align="center"> <strong>Give any text-only reasoning model a pair of eyes.</strong> </p>

<p align="center"> An MCP (Model Context Protocol) server that wraps vision-language models — like <strong>GLM-4.6V-Flash</strong> — as callable tools, so text-only LLMs (DeepSeek-R1, Qwen3-Thinking, etc.) can <em>see</em> images inside Cursor, Trae CN, Cline, Claude Desktop, and any MCP-compatible agent. </p>


How It Works

┌─────────────────────────────────────────────────────────┐
│                    Your IDE / Agent                     │
│                                                         │
│  Text-only model (e.g. DeepSeek-R1)                      │
│       │                                                 │
│       │  "What's in this screenshot?"                    │
│       ▼                                                 │
│  MCP Tool: see_image(path, question)                    │
│       │                                                 │
│       ▼  (stdio / MCP protocol)                         │
│  ┌──────────────────────────────────────────────────┐   │
│  │           GLM Vision MCP Server                  │   │
│  │                                                  │   │
│  │  Reads image → base64 → sends to vision model   │   │
│  │  (GLM-4.6V-Flash / GPT-4o / any VLM)            │   │
│  │                                                  │   │
│  │  Returns: "A login error dialog showing..."      │   │
│  └──────────────────────────────────────────────────┘   │
│       │                                                 │
│       ▼                                                 │
│  Text-only model continues reasoning with vision data   │
└─────────────────────────────────────────────────────────┘

You configure 3 things: model provider, model ID, and API key. The server handles everything else — image reading, base64 encoding, API calls, error handling.


Quick Start

1. Install

# From PyPI (once published) or from source:
pip install -e .

Or use directly with uv / pipx without installing:

# Using uv (recommended for MCP)
uv run glm-vision-mcp

2. Configure

The server reads configuration from environment variables:

Variable Required Default Description
VISION_API_KEY ✅ Yes — Your model provider's API key
VISION_MODEL_ID ✅ Yes glm-4.6v-flash The vision model to use
VISION_MODEL_PROVIDER ✅ Yes zhipu Provider name (see table below)
VISION_BASE_URL ❌ No provider default Custom API base URL
VISION_MAX_TOKENS ❌ No 2048 Max response tokens
VISION_TEMPERATURE ❌ No 0.4 Sampling temperature

3. Add to Your IDE

Cursor

Add to ~/.cursor/mcp.json (or .cursor/mcp.json in your project):

{
  "mcpServers": {
    "glm-vision": {
      "command": "python",
      "args": ["-m", "glm_vision_mcp"],
      "env": {
        "VISION_MODEL_PROVIDER": "zhipu",
        "VISION_MODEL_ID": "glm-4.6v-flash",
        "VISION_API_KEY": "your-zhipu-api-key-here"
      }
    }
  }
}

Trae CN

Add to Trae's MCP settings (设置 → MCP):

{
  "mcpServers": {
    "glm-vision": {
      "command": "python",
      "args": ["-m", "glm_vision_mcp"],
      "env": {
        "VISION_MODEL_PROVIDER": "zhipu",
        "VISION_MODEL_ID": "glm-4.6v-flash",
        "VISION_API_KEY": "your-zhipu-api-key-here"
      }
    }
  }
}

Cline (VS Code)

Add to ~/.cline/mcp_settings.json:

{
  "mcpServers": {
    "glm-vision": {
      "command": "python",
      "args": ["-m", "glm_vision_mcp"],
      "env": {
        "VISION_MODEL_PROVIDER": "zhipu",
        "VISION_MODEL_ID": "glm-4.6v-flash",
        "VISION_API_KEY": "your-zhipu-api-key-here"
      },
      "disabled": false,
      "autoApprove": []
    }
  }
}

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "glm-vision": {
      "command": "python",
      "args": ["-m", "glm_vision_mcp"],
      "env": {
        "VISION_MODEL_PROVIDER": "zhipu",
        "VISION_MODEL_ID": "glm-4.6v-flash",
        "VISION_API_KEY": "your-zhipu-api-key-here"
      }
    }
  }
}

Generic MCP Client (any MCP-compatible tool)

{
  "mcpServers": {
    "glm-vision": {
      "command": "python",
      "args": ["-m", "glm_vision_mcp"],
      "env": {
        "VISION_MODEL_PROVIDER": "zhipu",
        "VISION_MODEL_ID": "glm-4.6v-flash",
        "VISION_API_KEY": "your-api-key"
      }
    }
  }
}

Tip: If you installed via uv, use "command": "uv" and "args": ["run", "glm-vision-mcp"] instead.


Supported Providers

Provider VISION_MODEL_PROVIDER Default Base URL Example Models
Zhipu (智谱) zhipu https://open.bigmodel.cn/api/paas/v4 glm-4.6v-flash, glm-4v-plus
OpenAI openai https://api.openai.com/v1 gpt-4o, gpt-4o-mini
DeepSeek deepseek https://api.deepseek.com/v1 deepseek-vl
Moonshot (Kimi) moonshot https://api.moonshot.cn/v1 moonshot-v1-8k-vision
SiliconFlow siliconflow https://api.siliconflow.cn/v1 Qwen/Qwen2-VL-72B
Custom custom (you set VISION_BASE_URL) Any OpenAI-compatible VLM

Using a Custom Endpoint

Set VISION_BASE_URL to point to your own server (vLLM, Ollama, LM Studio, etc.):

{
  "env": {
    "VISION_MODEL_PROVIDER": "custom",
    "VISION_MODEL_ID": "your-model-name",
    "VISION_API_KEY": "any-or-empty",
    "VISION_BASE_URL": "http://localhost:8000/v1"
  }
}

Tools

The server exposes 4 tools. Your IDE's agent will automatically call them when it needs vision:

see_image — Core Vision Tool

Ask any question about an image.

see_image(image, question="What is in this image?")
  • image: File path, URL, or base64 string
  • question: What you want to know (default: "What is in this image?")

describe_image — Image Description

Generate a text description/caption.

describe_image(image, detail_level="detailed")
  • detail_level: "brief" | "detailed" | "exhaustive" (default: "detailed")

extract_text — OCR

Extract all visible text from an image.

extract_text(image, language_hint="Chinese")
  • language_hint: Optional — e.g. "Chinese", "English", "mixed"

analyze_chart — Chart & Diagram Analysis

Analyze charts, graphs, architecture diagrams, or UI screenshots.

analyze_chart(image, question="")
  • question: Optional specific question (default: general analysis)

Usage Example

Once configured, just talk to your IDE's agent normally:

You: "Look at the screenshot at /tmp/error.png — what's wrong?"

The agent will:

  1. Call see_image("/tmp/error.png", "What error is shown?")
  2. The MCP server sends the image to GLM-4.6V-Flash
  3. GLM returns: "The dialog shows a 'Connection Refused' error..."
  4. Your text-only model uses that answer to help you

Getting an API Key

Zhipu (智谱) — Free Tier Available

  1. Visit https://open.bigmodel.cn
  2. Sign up / log in
  3. Go to API Keys → Create new key
  4. Copy the key (format: xxxxxxxx.xxxxxxxx)

GLM-4.6V-Flash offers free quota — great for testing!

OpenAI

  1. Visit https://platform.openai.com/api-keys
  2. Create a new key

Development

Project Structure

glm-vision-mcp/
├── src/glm_vision_mcp/
│   ├── __init__.py
│   ├── __main__.py          # python -m glm_vision_mcp
│   ├── server.py            # MCP server + tool definitions
│   ├── config.py            # Env-var config loader
│   ├── utils.py             # Image encoding utilities
│   └── providers/
│       ├── base.py          # VisionProvider (shared HTTP logic)
│       ├── zhipu.py         # Zhipu GLM provider
│       ├── openai_compat.py # Generic OpenAI-compatible provider
│       └── registry.py      # Provider name → class mapping
├── examples/
│   └── quickstart.py
├── pyproject.toml
├── requirements.txt
└── README.md

Running Tests

pip install -e ".[dev]"
pytest

Adding a New Provider

  1. Create src/glm_vision_mcp/providers/my_provider.py:
from glm_vision_mcp.providers.base import VisionProvider

class MyProvider(VisionProvider):
    def _build_headers(self):
        # Custom auth if needed
        return {"X-Api-Key": self.config.api_key}
  1. Register in providers/registry.py:
_PROVIDERS["my_provider"] = MyProvider

FAQ

Q: Can I use this with a non-vision model? No — the configured model must support vision input (images). If you're unsure, GLM-4.6V-Flash is a good free option.

Q: Does it work with local images? Yes. Pass a file path and the server will read and base64-encode it automatically.

Q: How fast is it? Depends on the model provider. GLM-4.6V-Flash is very fast (typically 1-3 seconds per image).

Q: Can multiple images be analyzed at once? Currently each tool call handles one image. For multi-image comparison, call see_image multiple times or extend the tools.


License

MIT © 2026 xiayuyang750

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选