mata-kadalz
Lizard Eyes — a local vision MCP server backed by Qwen3-VL-4B, exposing a single vision.inspect tool for analyzing images via llama-server.
README
mata-kadalz 🦎
Lizard Eyes — local vision MCP server backed by Qwen3-VL-4B running on llama-server.
Gives any MCP client (opencode, Claude, Codex, ...) one tool: vision.inspect(image_path, task).
client -> vision.inspect -> mata-kadalz server.py (stdio or streamable HTTP)
-> HTTP POST http://<llama-server>:9931/v1/chat/completions
-> llama-server (Qwen3-VL-4B GGUF + mmproj)
mata-kadalz is a thin MCP layer. The inference runtime (llama-server from llama.cpp) and the
model are external dependencies you install yourself; this repo never downloads or bundles them.
Quick start
Install from source. This project is not yet published to PyPI, so the package is installed from this repository (see Setup guides for your platform).
- Install Python >= 3.11.
- Install llama-server (llama.cpp) and the Qwen3-VL-4B model + mmproj — external, from official sources (links in the Model and Setup guides sections).
- Start
llama-serverand confirm it is healthy:curl http://localhost:9931/health→{"status":"ok"}. - Install
mata-kadalzfrom this repo into a venv:git clone https://github.com/kadalzbaiq/mata-kadalz.git cd mata-kadalz bash scripts/install.sh # POSIX (Linux/macOS/WSL); Windows: see docs/SETUP-windows.md - Confirm the MCP server can reach llama-server:
.venv/bin/mata-kadalz --health. - Register
mata-kadalzin your MCP client — see Client registration.
Supported setups
| Host | llama-server runs on | Setup doc |
|---|---|---|
| Windows (native) | Windows | docs/SETUP-windows.md |
| Linux (native) | Linux | docs/SETUP-linux.md |
| macOS (native) | macOS | docs/SETUP-macos.md |
| WSL2 on Windows | Windows host (native) | docs/SETUP-hybrid-wsl.md |
| Any / remote / custom | anywhere reachable over HTTP | docs/SETUP-modular.md |
Model
- Model:
Qwen3VL-4B-Instruct-Q4_K_M.gguf(LLM, ~2.5 GB) - Vision encoder:
mmproj-Qwen3VL-4B-Instruct-F16.gguf(~800 MB) - Source:
Qwen/Qwen3-VL-4B-Instruct-GGUF- https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct-GGUF/resolve/main/Qwen3VL-4B-Instruct-Q4_K_M.gguf
- https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct-GGUF/resolve/main/mmproj-Qwen3VL-4B-Instruct-F16.gguf
Do not change the quant or mmproj combo without re-validating; this is the verified working setup.
Layout
server.py # MCP server (stdio + streamable HTTP), single file
config/config.json # optional overrides (empty = auto)
scripts/install.sh # install the mata-kadalz package only
docs/ # per-host setup guides
tests/ # pytest (no llama-server needed)
runtime/ # logs + image cache (gitignored)
Client registration
Replace /path/to/mata-kadalz/.venv/bin/mata-kadalz with the real path on your machine.
stdio (local — same machine as the images):
# claude
claude mcp add vision -- /path/to/mata-kadalz/.venv/bin/mata-kadalz
# codex
codex mcp add vision -- /path/to/mata-kadalz/.venv/bin/mata-kadalz
streamable HTTP (remote — server machine may differ from client):
# claude
claude mcp add --transport http vision http://127.0.0.1:9932/mcp
# codex
codex mcp add vision --url http://127.0.0.1:9932/mcp
opencode (config in opencode.json):
{
"mcp": {
"vision": {
"type": "local", // stdio
"command": ["/path/to/mata-kadalz/.venv/bin/mata-kadalz"]
}
// or "type": "remote", "url": "http://127.0.0.1:9932/mcp" // streamable HTTP
}
}
Usage
One tool: vision.inspect — takes image_path (absolute path on the machine where the MCP server runs) and task (what to analyze). Returns structured JSON:
{ "success": true, "summary": "...", "details": "", "warnings": [], "cache_hit": true }
On failure it returns is_error: true with a machine-readable code, e.g. IMAGE_NOT_FOUND, IMAGE_NOT_SUPPORTED, IMAGE_PATH_NOT_ALLOWED, LLAMA_SERVER_URL_NOT_SET, LLAMA_SERVER_TIMEOUT, LLAMA_BUSY, INVALID_VISION_RESPONSE.
The server also embeds a system prompt (SYSTEM_PROMPT) instructing the vision model how to structure its output; the prompt is sent on every inference request.
HTTP vs stdio — where files must live
- stdio (local): the MCP client and the server share one machine, so
image_pathis a path on that machine. - streamable HTTP (remote): the client and the server may be on different machines —
image_pathis resolved on the server machine, not the client's. SetVISION_IMAGE_ROOTSto restrict which directories the server will read from (strongly recommended for a network-exposed server).
Security for HTTP deployments
If you expose the server over the network:
- Set
VISION_IMAGE_ROOTSso the server can only read from the directories you choose. With it unset, the server can read any path on the host. - Bind to a safe interface.
--host 127.0.0.1(the default) only accepts local connections. For LAN/remote access, prefer a VPN or a firewall rule over binding0.0.0.0on a public interface. - No authentication is built in. Put the endpoint behind an authenticated reverse proxy or your VPN. The HTTP transport speaks raw MCP; there is no token/user layer.
- Allow inbound traffic only on the ports you use: 9931 (llama-server) and 9932 (mata-kadalz HTTP transport).
Configuration
Config precedence: DEFAULTS < config/config.json (empty values skipped) < environment variables.
| Key | Default | Notes |
|---|---|---|
LLAMA_SERVER_URL |
http://<gateway-ip>:9931 |
Auto-detects WSL gateway IP; set explicitly to override. If WSL gateway detection fails and no URL is set, inference returns LLAMA_SERVER_URL_NOT_SET |
VISION_RUNTIME_DIR |
<repo>/runtime/vision |
Relative paths resolve against repo root |
VISION_CACHE_DIR |
<repo>/runtime/vision/cache |
|
VISION_LOG_DIR |
<repo>/runtime/vision/logs |
|
VISION_TIMEOUT_SECONDS |
180 |
CPU inference takes 10–130 s per call |
VISION_MAX_IMAGE_SIZE |
20971520 |
20 MB |
VISION_MODEL_ID |
qwen3-vl |
|
VISION_IMAGE_ROOTS |
(empty = any path) | Restrict readable image dirs. JSON array or comma-separated, relative to repo root. Symlinks and .. escapes resolve and are rejected |
VISION_MAX_QUEUE |
4 |
Bounded inference queue; beyond this, calls fail fast with LLAMA_BUSY |
Example override in config/config.json:
{ "LLAMA_SERVER_URL": "http://192.168.64.1:9931" }
Restrict a network-exposed server to one directory:
{ "VISION_IMAGE_ROOTS": ["/srv/shared-images"] }
Health check
Confirm the MCP can reach llama-server before wiring up a client:
.venv/bin/mata-kadalz --health
Prints platform, resolved LLAMA_SERVER_URL, and a reachable: true/false health probe; exits 0 when reachable.
Caching
Requests are deduplicated by sha256(image) + task + model. A cache hit returns instantly without touching llama-server. Failed requests are never cached. Inference is serialized with a process-wide lock (one concurrent call at a time); the queue beyond the lock is bounded by VISION_MAX_QUEUE and returns LLAMA_BUSY when full.
Cache invalidation: changing VISION_MODEL_ID invalidates the cache automatically, because the model id is part of the cache key — stale answers from an older model are never served.
Cancellation: if the MCP client cancels a request mid-inference, the server releases its lock and queue slot immediately, discards the in-flight result (never caches a partial one), and the error propagates without crashing the server. The underlying llama-server call keeps running in the background; its result is ignored.
Self-check
echo '{"image_path":"/path/to/img.png","task":"describe"}' | .venv/bin/mata-kadalz --once
Reads one JSON request from stdin, runs it, prints the result, and exits: 0 on success or cache hit, 1 on any error (invalid input, missing file, unreachable llama-server, ...). Useful for scripting and cron-style smoke checks.
Test
.venv/bin/python -m pytest
No llama-server required — tests cover config, file validation, magic bytes, cache logic, image-root policy, cancellation, --once exit codes, WSL gateway detection, bounded queue, and a real streamable-HTTP session over uvicorn.
HTTP transport dependency
uvicorn is only needed for --transport http. It already ships transitively with the MCP SDK, but it is also declared as an explicit optional extra so the intent is unambiguous:
# from a source checkout (until PyPI publication)
pip install -e '.[http]' # or: bash scripts/install.sh then pip install -e '.[http]'
Until the package is published to PyPI,
pip install mata-kadalzandpip install 'mata-kadalz[http]'will not work. Use a source checkout.
License
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。