ocular
MCP server that provides vision capabilities to coding agents, enabling them to analyze screenshots, UI mockups, terminal errors, documents, tables, and charts through OpenAI-compatible vision models. Supports local stdio and remote HTTP deployments with structured JSON output and binary upload side channels.
README
ocular
Vision for coding agents.
ocular is an MCP server that lets text-first coding agents analyze screenshots, UI mockups, terminal errors, documents, tables, and charts through OpenAI-compatible vision models.
It is designed for both local stdio use and remote HTTP deployments. For remote agents, image bytes can travel through a binary upload side channel while MCP tool calls carry only a lightweight file_id, avoiding large inline base64 payloads.
Project status: early-stage and actively evolving. Feedback, bug reports, integrations, and real-world usage reports are welcome.
Why ocular?
Coding agents are good at reading source code but often lose context when the important evidence is visual: a broken layout, a terminal screenshot, an error dialog, a chart, or a design reference.
ocular turns those visual inputs into structured data an agent can reason about.
- Agent-oriented output — tools return structured JSON instead of prose-only descriptions.
- 8 focused vision tools — general analysis, OCR, UI inspection, error diagnosis, UI comparison, table extraction, chart analysis, and upload orchestration.
- OpenAI-compatible provider interface — point ocular at a compatible multimodal endpoint and model; reproducibly tested configurations are tracked in Provider compatibility.
- Remote-friendly uploads — binary
PUT /uploadflow for large images with content-addressedfile_idreferences. - Local or remote MCP — stdio for local clients, HTTP for hosted/private deployments.
- Caching and persistence — deduplicated uploads plus result caching for repeated agent workflows.
How it works
flowchart LR
A[Coding agent] -->|MCP tool call| B[ocular]
C[Image / screenshot] -->|binary upload or base64| B
B -->|OpenAI-compatible request| D[Vision model]
D -->|multimodal response| B
B -->|structured JSON| A
For remote HTTP deployments, the recommended path is:
image bytes -> PUT /upload -> file_id -> MCP vision tool -> structured result
See Architecture for the upload and caching model.
Demo
Want to see the full handoff from screenshot to coding-agent evidence? Read the end-to-end demo.
It walks through a remote image upload, a diagnose_error_screenshot call, the structured fields returned to the agent, and how that evidence is combined with repository context. Example model output is explicitly marked representative rather than presented as a benchmark.
Quick start
1. Install
The published npm package is ocular-mcp. It installs the CLI command ocular.
Global install:
npm install -g ocular-mcp
Or run it without a global install:
npx -y ocular-mcp
To build from source instead:
git clone https://github.com/xyun1996/ocular.git
cd ocular
npm install
npm run build
2. Configure a vision provider
ocular requires an OpenAI-compatible multimodal endpoint, API key, and model name:
OCULAR_BASE_URL=https://your-openai-compatible-endpoint.example/v1
OCULAR_API_KEY=your_api_key
OCULAR_MODEL=your_vision_model
For a local compatible endpoint, use that server's base URL and vision-capable model name. Compatibility depends on the endpoint/model combination; see Provider compatibility for the reproducible smoke-test procedure and verified configurations.
3. Run in stdio mode
With a global install:
ocular
Or:
npx -y ocular-mcp
The server communicates over stdio, so it may appear idle when started directly. In normal use an MCP client launches it and exchanges protocol messages over stdin/stdout.
4. Connect an MCP client
Claude Code example:
claude mcp add ocular \
-e OCULAR_BASE_URL=https://your-openai-compatible-endpoint.example/v1 \
-e OCULAR_MODEL=your_vision_model \
-e OCULAR_API_KEY=your_api_key \
-- npx -y ocular-mcp
Avoid putting long-lived API keys directly in shell history on shared machines. Use your client's environment/secret-management mechanism when available.
For a generic MCP client:
{
"mcpServers": {
"ocular": {
"command": "npx",
"args": ["-y", "ocular-mcp"],
"env": {
"OCULAR_BASE_URL": "https://your-openai-compatible-endpoint.example/v1",
"OCULAR_MODEL": "your_vision_model",
"OCULAR_API_KEY": "your_api_key"
}
}
}
}
See Claude Code setup for a fuller walkthrough.
Example workflows
Diagnose a screenshot
Ask your coding agent to inspect an error screenshot and extract the exact message, likely cause, and next checks.
{
"file_id": "e21ba723...",
"task": "Extract the exact error and suggest the next debugging checks",
"project_context": "Node.js TypeScript project"
}
Review a UI implementation
Use analyze_ui_screenshot to turn a screenshot into implementation-oriented observations about hierarchy, alignment, spacing, typography, contrast, and likely visual defects.
Compare expected vs actual UI
Use compare_ui_screenshots with a reference screenshot and an implementation screenshot to identify regressions and layout differences.
See Screenshot debugging example.
Tools
| Tool | Purpose |
|---|---|
analyze_image |
General structured image analysis |
extract_text_from_image |
OCR with reading-order/layout awareness |
analyze_ui_screenshot |
UI hierarchy, spacing, typography and accessibility review |
diagnose_error_screenshot |
Extract and diagnose terminal/browser/build errors |
compare_ui_screenshots |
Compare reference and implementation screenshots |
extract_table_from_image |
Extract table data into structured output |
analyze_chart_image |
Analyze chart labels, values, trends and uncertainty |
create_upload_session |
Return upload endpoint and instructions for remote clients |
Every vision tool accepts file_id; local workflows can also use inline image_base64 where appropriate.
Remote deployment
Set HTTP transport and authentication:
MCP_TRANSPORT=http
MCP_HTTP_HOST=127.0.0.1
MCP_HTTP_PORT=3000
MCP_HTTP_PATH=/mcp
MCP_AUTH_TOKEN=replace_with_a_long_random_token
MCP_AUTH_HEADER=authorization
MCP_AUTH_SCHEME=Bearer
Upload raw bytes:
curl --request PUT \
--data-binary @/path/to/image.png \
"https://your.host/upload" \
-H "Content-Type: image/png" \
-H "Authorization: Bearer your_mcp_auth_token"
The server returns a content-addressed file_id; pass that id to a vision tool instead of sending a large base64 string through MCP.
For reverse proxy and systemd examples, see Deployment.
Configuration
Common variables:
| Variable | Purpose |
|---|---|
OCULAR_BASE_URL |
OpenAI-compatible API base URL |
OCULAR_API_KEY |
Provider API key |
OCULAR_MODEL |
Vision-capable model name |
OCULAR_HEADERS |
Optional custom provider headers as JSON |
OCULAR_TEMPERATURE |
Generation temperature |
OCULAR_MAX_TOKENS |
Maximum generated tokens |
OCULAR_TIMEOUT_MS |
Provider timeout |
OCULAR_MAX_IMAGE_MB |
Maximum image size |
OCULAR_CACHE_ENABLED |
Enable result cache |
OCULAR_CACHE_DIR |
Cache directory |
OCULAR_UPLOADS_DIR |
Persistent upload directory |
OCULAR_UPLOAD_URL_BASE |
Public base URL used in upload instructions |
See .env.example for the full configuration surface.
Verification and benchmarks
Provider compatibility claims are based on real endpoint/model smoke tests, not on API naming alone. See Provider compatibility.
The repository also includes synthetic, redistributable visual fixtures for repeatable project-level measurements. See Benchmark fixtures. The benchmark measures execution, structural JSON output, and timing; it is not presented as a broad model-quality ranking.
Development
npm install
npm run build
npm test
npm run check
npm run dev
The repository includes tests for authentication, caching, image handling, MCP server behavior, provider payloads, tool execution, npm packaging, Registry metadata consistency, and release smoke checks.
Security and privacy
Do not commit provider API keys or MCP authentication tokens. Public HTTP deployments should sit behind HTTPS and a reverse proxy; the Node process should generally bind to a private interface.
See SECURITY.md for vulnerability reporting guidance.
Roadmap
Near-term areas where contributions are useful:
- Real-world MCP client integration and usage reports
- Provider/model compatibility verification
- Published fixture-based benchmark results from real endpoints
- Community-driven tool and prompt improvements
If you are using ocular in a real workflow, open a Usage report issue describing the client, provider/model, and use case. Public reports are useful even when nothing is broken and help keep compatibility/adoption claims grounded in real usage.
Contributing
Contributions are welcome. Start with CONTRIBUTING.md, run npm run check before opening a PR, and include reproduction details for behavior changes.
Release and registry
The npm package is ocular-mcp; the installed CLI is ocular. Release automation uses npm Trusted Publishing rather than a long-lived repository token. See Publishing.
ocular is published in the official MCP Registry as io.github.xyun1996/ocular. See MCP Registry for the live identity and release flow.
License
MIT — see LICENSE.
If ocular is useful in your agent workflow, a GitHub star helps other developers discover the project.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。