openai-image-mcp

openai-image-mcp

A minimal MCP server for generating images via OpenAI's GPT image model, supporting inline display or file output.

Category
访问服务器

README

openai-image-mcp

A minimal, production-ready MCP server that exposes OpenAI GPT Image generation as two well-defined tools: generate_image (returns the PNG inline as MCP image content) and generate_image_to_file (writes the PNG to disk for asset pipelines).

Part of the mcp-accelerator: it is the reference implementation of the accelerator's "local stdio server with externalized secrets" pattern for APIs that have no official hosted MCP server.

This documentation follows the structure of a TOGAF Architecture Definition Document (ADM phases A through F), so that both humans and AI agents can extract the intent, the contracts, and the constraints without reading the code first. If you are an AI agent, start with Section 8: Agent Usage Contract.


Table of Contents

  1. Architecture Vision (Phase A)
  2. Business Architecture (Phase B)
  3. Application Architecture (Phase C)
  4. Data Architecture (Phase C)
  5. Technology Architecture (Phase D)
  6. Security Architecture (cross-cutting)
  7. Implementation and Migration (Phases E and F)
  8. Agent Usage Contract
  9. Operations and Troubleshooting

1. Architecture Vision (Phase A)

1.1 Problem statement

AI assistants such as Claude can reason about images but cannot create them with OpenAI models out of the box. OpenAI offers no official hosted MCP server for its image API. Teams therefore either paste generated images around manually or build one-off integrations per project.

1.2 Vision

Provide image generation as a first-class MCP capability: any MCP host (Claude Code, Claude Desktop) gains a generate_image tool that produces brand assets, mockup imagery, placeholder art, or documentation illustrations on demand, inside the conversation, with the API key kept out of every config file.

1.3 Stakeholders and concerns

Stakeholder Concern How this architecture addresses it
Developer using Claude Code Generate assets without leaving the terminal Tools are available in-session after a one-line config entry
Project team Reusable setup across projects Config templates live in the mcp-accelerator, the server is generic
Security officer API keys must not leak via dotfiles or repos Key is resolved at process launch from the environment or the macOS Keychain, never stored in config
AI agent (Claude) Needs deterministic tool contracts Strict parameter validation, explicit error messages, contract documented in Section 8

1.4 Value proposition

  • One capability, two delivery modes: inline image content for conversational use, file output for build and asset pipelines.
  • Zero configuration drift: the server has no config file of its own. Behavior is fully determined by the tool parameters and one environment variable.
  • Small enough to audit in five minutes: a single Python module of about 75 lines.

2. Business Architecture (Phase B)

2.1 Business capability

"Generative image supply for software delivery": producing visual assets (app icons drafts, hero images, empty-state illustrations, test fixtures, blog graphics) as part of the normal engineering workflow rather than as a separate design request.

2.2 Core use cases

ID Use case Actor Tool used
UC-1 Generate an illustration during a conversation and inspect it immediately Developer via Claude generate_image
UC-2 Generate an asset directly into the project tree (for example assets/hero.png) Claude Code in an implementation task generate_image_to_file
UC-3 Batch-produce placeholder images for a UI prototype Claude Code, iterating generate_image_to_file
UC-4 Create documentation or listing graphics (App Store, README) Developer via Claude either

2.3 Value chain position

flowchart LR
    A[Requirement<br/>needs a visual] --> B[Prompt<br/>formulated by human or AI]
    B --> C[openai-image-mcp<br/>MCP tool call]
    C --> D[PNG asset]
    D --> E1[Inline review<br/>in conversation]
    D --> E2[Committed to repo<br/>asset pipeline]

3. Application Architecture (Phase C)

3.1 Component model

flowchart LR
    subgraph Host["MCP Host (Claude Code / Claude Desktop)"]
        LLM[Claude]
    end
    subgraph Server["openai-image-mcp (local process)"]
        FMCP[FastMCP runtime<br/>stdio transport]
        T1[Tool: generate_image]
        T2[Tool: generate_image_to_file]
        GEN[_generate_png<br/>validation + API call]
    end
    subgraph OpenAI["OpenAI Platform"]
        IMG[Images API<br/>model gpt-image-2]
    end
    FS[(Local filesystem)]
    KEY[/OPENAI_API_KEY<br/>env or Keychain/]

    LLM -- "JSON-RPC over stdio" --> FMCP
    FMCP --> T1 & T2
    T1 & T2 --> GEN
    GEN -- "HTTPS" --> IMG
    T2 -- "write PNG" --> FS
    KEY -. "read at process start" .-> Server

3.2 Tool catalog

Tool Purpose Returns
generate_image(prompt, size="auto") Generate a PNG and return it as MCP image content Image (base64 PNG, rendered inline by the host)
generate_image_to_file(prompt, output_path, size="auto") Generate a PNG and persist it to an absolute path Confirmation string with path and byte count

3.3 Interaction (runtime view)

sequenceDiagram
    participant C as Claude (MCP host)
    participant S as openai-image-mcp
    participant O as OpenAI Images API
    C->>S: tools/call generate_image_to_file(prompt, output_path, size)
    S->>S: validate prompt not empty, size allowed, path absolute
    S->>O: images.generate(model=gpt-image-2, output_format=png)
    O-->>S: b64_json payload
    S->>S: decode, create parent dirs, write file
    S-->>C: "Image saved: /path/file.png (N bytes)"

4. Data Architecture (Phase C)

4.1 Data flows and classification

Data Direction Classification Persistence
Prompt text Host to server to OpenAI Potentially confidential (leaves the machine) Not persisted by the server
API key Environment to server to OpenAI (HTTPS auth header) Secret Never written to disk by the server
Generated PNG OpenAI to server to host or filesystem Project asset Only persisted when generate_image_to_file is used

4.2 Data principles

  • The server is stateless: no cache, no logs of prompts or images, no temp files.
  • Prompts are transmitted to OpenAI. Do not include secrets, personal data, or unreleased confidential material in prompts.
  • File writes happen only at the exact output_path the caller supplies (plus creation of missing parent directories).

5. Technology Architecture (Phase D)

5.1 Technology stack

Layer Choice Rationale
Language / runtime Python >= 3.11 Team standard, first-class MCP SDK
MCP framework mcp[cli] (FastMCP), pinned <2 Decorator-based tools, stdio transport built in
API client official openai SDK Maintained, typed
Package manager uv Fast, lockfile-based, reproducible
Transport stdio Launched and owned by the MCP host, no open ports, no daemon
Image model gpt-image-2, PNG output Current OpenAI image model; change in main.py if needed

5.2 Deployment model

The server is not deployed anywhere. Each MCP host process spawns its own instance on demand as a child process and terminates it with the session. There is nothing to operate, monitor, or patch besides the Python dependencies.

6. Security Architecture (cross-cutting)

  • Fail-fast startup: the process refuses to start without OPENAI_API_KEY, so misconfiguration is caught at registration time, not at first tool call.

  • No secrets at rest: config templates reference the key via ${OPENAI_API_KEY} or resolve it at launch from the macOS Keychain:

    security add-generic-password -s OPENAI_API_KEY -a "$USER" -w "sk-your-key"
    
  • Input validation: empty prompts rejected, size restricted to an allowlist, output_path must be absolute (prevents surprises from host-dependent working directories).

  • Network egress: exactly one destination, api.openai.com over HTTPS.

  • Residual risk: generate_image_to_file writes wherever the caller points it. The MCP host's permission system is the guard rail for that; review file paths when approving tool calls.

7. Implementation and Migration (Phases E and F)

7.1 Installation

git clone https://github.com/rammc/openai-image-mcp.git ~/openai-image-mcp
cd ~/openai-image-mcp
uv sync

7.2 Claude Code (project scope .mcp.json)

{
  "mcpServers": {
    "openai-image": {
      "type": "stdio",
      "command": "uv",
      "args": ["--directory", "${HOME}/openai-image-mcp", "run", "python", "main.py"],
      "env": { "OPENAI_API_KEY": "${OPENAI_API_KEY}" }
    }
  }
}

Keychain variant (macOS, no shell profile changes needed):

{
  "mcpServers": {
    "openai-image": {
      "type": "stdio",
      "command": "sh",
      "args": [
        "-c",
        "OPENAI_API_KEY=\"$(security find-generic-password -s OPENAI_API_KEY -w)\" exec uv --directory \"$HOME/openai-image-mcp\" run python main.py"
      ]
    }
  }
}

7.3 Claude Desktop (claude_desktop_config.json)

Claude Desktop does not inherit your shell environment, so use the Keychain variant above (same JSON, without the "type" field).

7.4 Verification

export OPENAI_API_KEY="sk-your-key"
uv run python main.py   # must start without error, then Ctrl+C

In Claude Code, run /mcp and confirm that openai-image is connected and lists two tools.

7.5 Migration notes

  • Coming from a copy of this server with project-specific paths: only the --directory argument changes, the tool contract is identical.
  • Model upgrades (for example a future gpt-image-3): single constant in main.py, no contract change expected.

8. Agent Usage Contract

This section is written for AI agents (Claude and others) so the repository is directly consumable without reading the source.

8.1 Tool contracts

generate_image(prompt: str, size: str = "auto") -> MCP Image (PNG)
generate_image_to_file(prompt: str, output_path: str, size: str = "auto") -> str

Constraints, enforced server-side with descriptive errors:

  • prompt must be non-empty after trimming. Write precise, visual prompts: subject, style, composition, color, background. The prompt is sent verbatim to OpenAI.
  • size must be one of: "auto", "1024x1024" (square), "1536x1024" (landscape), "1024x1536" (portrait). Any other value raises an error, including values like "512x512".
  • output_path must be an ABSOLUTE path ending in a filename (use .png). Missing parent directories are created automatically.

8.2 Decision rule: which tool

  • The user wants to SEE or iterate on the image in conversation: use generate_image.
  • The image is an artifact of the task (belongs in a repo, a build, a document): use generate_image_to_file with a path inside the project, then reference the file. Do not use generate_image and re-save the payload manually.

8.3 Behavioral notes for agents

  • Calls are synchronous and take several seconds up to about a minute. Do not retry immediately on slowness.
  • Each call costs real OpenAI credits. Batch thoughtfully, do not regenerate an acceptable image just to confirm.
  • Validation errors (for example "The prompt must not be empty.") mean the argument was wrong, not that the service failed. Fix the argument and retry once.
  • For transparent backgrounds or specific aspect ratios beyond the three sizes, state it in the prompt; the API decides best-effort.

8.4 Example calls

generate_image(
  prompt="Flat vector illustration of a mountain trail at sunrise, warm orange and teal palette, minimal, no text",
  size="1536x1024"
)

generate_image_to_file(
  prompt="App icon draft: stylized letter M, deep red gradient background, modern, centered, no text besides the letter",
  output_path="/Users/me/project/assets/icon-draft.png",
  size="1024x1024"
)

9. Operations and Troubleshooting

Symptom Cause Fix
Server fails to start: "OPENAI_API_KEY is missing" Key not in the launch environment Export the variable or use the Keychain launch command
Tool error: "size must be one of ..." Unsupported size value Use one of the four allowed values
Tool error: "output_path must be an absolute path." Relative path passed Pass an absolute path
401 from OpenAI Invalid or revoked key Rotate the key, update Keychain entry
Server not listed in /mcp Wrong --directory path or uv not installed Verify the path and uv --version, then restart the session

License and status

MIT License, see LICENSE. Version 0.1.0. Issues and extension requests: open a GitHub issue in this repo.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选