media-context-mcp

media-context-mcp

Give your AI assistant eyes and ears — analyze any video, audio, or image, entirely on your machine.

Category
访问服务器

README

<p align="center"> <img src="./assets/banner.svg" alt="media-context-mcp — local MCP server to analyze video, audio and images for AI assistants" width="100%"> </p>

<p align="center"> <a href="https://www.npmjs.com/package/media-context-mcp"><img src="https://img.shields.io/npm/v/media-context-mcp.svg" alt="npm"></a> <a href="https://github.com/vishalguptax/media-context-mcp/actions/workflows/ci.yml"><img src="https://github.com/vishalguptax/media-context-mcp/actions/workflows/ci.yml/badge.svg" alt="CI"></a> <a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-blue.svg" alt="license"></a> </p>

<p align="center"> Give your AI assistant eyes and ears — analyze any <b>video, audio, or image</b>, entirely on your machine. </p>

<p align="center"> <a href="#-install">Install</a> · <a href="#-examples">Examples</a> · <a href="#-tools">Tools</a> · <a href="./docs/usage.md">Usage guide</a> · <a href="https://www.npmjs.com/package/media-context-mcp">npm</a> · <a href="https://lobehub.com/mcp/vishalguptax-media-context-mcp">LobeHub</a> </p>


Your assistant can read text and look at a picture, but it can't watch a video or listen to audio. media-context-mcp fills that gap: point it at a file or a URL and it hands back clean, model-ready context — sampled frames, a transcript, or the text on screen — without sending anything to the cloud.

# 1. add it to your client (Claude Code shown — see Install for others)
claude mcp add media-context -- npx -y media-context-mcp
# 2. install the media tools it uses
npx media-context-mcp setup

Then just ask: “Summarize demo.mp4.”

✨ Features

  • Any source — video, audio, or images; a local file or a URL (YouTube, Vimeo, and 1000+ more).
  • See video — a quick montage overview, full-resolution stills, scene-change shots, or a dense filmstrip that catches glitches lasting a fraction of a second.
  • Hear audio — turn speech in a clip, voice note, or podcast into text.
  • Read screens — pull the exact text off a UI, an error dialog, or a screenshot.
  • Cheap by design — frames are tiled and downscaled, so a long clip costs a couple of images, not hundreds.
  • Private & local — runs on your machine. No API keys, no uploads.
  • Works everywhere — any MCP client: Claude, Cursor, VS Code, and more.

🚀 Install

1. Add the server to your client

The launch command is always npx -y media-context-mcp.

<details open> <summary><b>Claude Code</b></summary>

claude mcp add media-context -- npx -y media-context-mcp

Or install it as a plugin (one command, bundles the server):

/plugin marketplace add vishalguptax/media-context-mcp
/plugin install media-context

</details>

<details> <summary><b>Claude Desktop</b></summary>

Settings → Developer → Edit Config (claude_desktop_config.json):

{
  "mcpServers": {
    "media-context": { "command": "npx", "args": ["-y", "media-context-mcp"] }
  }
}

</details>

<details> <summary><b>Cursor</b> · <b>Windsurf</b> · <b>Cline</b> · other clients</summary>

Add to the client's MCP config (~/.cursor/mcp.json, ~/.codeium/windsurf/mcp_config.json, Cline settings, …):

{
  "mcpServers": {
    "media-context": { "command": "npx", "args": ["-y", "media-context-mcp"] }
  }
}

</details>

<details> <summary><b>VS Code (GitHub Copilot, agent mode)</b></summary>

Create .vscode/mcp.json — VS Code uses the servers key:

{
  "servers": {
    "media-context": { "command": "npx", "args": ["-y", "media-context-mcp"] }
  }
}

</details>

<details> <summary><b>Codex CLI</b></summary>

~/.codex/config.toml:

[mcp_servers.media-context]
command = "npx"
args = ["-y", "media-context-mcp"]

</details>

Global vs per-project — install once for all projects, or commit a project-scoped config so your team shares it: claude mcp add … --scope project (writes .mcp.json), or a .cursor/mcp.json / .vscode/mcp.json in the repo.

2. Install the media tools

One command installs what the server uses, via your OS package manager:

npx media-context-mcp setup          # ffmpeg, URL download, on-screen text
npx media-context-mcp setup --audio  # also enable transcription

The server then finds the tools automatically — including common off-PATH spots (Tesseract in Program Files, Whisper in a Python Scripts folder), so transcripts and OCR work without extra config. check_media_deps shows what's ready; setup --uninstall removes the tools again.

<details> <summary>Install by hand / point at a custom path</summary>

The package ships no binaries. Only ffmpeg is required; the rest are optional, one feature each.

Tool For Install
ffmpeg + ffprobe required winget install Gyan.FFmpeg · brew install ffmpeg · apt install ffmpeg
yt-dlp URLs winget install yt-dlp.yt-dlp · brew install yt-dlp · pip install -U yt-dlp
tesseract on-screen text winget install UB-Mannheim.TesseractOCR · brew install tesseract · apt install tesseract-ocr
whisper transcription pip install -U openai-whisper

If a tool lives somewhere unusual, point at it with FFMPEG_BIN / YTDLP_BIN / WHISPER_BIN / TESSERACT_BIN (env vars, e.g. in your client's config env block). </details>

💬 Examples

Just ask your assistant in plain language — it picks the right options for you.

  • “Summarize demo.mp4.” — a quick overview from sampled frames.
  • “What error does the app show at the end of bug.mp4?” — reads the on-screen text.
  • “Transcribe standup.m4a and list the action items.” — speech to text.
  • “Summarize https://youtu.be/VIDEO_ID and include the transcript.” — fetches and transcribes.
  • “In slider.mp4, find the frame where the slider flickers around 0:06.” — catches a sub-second glitch.

Finer control — modes, cropping, language, sampling rate — is in the usage guide.

🧰 Tools

The server exposes two tools, which your assistant calls automatically.

Tool What it does
analyze_media Turn a video, audio, or image — file or URL — into model-readable context. Auto-detects the type: video → frames, stills, scene montages, or a dense filmstrip; audio → a transcript; image → the picture plus optional text recognition. Supports cropping, time windows, language, and sampling rate.
check_media_deps Report which optional capabilities (URL fetching, transcription, text recognition) are ready, with setup hints.

Everything runs locally, and each call cleans up its temporary files when it returns.

❓ FAQ

Can Claude (or any LLM) watch a video? Not directly — models take images and text, not video. This server extracts frames and transcripts so your assistant can analyze it.

How do I give Claude Code, Cursor, or VS Code video context? Add the server (see Install), then ask in plain language — it works in any MCP client.

Can it convert video or audio to text? Yes — it samples frames for the model to read and transcribes speech locally.

Does it work offline, without an API key? Yes. Everything runs on your machine; nothing is uploaded and no keys are required.

Does it support YouTube and other links? Yes — any yt-dlp-supported URL.

Is it free? Yes, open source under Apache-2.0.

🛠️ Development

npm install
npm run build
npm test

Tests cover the pipeline end-to-end; the integration ones skip themselves when the optional tools aren't installed. Issues and PRs welcome.

📄 License

Apache-2.0 © Vishal Gupta — free and open, use it however you like.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选