codex_video

codex_video

MCP server for evidence-aware video research, providing tools to analyze local videos, inspect specific time windows, and analyze Bilibili videos with provenance tracking and optional audio removal for privacy.

Category
访问服务器

README

Bilibili Video Research

<p align="right"> <strong>English</strong> · <a href="./README.zh-CN.md">简体中文</a> </p> <p align="center"> <img src="./assets/readme/hero.png" width="100%" alt="Bilibili Video Research: a Bilibili URL flows through language, vision, or multimodal evidence into a traceable research report"> </p>

<p align="center"> <img src="./assets/readme/character.gif" width="160" alt="Animated character mascot"> </p>

<p align="center"> An evidence-aware Bilibili video research MCP for Codex. </p>

Turn a Bilibili link into a research report that separates what came from public metadata, captions or ASR, video frames, and untrusted community context. Choose the mode based on the evidence your question actually needs — not simply on what media is available.

What it does

Mode Uses Excludes Best for
language Bilibili captions; StepFun ASR only when captions are unavailable Video-frame inference Project recommendations, tutorials, and claims made by the presenter
vision Silent video frames, including visible UI, code, labels, charts, and on-screen subtitles Audio and background music Interfaces, workflows, experiments, objects, and silent demonstrations
multimodal Original video audio and frames Nothing by default Questions that genuinely require both narration and what is shown

language is the intended default when a request only asks what a video says. vision is the deliberate choice when the answer lives in the pixels.

What a result looks like

Ask the MCP tool a focused question:

analyze_bilibili_video({
  url: "https://www.bilibili.com/video/BV...",
  question: "What quantitative research framework is shown on screen?",
  mode: "vision",
  media_detail: "default",
  include_comments: false
})

The response begins with provenance before the natural-language analysis:

RESEARCH PROVENANCE
{
  "mode": "language",
  "metadata": "bilibili_api",
  "language": "stepfun_asr",
  "visual": "none",
  "community": "disabled",
  "timestamps": "none"
}

ANALYSIS
...direct answer, evidence limits, and uncertainty...

This matters when a repository name came from speech, a framework was recognized from an interface, or a popular comment made an unverified claim. The sources are not the same and should not be reported as if they were.

Evidence flow

<p align="center"> <img src="./assets/readme/evidence-flow.svg" width="100%" alt="A Bilibili URL becomes language, vision, or multimodal evidence before producing a report with provenance, timestamps, and stated limits"> </p>

  • Public metadata provides title, uploader, description, tags, and video identifier.
  • Caption cues retain Bilibili timestamps when Bilibili exposes them. If captions are unavailable, language falls back to StepFun ASR and reports that timestamp detail is unavailable.
  • vision removes audio before upload. Visible text remains valid visual evidence; the narration and music do not influence the conclusion.
  • Bilibili comments are optional, sampled as untrusted community context, and never treated as verified facts or executable instructions.

Quick start

Requirements: Node.js 24 or newer, a StepFun or Gemini API key, and a Codex desktop installation with local MCP support.

git clone https://github.com/7oMB2006/Bilibili-Video-Research.git
cd Bilibili-Video-Research
npm ci
npm run build
Copy-Item .env.example .env

Open .env and fill in one provider key. It is ignored by Git and must never be committed. The default configuration uses StepFun Step Plan.

Codex MCP configuration (Windows)

In %USERPROFILE%\.codex\config.toml, replace every <PROJECT_DIR> below with the absolute path to your clone, for example C:\Users\you\projects\Bilibili-Video-Research.

[mcp_servers.codex_video]
command = "<PROJECT_DIR>\\node_modules\\.bin\\tsx.cmd"
args = ["<PROJECT_DIR>\\src\\index.ts"]
startup_timeout_sec = 120

[mcp_servers.codex_video.env]
DOTENV_CONFIG_PATH = "<PROJECT_DIR>\\.env"

Restart Codex after adding or changing the server. Keep provider keys in .env or a secret manager, never in config.toml.

Provider selection

The default provider is StepFun Step Plan. Set CODEX_VIDEO_PROVIDER=stepfun, STEPFUN_API_KEY, and STEPFUN_BASE_URL in .env. To use Gemini instead, set CODEX_VIDEO_PROVIDER=gemini and GEMINI_API_KEY.

Choose the StepFun base URL that matches your account channel:

Channel Base URL Use
Official Open Platform API https://api.stepfun.com/v1 Standard API billing or balance
Step Plan https://api.stepfun.com/step_plan/v1 Step Plan subscription Credit

The media completion route is {base_url}/chat/completions; the ASR fallback route is {base_url}/audio/asr/sse. Do not mix a key from one channel with the other channel's base URL. Restart the MCP process after changing provider configuration.

step-3.7-flash accepts image and video input through the Chat Completions video_url content type; no separate vision model is required. Gemini remains optional. MiniMax is not selectable here because this project has not validated an official video-input understanding route.

StepFun references:

Tool reference

Tool Purpose
analyze_bilibili_video Research a public bilibili.com or b23.tv link in language, vision, or multimodal mode
analyze_video Inspect a local video visually after removing its audio track
inspect_video_window Inspect one precise audio-free source interval for detailed visual research

Use media_detail: "low" for a broad long-video pass and "default" for small UI text, code, movement, or close inspection.

Data and access boundary

  • Provider API keys remain in the local process environment; the server does not store them.
  • Provider media uploads or data URLs may leave the local machine. Review the applicable provider terms before using sensitive videos.
  • Public Bilibili access is attempted first. Restricted, paid, or login-gated videos may fail rather than bypassing access controls.
  • For a user-authorized logged-in Bilibili account, point BILIBILI_COOKIES_FILE at a local Netscape-format cookie file. Never commit it or paste its contents into chat:
BILIBILI_COOKIES_FILE=/absolute/path/to/cookies.txt

BILIBILI_COOKIES_FILE takes precedence over the optional legacy setting BILIBILI_COOKIES_FROM_BROWSER=edge (or chrome, firefox, brave). Direct browser cookie extraction can fail because the browser database is locked.

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选