vidtheque

vidtheque

Enables search across videos you've watched via transcripts, on-screen text, and frames, citing exact timestamps. Point it at videos, channels, or playlists; it indexes everything locally and answers queries with deep links to the exact second.

Category
访问服务器

README

vidtheque

Self-hosted MCP server that turns videos you've watched into a searchable multimodal corpus — transcript, on-screen text, and frame search with timestamped citations.

Point it at a video, a channel, or a playlist; it transcribes with word-level timestamps, OCRs what is on screen, embeds keyframes, and keeps the whole thing in a local index. Your MCP client then searches across everything you have ever indexed and cites answers with deep links that land on the exact second (https://youtu.be/ID?t=123).

Early development. Nothing here is stable yet: no releases, no published images, schemas and endpoints change without notice. The GPU worker skeleton and the deployment scaffolding are in; the MCP server itself is not written yet. Watch the repo rather than depending on it.

Architecture

Two services, one repo, HTTP between them — never a shared Python import.

                   ┌──────────────────────────────────────────┐
  MCP client       │ mcp/  (CPU, multi-arch — runs on a Pi)    │
  (Claude, …)      │                                          │
        │          │   OAuth ── yt-dlp ── pipeline orchestr.   │
        │  MCP     │       SQLite + sqlite-vec + FTS5          │
        └─────────▶│       keyframe JPEGs, job queue           │
                   └───────────────────┬──────────────────────┘
                                       │  HTTP — OpenAI shapes where they fit
                                       │  /v1/audio/transcriptions
                                       │  /v1/embeddings
                                       │  /v1/embeddings/image
                                       │  /v1/embeddings/frame-query
                                       │  /v1/ocr
                                       ▼
                   ┌──────────────────────────────────────────┐
                   │ worker/  (GPU, single box)               │
                   │                                          │
                   │   LifecycleManager — one job queue,      │
                   │   load-on-demand, idle-TTL unload,       │
                   │   NVML VRAM check, acquire/release hooks │
                   │                                          │
                   │   STTBackend        EmbedBackend         │
                   │   whisperX          Qwen3-Embedding-0.6B │
                   │                                          │
                   │   ImageEmbedBackend    OCRBackend        │
                   │   SigLIP 2 NaFlex      RapidOCR          │
                   └──────────────────────────────────────────┘

  optional: cloudflared ── tunnels the mcp service to a public hostname
            (compose profile `tunnel`)

The worker is a stateless inference API. No GPU? Skip the worker service entirely and point WORKER_URL at any OpenAI-compatible provider — the endpoints are the contract, not the implementation. Transcripts, metadata and OCR text are embedded by Qwen3-Embedding-0.6B (1024 dims); keyframes by SigLIP 2 NaFlex so400m (1152 dims), whose text tower is what turns a natural-language query into something comparable to a frame. Two spaces, never mixed, which is why they never share an endpoint.

Quickstart

git clone https://github.com/T0mSIlver/vidtheque.git
cd vidtheque/deploy
cp .env.example .env      # read it: every knob is documented there
docker compose up -d      # mcp + worker

With a Cloudflare tunnel for remote access:

TUNNEL_TOKEN=… docker compose --profile tunnel up -d

Check the worker:

curl localhost:8081/healthz
curl localhost:8081/status     # loaded models, VRAM, queue depth

Development

Requires uv and Python 3.12.

uv sync            # workspace: mcp + worker + dev tools (CPU-only deps)
make test          # pytest, CPU-only, no model downloads
make bench         # backend-vs-backend comparisons on your own hardware

Heavy inference dependencies live in the worker's [gpu] extra, so CI and a laptop checkout install cleanly without CUDA:

uv sync --extra gpu     # whisperX, sentence-transformers, transformers, RapidOCR

Layout

Path What
mcp/ MCP server (placeholder — framework choice pending)
worker/ GPU inference worker: FastAPI + backend registry + lifecycle
deploy/ docker-compose, .env.example, tunnel wiring
bench/ benchmark harness — backend comparisons on real hardware
research/ design notes and landscape survey (private-ish working notes)

License

MIT — see LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选
mcp-server-qdrant

mcp-server-qdrant

这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。

官方
精选
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选