vidtheque
Enables search across videos you've watched via transcripts, on-screen text, and frames, citing exact timestamps. Point it at videos, channels, or playlists; it indexes everything locally and answers queries with deep links to the exact second.
README
vidtheque
Self-hosted MCP server that turns videos you've watched into a searchable multimodal corpus — transcript, on-screen text, and frame search with timestamped citations.
Point it at a video, a channel, or a playlist; it transcribes with word-level
timestamps, OCRs what is on screen, embeds keyframes, and keeps the whole thing
in a local index. Your MCP client then searches across everything you have ever
indexed and cites answers with deep links that land on the exact second
(https://youtu.be/ID?t=123).
Early development. Nothing here is stable yet: no releases, no published images, schemas and endpoints change without notice. The GPU worker skeleton and the deployment scaffolding are in; the MCP server itself is not written yet. Watch the repo rather than depending on it.
Architecture
Two services, one repo, HTTP between them — never a shared Python import.
┌──────────────────────────────────────────┐
MCP client │ mcp/ (CPU, multi-arch — runs on a Pi) │
(Claude, …) │ │
│ │ OAuth ── yt-dlp ── pipeline orchestr. │
│ MCP │ SQLite + sqlite-vec + FTS5 │
└─────────▶│ keyframe JPEGs, job queue │
└───────────────────┬──────────────────────┘
│ HTTP — OpenAI shapes where they fit
│ /v1/audio/transcriptions
│ /v1/embeddings
│ /v1/embeddings/image
│ /v1/embeddings/frame-query
│ /v1/ocr
▼
┌──────────────────────────────────────────┐
│ worker/ (GPU, single box) │
│ │
│ LifecycleManager — one job queue, │
│ load-on-demand, idle-TTL unload, │
│ NVML VRAM check, acquire/release hooks │
│ │
│ STTBackend EmbedBackend │
│ whisperX Qwen3-Embedding-0.6B │
│ │
│ ImageEmbedBackend OCRBackend │
│ SigLIP 2 NaFlex RapidOCR │
└──────────────────────────────────────────┘
optional: cloudflared ── tunnels the mcp service to a public hostname
(compose profile `tunnel`)
The worker is a stateless inference API. No GPU? Skip the worker service
entirely and point WORKER_URL at any OpenAI-compatible provider — the
endpoints are the contract, not the implementation. Transcripts, metadata and
OCR text are embedded by Qwen3-Embedding-0.6B (1024 dims); keyframes by
SigLIP 2 NaFlex so400m (1152 dims), whose text tower is what turns a
natural-language query into something comparable to a frame. Two spaces, never
mixed, which is why they never share an endpoint.
Quickstart
git clone https://github.com/T0mSIlver/vidtheque.git
cd vidtheque/deploy
cp .env.example .env # read it: every knob is documented there
docker compose up -d # mcp + worker
With a Cloudflare tunnel for remote access:
TUNNEL_TOKEN=… docker compose --profile tunnel up -d
Check the worker:
curl localhost:8081/healthz
curl localhost:8081/status # loaded models, VRAM, queue depth
Development
Requires uv and Python 3.12.
uv sync # workspace: mcp + worker + dev tools (CPU-only deps)
make test # pytest, CPU-only, no model downloads
make bench # backend-vs-backend comparisons on your own hardware
Heavy inference dependencies live in the worker's [gpu] extra, so CI and a
laptop checkout install cleanly without CUDA:
uv sync --extra gpu # whisperX, sentence-transformers, transformers, RapidOCR
Layout
| Path | What |
|---|---|
mcp/ |
MCP server (placeholder — framework choice pending) |
worker/ |
GPU inference worker: FastAPI + backend registry + lifecycle |
deploy/ |
docker-compose, .env.example, tunnel wiring |
bench/ |
benchmark harness — backend comparisons on real hardware |
research/ |
design notes and landscape survey (private-ish working notes) |
License
MIT — see LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。