Framesleuth
Local bug-reproduction video analysis tool that processes bug recordings frame-by-frame, producing a structured Bug Context Bundle, and exposes this capability over MCP for use by coding agents like VS Code and Claude.
README
<p align="center"> <img src="docs/logo/framesleuth-logo-256.png" alt="Framesleuth" width="128" height="128" /> </p>
Framesleuth
Local bug-reproduction video analysis, exposed over MCP.
Framesleuth takes a bug-recording video (plus optional browser sidecars), understands it frame-by-frame, and produces a structured Bug Context Bundle. It is MCP-ready, so any MCP client — a VS Code agent, another coding agent, or a custom system — can drive the analysis and consume the result to fix the bug directly.
Capture happens in a separate Chrome extension, inkwell, which records the bug and posts the video + sidecars to this agent's local API. This repo is the analysis agent only.
Everything runs locally. Nothing leaves your machine.
Quick start
Want to fix a bug from a video inside VS Code? Follow Use with VS Code & Claude (MCP) — connect the bundled MCP server and go from a recording to a grounded fix.
Try it end-to-end in 2 minutes (no models required)
Framesleuth degrades gracefully: with no vision model or ffmpeg installed, it still produces a Bug Context Bundle from the browser sidecars (console errors, failed network requests, clicks). This is enough to record a bug in Chrome and get a structured report.
uv venv && source .venv/bin/activate
uv pip install -e ".[dev]"
# 1. Start the backend (binds 127.0.0.1:8010 from config; 8000 is left for inkwell)
framesleuth-api # or: uvicorn framesleuth.service.api:app --port 8010
# 2. Install the inkwell Chrome extension (separate repo) to capture a bug:
# https://github.com/santoshshinde2012/inkwell
# Record → reproduce the bug → Stop. It posts the video + sidecars here.
# 3. Or drive the agent directly with your own video file — over the HTTP API
# (see the Postman collection / runbook) or the videobug MCP server
# (see docs/use-with-vscode-and-claude.md).
Check GET /v1/healthz: status is healthy when the vision + coder models
are up, or unhealthy when running sidecar-only (with vlm/coder reported
unavailable and storage ready). Add the model servers (below) to enable
frame-level OCR/visual understanding.
Prerequisites (full pipeline)
- Python 3.11+
- A local VLM server (llama.cpp
:8080or Ollama:11434) for frame understanding - 8GB+ RAM (for models)
ffmpeg is not a separate prerequisite — frame/audio decoding uses PyAV, which bundles its own ffmpeg libraries. (
ffprobe, if present, is used opportunistically to detect whether a recording has an audio stream.)
Setup
# Clone and navigate
git clone https://github.com/santoshshinde2012/framesleuth.git
cd framesleuth
# Create environment and install
uv venv && source .venv/bin/activate
uv pip install -e ".[dev]"
# Download models (one-time, ~10-20GB)
python scripts/download_models.py
# Copy environment template
cp .env.example .env
# Edit .env to match your setup (see runbook.md for details)
# Start services (Docker Compose). Run it — do NOT `source` it.
./scripts/dev_up.sh
# No Docker? Skip this and use the Ollama path in "Start & stop the stack" below.
Start & stop the stack (Ollama — verified working)
scripts/dev_up.sh uses Docker Compose. If you don't run Docker, use the
engine-agnostic Ollama path below. .env.example ships the llama.cpp defaults
(VLM_URL=http://127.0.0.1:8080, VLM_MODEL=Qwen/Qwen3-VL-8B-Instruct-GGUF); for
the Ollama path, copy it to .env and set VLM_URL=http://127.0.0.1:11434 and
VLM_MODEL=qwen2.5vl. The vision model is the
piece that powers frame-level understanding; without it reachable, analyses come
back degraded (no on-screen evidence read).
Start
# 1. Vision model — start Ollama and pull the VLM once (~6 GB; or qwen2.5vl:3b)
ollama serve & # skip if Ollama is already running
ollama pull qwen2.5vl
# 2. Backend — run from the repo root so it loads .env (binds 127.0.0.1:8010)
source .venv/bin/activate
framesleuth-api # or: uvicorn framesleuth.service.api:app --port 8010
# 3. Verify BEFORE recording — both should report ready
curl -s http://127.0.0.1:11434/v1/models | grep -q qwen2.5vl && echo "VLM ready"
curl -s http://127.0.0.1:8010/v1/healthz # expect status: healthy, vlm: ready
When /v1/healthz shows vlm: ready, recordings analyze with a real
classification (analysis_quality.level = full/partial). If the VLM is down,
you get the degraded "evidence was thin" report instead. Record with narration
so the audio transcript (asr) stage contributes too.
Stop
# Stop the backend: Ctrl+C in its terminal, or
pkill -f framesleuth-api
# Stop Ollama (optional — leaving it running keeps the model warm)
pkill -f "ollama serve" # macOS app users: quit Ollama from the menu bar
Architecture
Bug video (mp4/webm) + sidecars
↓
Local Analysis Service (pipeline)
├─ Preprocess (PyAV: duration/fps/dims)
├─ Transcript (faster-whisper)
├─ Keyframes (visual-delta change scoring)
├─ Understanding (Qwen3-VL)
├─ Fusion + Classification
├─ Extraction → Bug Context Bundle
├─ Summarize (skill/system-prompt-driven)
└─ Grounding (workspace search)
↓
Bug Context Bundle
↓
MCP server + local HTTP API
└─ consumed by any MCP client (VS Code agent, other agents, inkwell extension)
Features
- Frame-by-frame understanding using Qwen3-VL vision model
- Automatic keyframe selection via frame-to-frame visual-delta change scoring
- Error detection and extraction from console, OCR, and UI state
- Redaction-first design — sensitive data (passwords, tokens) redacted before models see it
- No data leaves your machine — fully local, no telemetry or cloud APIs
- Engine-agnostic — swap Ollama, llama.cpp, or vLLM via config only
- Structured output — canonical Bug Context Bundle with evidence citations
- Configurable response — pick a summary skill and an action mode
(
fix/explain/triage/test/report/reproduce, auto-picked from the classification), plus a machine-readablesuggested_actionsmenu and on-demand artifact renderers (markdown / GitHub issue / test plan) - Resilient — handles no-audio videos, weak local models, low-confidence cases
Project structure
framesleuth/
├── framesleuth/ # Main package
│ ├── config.py # Typed config (pydantic-settings)
│ ├── schemas.py # Data contracts (Bug Context Bundle, enums)
│ ├── errors.py # Exception taxonomy
│ ├── logging_config.py # Structured JSON logging, job-id correlation
│ ├── prompts.py # VLM / classify / summary / fix prompt templates
│ ├── skills.py # Built-in summary skills (summary, bug_report, ...)
│ ├── actions.py # Action modes (fix/explain/triage/...) + suggested-actions menu
│ ├── render.py # Artifact renderers (markdown / GitHub issue / test plan)
│ ├── clients/ # VLM, coder HTTP clients (OpenAI-compatible)
│ ├── pipeline/ # preprocess, asr, scenes, understand, fusion, classify, bug_extract, redact, summarize, sidecars, grounding
│ ├── orchestrator/ # graph.py — linear async stage pipeline
│ ├── jobs/ # store.py — SQLite job state + bundle index
│ ├── service/ # FastAPI HTTP endpoints
│ └── mcp_server/ # videobug MCP server (VS Code + any MCP client)
├── tests/ # pytest tests + fixtures
├── scripts/ # download_models.py, dev_up.sh
├── postman/ # HTTP API collection + environment
├── docs/ # capabilities, use-with-vscode-and-claude, web-integration
└── pyproject.toml # Dependencies and tool config
Development
Run tests
pytest tests/ -v --cov=framesleuth
Code quality
ruff check framesleuth tests
black --check framesleuth tests
mypy --strict framesleuth
Set up pre-commit hooks
pre-commit install
A short, focused set:
- Capabilities — the single reference: every input, output, skill, action, renderer, HTTP endpoint, and MCP tool
- Use with VS Code & Claude (MCP) — connect the
videobugMCP server to Copilot, Claude Code, and Claude Desktop - Web App Integration (end-to-end) — embed Framesleuth behind your own backend with an agent loop
- Postman Collection — exercise the HTTP API end-to-end (import or run headless with Newman)
- Runbook & Troubleshooting — setup, health checks, and common issues
License
Apache-2.0
Capture client (inkwell)
Bug capture lives in a separate Chrome extension,
inkwell. It records the bug, collects
browser sidecars (console errors, failed requests, clicks), and posts the video + sidecars
to this agent's local API. The agent's CORS is already scoped to chrome-extension://
origins and the loopback bind, so inkwell works against a locally running backend with no
extra setup.
Status: Backend + pipeline + MCP server completed.
Questions? Open an issue or check runbook.md for common questions.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。