Sensory-Grounding MCP
Gives AI coding agents a closed-loop verification cycle for visual, audio, and video output, with enforcement hooks that make verification mandatory.
README
Sensory-Grounding MCP
Give your AI coding agents real eyes and ears — and stop them from telling you "it's done" when it isn't.
A Model Context Protocol server that gives Claude Code, Codex, Antigravity, and any other MCP-compatible coding agent a closed-loop verification cycle for visual, audio, and video output — plus the enforcement hooks that make verification mandatory, not optional.
Built by Luclino Cruz for LuCross, a one-person, AI-agent-operated holding company, as part of the internal harness that governs every agent in its ecosystem. Released publicly because it solves a problem every agentic coding setup runs into, and there was no reason to keep it private.
The problem this solves
You ask an agent to fix a layout bug, re-render a thumbnail, or clean up a narration track. It edits the file, the tool call returns with no error, and the agent reports: "Done — looks great now."
Except it doesn't. The heading still overflows. The narration still clips. The video still has the wrong aspect ratio. The absence of an error is not proof of correctness — and an LLM has no way to see its own output unless you explicitly hand it eyes.
This is not a prompting problem. Telling a model "please actually check your work" in a system prompt does not work reliably — the research is unambiguous on this (Huang et al., ICLR 2024; Gou et al., CRITIC, ICLR 2024): self-correction without an external, deterministic signal degrades performance rather than improving it. What actually works is forcing a verifier into the loop — something outside the model's own judgment that the model cannot talk its way around.
How it works
Three layers, each doing exactly one job:
┌──────────────────────────────────────────────────────────────┐
│ ENFORCEMENT — Claude Code hooks (outside the agent's control) │
│ PostToolUse (Write|Edit) → flags the file as pending review │
│ Stop → blocks turn completion (exit code 2) until resolved │
└───────────────────────────┬──────────────────────────────────┘
│
┌────────────────────────────▼──────────────────────────────────┐
│ CRITIC — the agent itself, in the same session │
│ Looks at the screenshot/spectrogram/frames the sensor returns │
│ Assigns a 0-100 score against the task's success criteria │
└────────────────────────────┬──────────────────────────────────┘
│
┌────────────────────────────▼──────────────────────────────────┐
│ SENSORS — 3 MCP tools, stateless, deterministic │
│ visual_inspector · acoustic_inspector · video_inspector │
└──────────────────────────────────────────────────────────────┘
The critical piece most "AI eyes" projects skip is the enforcement layer.
A visual_inspector tool that the agent can call but doesn't have to
call changes nothing — an agent under time pressure will skip it and report
success anyway. This project wires the verification cycle into Claude Code's
own hook system so the agent cannot end its turn with an unresolved,
un-improved visual/audio/video change. No new orchestrator, no external
critic model, no extra API cost beyond what you already pay for the agent
itself.
The 6 tools
| Tool | Layer | What it does |
|---|---|---|
visual_inspector |
Sensor | Renders HTML/CSS/SVG via Playwright, or analyzes PDF/DOCX/PPTX layout via Docling. Returns a screenshot + a structured layout report. Content-hash cached. |
acoustic_inspector |
Sensor | Spectrogram + RMS/clipping/duration telemetry via librosa. Optional audio_pt profile flags silences >2s. |
video_inspector |
Sensor | Key-frame sampling via ffmpeg scene-detection — cheaper and more effective than asking a model to "watch" a whole video. |
log_pending |
State | Records that a file needs inspection. Called by the PostToolUse hook, not the agent. |
mark_inspected |
State | The agent calls this after looking at a sensor's output and scoring it 0-100. Refuses to mark a pass unless the new score is strictly higher than the file's last known score — the Forced Optimization rule (see below). |
finalize_task |
State | Checked by the Stop hook. Any file with an unresolved or regressed inspection blocks turn completion. |
Forced Optimization
Borrowed from ReLook (Li et al., Tencent, ACL 2026):
an edit is only accepted if score_new > score_previous for that exact file.
A regression that "at least didn't make it much worse" is not accepted —
the agent must revert or try a different approach. This turns iteration into
a monotonically improving trajectory instead of a random walk.
Quick start
git clone https://github.com/luclinocruz/lucross-sensory-grounding-mcp.git
cd lucross-sensory-grounding-mcp
python -m venv .venv
.venv/Scripts/activate # or: source .venv/bin/activate on Linux/macOS
pip install -r requirements.txt
playwright install chromium
There are two ways to install this — pick based on how broadly you want the gate to apply. Both use the exact same server and hook script; only the config differs.
Option A — Project-scoped (one repo, explicit root)
Best for: shared/team repos, or any project where you want the sandbox root spelled out with zero ambiguity.
.mcp.json in your project root:
{
"mcpServers": {
"sensory-grounding": {
"command": "/absolute/path/to/lucross-sensory-grounding-mcp/.venv/bin/python",
"args": ["/absolute/path/to/lucross-sensory-grounding-mcp/server.py"],
"cwd": "/absolute/path/to/your/project",
"env": { "SGS_SANDBOX_ROOT": "/absolute/path/to/your/project" }
}
}
}
.claude/settings.json in the same project:
{
"hooks": {
"PostToolUse": [{
"matcher": "Write|Edit",
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage post-edit --root /absolute/path/to/your/project" }]
}],
"Stop": [{
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage on-stop --root /absolute/path/to/your/project" }]
}]
}
}
With SGS_SANDBOX_ROOT and --root both set, this behaves exactly as
described above — fixed root, no ambiguity, easy to reason about in a repo
other people also work in.
Option B — Global / user-scope (one install, protects every project)
Best for: your own machine, where you want the gate to apply automatically to any project you open with Claude Code, without configuring each one.
claude mcp add sensory-grounding --scope user -- \
/absolute/path/to/lucross-sensory-grounding-mcp/.venv/bin/python \
/absolute/path/to/lucross-sensory-grounding-mcp/server.py
Then in your global ~/.claude/settings.json (not a per-project one):
{
"hooks": {
"PostToolUse": [{
"matcher": "Write|Edit",
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage post-edit" }]
}],
"Stop": [{
"hooks": [{ "type": "command",
"command": "/absolute/path/to/.venv/bin/python /absolute/path/to/enforce_inspect.py --stage on-stop" }]
}]
}
}
Note there's no --root and no SGS_SANDBOX_ROOT in this mode. Without
them, the server and the hook both fall back to resolving the sandbox root
from the current session's working directory — the server via Path.cwd()
at the time it's invoked, the hook via the cwd field Claude Code includes
in every PostToolUse/Stop payload. In practice this means: whichever
project directory a given Claude Code session is running in becomes that
session's sandbox root, automatically, with the exact same path-containment
check as Option A — nothing outside that directory is ever reachable.
Verify it after installing, once, in two different project directories —
confirm a file edit in project A doesn't show up as pending in project B,
and that the Stop hook blocks/releases correctly in each. The cwd-based
fallback is the intended mechanism, but exact subprocess cwd inheritance
can vary slightly across Claude Code versions/platforms, so this is worth
one empirical check on your setup rather than trusting it blindly.
Trade-off to know before you flip this on globally: it now fires in
any directory, for any task — including work that has nothing to do
with the project you built this for. A throwaway edit to a .png in a
random folder will trigger the same block as a real regression. If that
gets noisy, Option A (installed only in the repos that matter to you) avoids
the false positives entirely.
Either way: any Write/Edit touching a tracked extension
(.html .css .svg .pdf .docx .pptx .wav .mp3 .m4a .mp4 .mov) now requires a
resolved, improved inspection before the agent can end its turn.
Tell the agent what's expected of it
Add to your project's CLAUDE.md (or equivalent system prompt):
Any change to a visual/audio/video file ends with calling the matching sensor tool, comparing the result against the task's success criteria, and calling
mark_inspectedwith a 0-100 score. Only accept your own change if the new score is strictly higher than the previous one — otherwise revert and try a different approach. Never report a task complete without this.
State model
State is scoped per project, not per session — deliberately. An agent
has no reliable way to know its own Claude Code session ID, so pending
inspections live in <project>/.sgs_inspection_log.json, keyed by absolute
file path, with a persistent per-file score baseline. This also means the
gate is meaningfully project-level: "this project has an open, unresolved
visual regression" is a fact about the project, not about one conversation.
No SQLite, no external database — one JSON file, human-readable, git-ignored by default.
Why not just prompt the model harder?
Because that's exactly the failure mode the underlying research describes.
Self-correction without a deterministic, external verifier reliably
underperforms not correcting at all in some tasks (Huang et al.).
The fix isn't a better-worded instruction — it's a verifier the model
literally cannot argue its way past. That's what the Stop hook's exit code
2 gives you: Claude Code refuses to end the turn and hands the reason back to
the model, verbatim.
Design notes
- The critic is the agent itself — not a separate paid vision model. The
sensor returns an
ImageContentblock; whichever model is already running the session evaluates it, inside the same context. Zero marginal API cost. - Sensors are stateless and know nothing about each other. Swap the rendering engine, the audio backend, anything — the interface doesn't change.
- Sandbox containment is never optional, only its root is configurable.
Project-scoped installs pin
SGS_SANDBOX_ROOTexplicitly; global installs derive it from the current session's working directory instead. Either way, the containment check itself is identical and always enforced — there is no mode where an arbitrary path outside the resolved root is reachable.
License
Apache License 2.0 — see LICENSE. Use it, fork it, ship it inside a commercial product, whatever you need. If you build something interesting on top of it, a mention is appreciated but never required.
Author
Luclino Cruz — github.com/luclinocruz
Built as part of the harness for LuCross, a solo-operated, AI-agent-driven company. If this saved you a debugging session, a star helps other people find it.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。