mcp-eval-harness
MCP-based code evaluation harness — sandboxed execution + LLM quality scoring.
README
mcp-eval-harness
MCP-based code evaluation harness — sandboxed execution + LLM quality scoring.
An MCP server that exposes 4 evaluation tools (no AI on the server itself — it's just a tool provider). Any MCP-compatible client can connect and use these tools: the included AI agent, Cursor, Claude Desktop, or anything that speaks MCP.
Tools
| Tool | Description |
|---|---|
run_code |
Execute code in a sandboxed Docker container (or local fallback). Returns stdout, stderr, exit code. |
run_tests |
Run code against test cases. Returns pass/fail for each test with actual vs expected output. |
score_answer |
LLM-based quality scoring (0-10) against a rubric with breakdown by correctness, efficiency, readability, edge cases. |
detect_hallucination |
Check if text contains claims not supported by provided context. Returns confidence and unsupported claims. |
Architecture
┌─────────────────────┐ MCP (stdio) ┌─────────────────────┐
│ │ ◄───────────────────── │ │
│ MCP Client │ tool calls │ MCP Server │
│ (AI agent) │ ────────────────────► │ (tool provider) │
│ │ results │ │
│ Has the brain — │ │ No AI here — │
│ decides what to │ │ just runs code, │
│ evaluate & when │ │ executes tests, │
│ │ │ calls scoring APIs │
└─────────────────────┘ └─────────────────────┘
Server = pure tool provider. Exposes run_code, run_tests, score_answer, detect_hallucination over MCP stdio. No decision-making, no orchestration.
Client = the brain. An AI agent (GPT-4o-mini) that decides which tools to call, in what order, and synthesizes a final evaluation report.
Transport = stdio. The client spawns the server as a subprocess and communicates via stdin/stdout. This is how MCP stdio works — same as how Cursor or Claude Desktop talk to MCP servers.
Use with Any MCP Client
The server isn't locked to our AI agent. You can use it with:
Cursor — Add to your MCP config:
{
"mcpServers": {
"eval-harness": {
"command": "npx",
"args": ["tsx", "server/index.ts"],
"cwd": "/path/to/mcp-eval-harness"
}
}
}
Claude Desktop — Add to claude_desktop_config.json:
{
"mcpServers": {
"eval-harness": {
"command": "npx",
"args": ["tsx", "server/index.ts"],
"cwd": "/path/to/mcp-eval-harness"
}
}
}
MCP Inspector (for debugging):
npx @modelcontextprotocol/inspector npx tsx server/index.ts
<!-- Yes, this uses Docker not Firecracker microVMs. It's a demo, not a bank. In prod, use E2B or self-hosted Firecracker. -->
Stack
| Package | Purpose |
|---|---|
@modelcontextprotocol/sdk |
MCP server + client protocol |
@ai-sdk/mcp |
Vercel AI SDK MCP client adapter |
ai |
Vercel AI SDK core — generateText, generateObject |
@ai-sdk/openai |
OpenAI provider (direct API key, NOT Vercel Gateway) |
zod |
Schema validation for tool inputs and structured outputs |
dotenv |
Environment variables |
Setup
git clone https://github.com/salmankhan-prs/mcp-eval-harness.git
cd mcp-eval-harness
pnpm install
cp .env.example .env
# Add your OpenAI API key to .env
Pull Docker images for sandboxed execution (optional — falls back to local if Docker isn't running):
docker pull node:22-alpine
docker pull python:3.12-alpine
Usage
pnpm eval examples/problems.json
Example Output
Evaluating: Two Sum
Language: javascript
============================================================
Test Results: 3/3 passed
Quality Score: 8.5/10
Correctness: 9/10
Efficiency: 9/10
Readability: 8/10
Edge Cases: 8/10
Hallucination Check: No hallucinations detected
Verdict: PASS
Project Structure
mcp-eval-harness/
├── README.md
├── package.json
├── tsconfig.json
├── .env.example
├── .gitignore
├── server/
│ ├── index.ts # MCP server — registers tools, starts stdio transport
│ └── tools/
│ ├── run-code.ts # Docker-sandboxed code execution (hacky but works)
│ ├── run-tests.ts # Test runner — executes code against test cases
│ ├── score-answer.ts # LLM-based quality scoring against rubric
│ └── detect-hallucination.ts # LLM-based hallucination detection
├── client/
│ └── index.ts # AI agent that orchestrates evaluation via MCP
└── examples/
└── problems.json # Sample coding problems with test cases
How Each Tool Works
run_code — Detects if Docker is available. If yes, runs code in a container with --network=none, memory/CPU/PID limits. If Docker is unavailable, falls back to local child_process with a 10s timeout. In production you'd swap this for a Firecracker microVM pool.
run_tests — Wraps the candidate solution + each test input into a runnable script. Calls run_code for each test case and compares stdout to expected output. Convention: solution must define a solution() function.
score_answer — Sends the problem + solution to GPT-4o-mini with a structured output schema (Zod). Returns an overall score and per-criterion breakdown. This is what ReasonCore's expert evaluators do, but automated.
detect_hallucination — Sends a claim + context to GPT-4o-mini. Returns whether unsupported claims exist, a confidence score, and specific unsupported claims.
Prerequisites
- Node.js 22+
- Docker Desktop running (optional — falls back to local execution)
- OpenAI API key in
.env
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。