evalgate
Eval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a benchmark number that won't survive a second look.
README
evalgate
Cheap statistical checks for AI eval claims — run them before you publish.
Most benchmark headlines overstate themselves in one of a few nameable ways. evalgate is four tiny, dependency-free checks, one per failure mode — the same checks behind a set of independent eval-integrity audits that caught these mistakes in published work.
Pure Python, zero dependencies, runs anywhere.
pip install git+https://github.com/ipezygj/evalgate
Use it as an MCP tool (for agents)
If you're an AI agent — or you run one (Claude, Cursor, Claude Code, Windsurf…) — evalgate ships an MCP server so the model can call these checks itself before it trusts, reports, or acts on any eval number: a benchmark score, a leaderboard #1, an LLM-as-judge verdict, or a claimed trend.
pip install "eval-integrity[mcp] @ git+https://github.com/ipezygj/evalgate" # once on PyPI: pip install "eval-integrity[mcp]"
Add it to your MCP client (e.g. claude_desktop_config.json / Cursor / Claude Code):
{
"mcpServers": {
"evalgate": { "command": "evalgate-mcp" }
}
}
<sub>MCP registry identity — mcp-name: io.github.ipezygj/evalgate</sub>
Tools the agent gets, each with a "call this when…" description it can reason about:
| tool | the agent calls it before… |
|---|---|
check_top_rank |
claiming a model is #1 / SOTA on a benchmark (is the top rank real or a tie?) |
check_subset_win |
trusting a "best on subset/metric X" claim (look-elsewhere correction) |
check_judge_bias |
trusting an LLM-as-judge / A-B preference result (length / self-preference / position bias) |
check_resolution |
calling one of two close models better (can the benchmark even tell them apart?) |
check_trend_fragility |
reporting a fitted trend / scaling exponent (does one point carry it?) |
The five checks above work on summary numbers (scores, p-values, win counts). If the agent has the raw per-item results — which items each model solved, or the head-to-head battles — three deeper tools do the real audit instead of an approximation:
| tool | the agent calls it when… |
|---|---|
audit_leaderboard |
it has per-item results ({model: [solved item-ids]}) — the real version of check_top_rank: bootstrapped rank confidence intervals, the paired-McNemar tie group, resolvable tiers, split-half stability |
audit_preferences |
a ranking comes from pairwise votes ([winner, loser] battles) — Bradley-Terry rank CIs and a Condorcet check that preferences are transitive, not rock-paper-scissors cycles |
check_dimensions |
deciding whether one number fairly summarizes a multi-skill benchmark — counts the latent skills in the result matrix (eigenspectrum vs a shuffled null) |
The point: an agent that produces an eval number should sanity-check it, and now it can — in one
call, with a plain verdict and a recommendation. Reproducible, zero-dependency checks (the mcp
extra is only for the server transport).
The four checks
1. "We lead on subset X" — corrected for look-elsewhere
Report the subset/metric/checkpoint where a model looks best and you are reporting the maximum of many noisy tests. Correct for how many you could have picked.
evalgate correct --p 0.009 --n 23
# raw p=0.009 over 23 tests -> sidak p=0.19 (does NOT survive correction at alpha=0.05)
from evalgate import correct_best_of
correct_best_of(0.009, n_tested=23).significant # False
(A real RewardBench "best subset" win: raw p=0.009 → p=0.19 after correcting for the 23 subsets. Not a finding.)
2. Is the judge winning, or just longer / first / same-family?
An LLM-as-judge that "prefers" your model may be preferring the longer answer, the first-listed one, or its own family. Feed it the count and test against chance.
evalgate bias --wins 68 --n 100 --label "longer answer wins"
# longer answer wins: 68/100 = 68.0% (p=0.0004) -> BIAS
from evalgate import bias_rate
bias_rate(68, 100).biased # True
(A widely-used GPT-4 judge preferred the longer answer 68% of the time and its own model family 71.5% — both at p≈0.)
3. Does one data point flip your slope?
A scaling exponent or trend that hangs on a single high-leverage point isn't one. Leave each point out and refit.
evalgate loo examples/points.txt --power-law --threshold 1.0
# slope=1.08, leave-one-out range [0.87, 1.26] -> CROSSES 1 (most influential point: index 5)
from evalgate import leave_one_out, power_law_exponent
leave_one_out(xs, ys, fit=power_law_exponent, threshold=1.0).crosses_threshold # True
(A reported "super-linear" grokking exponent, α=1.13, fell to 0.97 — with a better fit — when one point was dropped.)
4. Is the gap bigger than the sample can resolve?
A leaderboard orders two models by a two-point accuracy gap on a finite test set. Ask whether that gap is even detectable at this sample size — or smaller than the minimum detectable effect, i.e. a coin flip.
evalgate power --n 200 --p1 0.85 --p2 0.83
# gap=+0.02 on n=200 (NOT significant, p=0.585); MDE at 80% power=0.103 -> UNDERPOWERED (gap < MDE)
from evalgate import power_check
power_check(200, 0.85, 0.83).resolvable # False — 2pp over 200 items can't be resolved
power_check(2000, 0.85, 0.80).resolvable # True — 5pp over 2000 items can
(Frontier models on a fixed benchmark routinely sit a task or two apart — inside the MDE — so the #1 rank is noise. More votes, not a better model, resolves the tie.)
Library API
from evalgate import (
correct_best_of, sidak, bonferroni, # look-elsewhere
bias_rate, binomial_test, # judge / metric bias
leave_one_out, ols_slope, power_law_exponent, # fragility + fits
power_check, min_detectable_effect, # power / minimum detectable effect
)
Every function returns a small dataclass that prints a one-line verdict and exposes the numbers (.corrected_p, .p_value, .loo_min …) so you can gate CI on them.
Reproduce the case studies:
python -m evalgate.checks # -> evalgate selftest: OK (reproduced all 3 case studies + power check)
Why this exists
These are textbook checks — the value isn't the math, it's running all of them, adversarially, on a number you're too close to. evalgate is the open, do-it-yourself version. When a launch, a paper, or a fundraise rides on a figure and you want it audited independently first, that's the paid practice.
Want the full checklist and the client-grade report template that wrap these checks? The Eval Integrity Kit — the 9-check audit checklist, the report template I ship to clients, an evalgate quickstart, and three worked case studies.
The fuller story — why AI benchmark scores and trading backtests overpromise, and how to catch them — is in the book Measured, Not Believed (pay what you want).
License
MIT — see LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。