agent-guardrail
MCP server for AI agent guardrails — validates and signs agent decisions (OAA tokens) so third parties can verify them without server access
README
Guardrail
<!-- mcp-name: io.github.rudimentall1/agent-guardrail -->
A policy firewall for AI agent tool calls.
Your agent wants to run a shell command, send an email, or move money. Guardrail checks that request against rules you wrote, before it happens, and either lets it through, asks a human, or blocks it — with a plain- English reason every time.
60-second quickstart
git clone <this repo> && cd agent-guardrail
pip install -r requirements.txt
python3 cli.py check --agent trading-agent-001 --tool wallet.transfer \
--args '{"amount": 9999, "to": "0xabc"}'
Or, once published, pip install guardrail-mcp gives you a guardrail
command directly — same output, no repo checkout required (falls back to
the policy bundled in the package if you don't point --policy at your
own file):
guardrail check --agent trading-agent-001 --tool wallet.transfer \
--args '{"amount": 9999, "to": "0xabc"}'
{
"decision": "BLOCK",
"matched_rules": [
{"rule": "numeric_cap_exceeded", "severity": "BLOCK",
"message": "amount=9999.0 exceeds cap 5 for 'wallet.transfer' (unknown agent)"}
]
}
That's it — no server, no account, no API key. policies/default.yaml is
the file that decided this; open it and change the numbers to match your
own rules.
Why this, not another "AI risk scoring" tool
Most "AI agent security" projects (including an earlier project of mine) lean on statistical risk scores computed from data nobody can actually verify at build time — wallet age, "reputation," contract "risk" — which either requires paid data feeds you don't have yet, or quietly becomes mock data pretending to be real. Fine for prototyping, dishonest to ship.
Guardrail only makes claims it can back up. Every check is a deterministic rule — a blocklist entry, a regex match, a numeric cap, a rate limit — evaluated against a policy file you write and can audit yourself, backed by a real, persistent audit log (SQLite) you can query. Nothing here pretends to know something it doesn't.
It's also not blockchain-specific. Shell execution, email, HTTP requests, file deletion, database writes, crypto transactions — same engine, same policy file, same rules.
Three ways to use it
1. CLI — for testing a policy by hand
Shown above. No setup, instant feedback while you write rules.
2. MCP server (mcp_server.py) — the easy on-ramp, advisory
Exposes guardrail_check, guardrail_record_outcome, and
guardrail_agent_history as MCP tools any MCP-compatible agent (Claude
Desktop, Claude Code, custom MCP clients) can call.
{
"mcpServers": {
"guardrail": {
"command": "python3",
"args": ["/absolute/path/to/agent-guardrail/mcp_server.py"],
"env": { "GUARDRAIL_POLICY": "/absolute/path/to/agent-guardrail/policies/default.yaml" }
}
}
}
Then tell your agent (in its system prompt) to always call
guardrail_check before spending money, deleting data, messaging someone
externally, or running code.
Be clear-eyed about its limit: like any MCP tool, nothing stops the calling model from just not invoking it. This only helps if the agent is instructed to always check first — for a guarantee it can't skip, see #3.
3. guardrail.decorator.enforce — the real guarantee
Wraps the actual Python function that performs a tool's side effect. The check runs in your code, before that function executes — the model never gets a chance to call the real function directly.
from guardrail.decorator import enforce, BlockedActionError
@enforce(engine, tool_name="send_email")
def send_email(agent_id: str, to: str, subject: str, body: str):
... # only runs if the decision is ALLOW, or WARN-and-confirmed
Use this if you're building your own agent loop (LangChain, CrewAI, a
custom MCP host, a Slack bot with tool access). Run python3 examples/example_agent_usage.py to see it block a real function call.
Getting a human to actually confirm a WARN
on_warn is the hook — Guardrail ships two ready-made implementations:
Local web UI (guardrail/confirmation/web_ui.py) — a tiny built-in
server (stdlib only, no Flask) with Approve/Reject buttons. The wrapped
function blocks until someone clicks one, or times out (fails closed —
timeout means reject, not "allow by default").
from guardrail.confirmation.web_ui import ConfirmationServer
confirmation = ConfirmationServer(port=8787, timeout_seconds=300)
confirmation.start(open_browser=True)
@enforce(engine, tool_name="wallet.transfer", on_warn=confirmation.request_confirmation)
def transfer(...): ...
Try it live: python3 examples/example_web_confirmation.py, then open
http://localhost:8787.
Terminal prompt (guardrail/confirmation/cli_ui.py) — for scripts and
local testing where a browser is overkill:
from guardrail.confirmation.cli_ui import cli_confirm
@enforce(engine, tool_name="wallet.transfer", on_warn=cli_confirm)
def transfer(...): ...
Neither is required — on_warn is just a function (decision) -> bool,
so a Slack message, a ticket, or anything else you already use works too.
Writing a policy
Policies are plain YAML — see policies/default.yaml for a real, working
starting point (11 confirmation-gated tools, 10 destructive-pattern
checks, numeric caps, domain rules, rate limits, all commented).
| Rule type | What it checks |
|---|---|
blocked_tools |
Tool names that are never allowed |
confirmation_required_tools |
Tool names that always produce WARN |
argument_patterns |
Regex against the JSON-serialized call arguments — destructive shell commands, SQL, leaked credentials, path traversal, SSRF, force-pushes, regardless of which tool carries them |
numeric_caps |
Per-tool numeric field caps, tighter for agents with no history |
domain_rules |
Allow/deny lists on a URL or email-recipient field, per tool |
rate_limits |
Sliding-window call limits per (agent, tool), backed by SQLite |
No code changes needed to adjust any of this — edit the YAML, restart the process (or the MCP server).
Running the tests
pip install -r requirements.txt
PYTHONPATH=. python3 -m unittest discover -s tests -v
46 tests: rule evaluation, the full engine pipeline (real SQLite-backed
rate limiting and audit persistence), the enforce decorator (proving a
BLOCK genuinely prevents the wrapped function from running), the
hand-rolled MCP server's JSON-RPC handling over an actual stdio pipe, the
confirmation web UI over real HTTP requests against a live server, and a
dedicated suite that checks the shipped policies/default.yaml — not
just synthetic test policies — actually catches what it claims to.
What's honestly still missing
- Single-process SQLite by default. Fine for one agent process; for multiple replicas sharing rate limits/audit history, point every process at the same file on shared storage, or swap in a real database (the storage classes are small and easy to re-target).
- No built-in secrets/PII redaction in the audit log. Arguments are
stored as-submitted. If your tools take sensitive arguments, redact
before calling
evaluate(), or extendAuditLogto redact specific fields before persisting. - The default policy is a reasonable starting point, not a complete
threat model. It catches well-known destructive shell/SQL patterns
and obvious credential formats — extend
argument_patternsfor whatever your agents actually touch. - The confirmation web UI has no auth. It binds to
127.0.0.1by design (not exposed on the network), but anyone with local access to that port can approve/reject. Fine for a single developer's machine; put it behind your own auth if multiple people share the host.
None of these are mocked or faked — they're just not built yet, and they're the honest next steps if you adopt this.
Publishing this / getting people to actually use it
See PUBLISHING.md for a concrete checklist: MCP directories to submit
to, what a listing needs, and what "done" looks like.
Related projects
Same author, same principle applied elsewhere:
- agentic-wallet-guardian-v3 - a security decision layer for AI agents transacting on-chain. MIT, 101 tests.
- x402-attest - cryptographically signed (Ed25519), independently verifiable attestations for agent-to-agent payment policy decisions. Early proof of concept.
- open-agent-attestation - vendor-neutral open spec (JWT+EdDSA) for signing agent policy decisions, verifiable by anyone. x402-attest above uses a custom format; this is the generalized version. Draft v0.1.
Project layout
guardrail/
__main__.py CLI implementation — also the `guardrail` console command
mcp_server.py MCP stdio server — also the `guardrail-mcp-server` console command
core/
models.py ActionRequest, RuleMatch, GuardrailDecision (stdlib only)
policy.py Policy loader (the one place PyYAML is used)
rules.py Deterministic rule evaluators
storage/
rate_limiter.py SQLite-backed sliding-window rate limiter
audit.py SQLite-backed persistent audit log
engine.py GuardrailEngine — orchestrates rules + rate limit + audit
decorator.py enforce() — the unbypassable integration point
confirmation/
web_ui.py Local web UI for human approve/reject (stdlib http.server)
cli_ui.py Terminal-prompt confirmation
policies/default.yaml Copy of the default policy bundled into the installed package
policies/default.yaml Canonical, editable default policy (git-clone workflow)
cli.py Thin shim -> guardrail/__main__.py (for `python3 cli.py`)
mcp_server.py Thin shim -> guardrail/mcp_server.py (for `python3 mcp_server.py`)
pyproject.toml Package metadata — `pip install .` gives you `guardrail` + `guardrail-mcp-server`
.github/workflows/ci.yml Runs the test suite + policy validation + package build on every push
examples/
example_agent_usage.py Decorator basics
example_web_confirmation.py Real browser-based approve/reject, live
tests/ 46 unit tests, all runnable with just PyYAML installed
CONTRIBUTING.md How to add a rule type, ground rules
CHANGELOG.md Version history
PUBLISHING.md How to actually get this in front of people
landing/index.html Static one-page site (open directly or host on GitHub Pages)
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。