context-firewall
Aggregator MCP proxy that collapses N downstream MCP servers into 4 meta-tools with progressive tool discovery, and compresses large tool outputs (HTML→Markdown, JSON summarization) with full-output retrieval via read_more and a per-session token-savings report.
README
Context Firewall
Turn 50+ MCP tools into 4, and shrink large tool outputs by 60–95% (real HTML/JSON, measured — see benchmark) — for any MCP client, any model. Any output still over your configured token budget after compression is hard-truncated to that budget, with the full original retrievable via read_more.
Context Firewall is a local MCP proxy that sits between your AI agent (Claude Code, Claude Desktop, Cursor, Cline, ...) and every downstream MCP server you've configured. It exposes exactly 4 tools to the client, no matter how many tools the downstream servers actually have, and compresses large tool outputs (raw HTML, base64 blobs, giant JSON) before they ever reach the model's context window.
Measured results
| Metric | Result |
|---|---|
| Tool collapse | 122 → 4 exposed meta-tools (5 real downstream servers incl. official GitHub github-mcp-server, 85 tools) |
| Tool-definition savings | ~28,600 tokens (estimated, chars ÷ 3.5) — 102,158 raw definition chars vs. 2,146 exposed |
| Output compression | 60–95%+ on real-world HTML/JSON, e.g. a live Wikipedia page via the fetch tool: 232,391 → 6,907 chars (97.0%, measured); a GitHub issues JSON payload via the jsonSummary stage: 186,810 → 3,480 chars (98.1%, measured) |
All figures measured against real downstream MCP servers, not synthetic data — see STATE.md ("P1-1 — github server portion" section) for full methodology.
What it does
- Progressive tool disclosure — instead of loading every downstream tool's full schema at startup, the client sees 4 meta-tools (
list_tool_categories,search_tools,invoke_tool,read_more) and only pays the token cost of a tool's full schema when it actually searches for it. - Output compression pipeline — large tool results are run through base64 stripping, HTML→Markdown conversion, JSON structure-aware summarization, and finally character-budget truncation, in that order, before being returned.
- Full output retrievable via
read_more— nothing is silently thrown away. Every compressed output is stored in full (in memory, opaque handle) and can be paged back withread_more(handle, offset, length). - Session savings report — on shutdown, prints a shareable terminal card (and optional Markdown file) showing tool-definition and output-token savings for the session, plus a breakdown of which tools saved the most.
Quickstart
npx context-firewall --config context-firewall.json
Minimal context-firewall.json (the downstreams block mirrors the mcpServers format you already know):
{
"downstreams": {
"filesystem": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-filesystem", "/path/to/project"]
},
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" }
}
}
}
(${GITHUB_TOKEN} above is just this example's environment variable name — pick whatever's already set in your shell; it gets expanded into GITHUB_PERSONAL_ACCESS_TOKEN, the env var name the downstream server itself actually reads.)
Which GitHub server? Two options, different tool counts:
-
@modelcontextprotocol/server-github(used above) — the original npm package, 26 tools, onenpx -yline, zero extra setup. Archived/no longer maintained upstream, but still functional. -
github/github-mcp-server— the actively-maintained official server, 44 tools (default toolset) to 85 (GITHUB_TOOLSETS=all). Ships as a Go binary or Docker image, not an npm package:"github": { "command": "docker", "args": ["run", "-i", "--rm", "-e", "GITHUB_PERSONAL_ACCESS_TOKEN", "ghcr.io/github/github-mcp-server"], "env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" } }or, running a locally-built/downloaded binary directly:
"command": "/path/to/github-mcp-server", "args": ["stdio"](sameenvblock; addGITHUB_TOOLSETSto scope which of the 85 tools are exposed).
Client setup
Set Context Firewall as your only MCP server — move every downstream server you currently configure directly (filesystem, github, everything, ...) into context-firewall.json's downstreams block instead. Your agent then sees 4 tools instead of the sum of every downstream server's tool count; pointing the client at Context Firewall alongside your existing servers doesn't give you the tool-collapse or compression benefit.
Each client below takes the same server entry:
{
"mcpServers": {
"context-firewall": {
"command": "npx",
"args": ["-y", "context-firewall", "--config", "/absolute/path/to/context-firewall.json"]
}
}
}
Claude Code
Project-scoped .mcp.json in your repo root (shown above), or via the CLI:
claude mcp add --transport stdio context-firewall -- npx -y context-firewall --config /absolute/path/to/context-firewall.json
Claude Desktop
claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json; Windows: %APPDATA%\Claude\claude_desktop_config.json) — same mcpServers block as above.
Cursor
.cursor/mcp.json (project-scoped) or ~/.cursor/mcp.json (global) — same mcpServers block as above.
Cline
cline_mcp_settings.json (VS Code extension storage; macOS: ~/Library/Application Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json) — same mcpServers block as above.
Compatibility status
| Client | Status |
|---|---|
| Claude Code | tested in real agent sessions — autonomous list → search → invoke → read_more workflow verified end-to-end |
| Claude Desktop | protocol-verified* |
| Cursor | config format documented, community testing welcome |
| Cline | config format documented, community testing welcome |
* Verified via MCP protocol integration tests (176 automated tests, including full stdio protocol round-trips against real downstream servers). Real-client reports welcome.
Configuration
downstreams
Each entry is either a stdio server (same shape as mcpServers) or a Streamable HTTP server:
{
"downstreams": {
"local-tool": { "command": "npx", "args": ["-y", "some-mcp-server"], "env": { "TOKEN": "${TOKEN}" } },
"remote-tool": { "url": "https://mcp.example.com/mcp", "transport": "streamable-http" }
}
}
${VAR_NAME} in any string value is expanded from the environment; a missing variable fails config load with a readable error.
compression
Policy resolution order is default < perServer < perTool (later overrides earlier, field by field):
{
"compression": {
"default": {
"maxOutputTokens": 2000,
"htmlToMarkdown": true,
"stripBase64": true,
"jsonSummary": true,
"bypass": false
},
"perServer": { "github": { "maxOutputTokens": 4000 } },
"perTool": { "filesystem/read_file": { "maxOutputTokens": 8000 } }
}
}
| Field | Type | Default | Meaning |
|---|---|---|---|
maxOutputTokens |
number | 2000 |
Soft budget (chars ≈ tokens × 3.5) a compressed output is truncated to as a last resort. |
htmlToMarkdown |
boolean | true |
Convert detected HTML markup to Markdown. |
stripBase64 |
boolean | true |
Replace base64 blobs (data URIs and bare blocks) with a read_more handle. |
jsonSummary |
boolean | true |
Collapse homogeneous JSON arrays and trim long string fields, keeping valid JSON. |
bypass |
boolean | false |
Skip the whole pipeline for this server/tool - output passes through untouched. |
report
{
"report": {
"enabled": true,
"markdownPath": "./context-firewall-report.md"
}
}
| Field | Type | Default | Meaning |
|---|---|---|---|
enabled |
boolean | true |
Print the session report to stderr on shutdown. |
markdownPath |
string | (none) | If set, also write the report as a Markdown file at this path. |
callToolTimeoutMs
Top-level (not nested under compression). Per-invoke_tool timeout in milliseconds passed to the downstream MCP SDK client; a downstream that hangs without responding causes invoke_tool to return an isError result once this elapses, instead of blocking. Defaults to the SDK's own default (60,000ms) when unset.
{ "callToolTimeoutMs": 30000 }
How it works
The client calls list_tool_categories() to see what's connected and what it's roughly capable of, search_tools(query) to pull the full input schema for candidate tools, invoke_tool(server, tool, args) to actually run one (compressed on the way back), and read_more(handle, offset, length) to page through anything that got compressed. Compression, when it runs, always applies in the same order: strip base64 → HTML to Markdown → JSON structure summary → truncate to budget. Security-relevant outputs (errors, permission/warning/confirmation messages) are never silently compressed - they pass straight through, only hard-capped at 50,000 characters to prevent a single runaway error dump from blowing out the caller's context.
Positioning
Context Firewall complements Anthropic's Tool Search Tool, it doesn't compete with it. Tool Search solves tool definition bloat at startup (the schemas loaded into context before any tool is even called), and is Claude-specific. Context Firewall compresses tool outputs at call time - the half of the problem Tool Search doesn't touch - and works with any MCP client and any model, not just Claude.
A note on token counts
Every token count in this project (truncation budgets, the session report) is estimated as chars / 3.5, never an exact model-specific count. There is no code path that sends your tool output content to an external API to get an exact count - that would conflict with the safety posture below. The session report is always labeled "(estimated)" for this reason.
Safety
- Tool arguments and output content are never written to logs or the session report - only server/tool names and character/token counts.
- Security-relevant outputs (errors, permission denials, warnings, confirmations) are never silently compressed.
- Downstream tool descriptions are treated as untrusted input and only ever displayed, never executed.
- Downstream tool descriptions are passed through verbatim, unsanitized -
search_toolsdoes not strip or filter prompt-injection text a malicious downstream might put there. The trust boundary is which downstream servers you choose to configure, not this gateway.
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。