ToolForge MCP Server
Enables AI agents to securely discover, execute, and observe tools with role-based access control and audit logging. Serves tools over MCP stdio and HTTP for integration with Claude Desktop, Cursor, and other clients.
README
<div align="center">
⚒️ ToolForge
Agent Tool Infrastructure & MCP Platform
A production-grade runtime that gives AI agents governed, observable, and secure access to tools.
</div>
What is ToolForge?
Modern AI agents call tools as raw function calls — with no governance, no observability, and no reliability guarantees. ToolForge fills that gap.
It acts as a secure execution gateway between your AI agent and the outside world, providing a full lifecycle for every tool invocation:
| Stage | What it does |
|---|---|
| 🔍 Discover | Semantic registry with semver, categories, and full-text search |
| ✅ Validate | Auto-generated JSON Schema from Python type hints via Pydantic v2 |
| 🔐 Authorize | Capability-based RBAC permissions checked before every execution |
| ⚡ Execute | Async-first runtime with timeouts, retries, and sandboxing |
| 📊 Observe | Structured ToolResult — execution ID, latency, retry count, error metadata |
| 🌐 Expose | MCP stdio & HTTP, LangGraph adapter, OpenAI function calling schemas |
Architecture
flowchart LR
subgraph Clients["Clients"]
A1["🤖 AI Agent"]
A2["🖥️ Claude Desktop"]
A3["⚡ Cursor / IDE"]
A4["🌐 REST API"]
end
subgraph Gateway["ToolForge Gateway"]
direction TB
B["🔑 Auth Layer\nAPI Key · JWT"]
C["🛡️ RBAC Engine\nRoles · Capabilities"]
D["⏱️ Rate Limiter\nSliding Window"]
B --> C --> D
end
subgraph Runtime["Execution Runtime"]
direction TB
E["📋 Input Validation\nJSON Schema · Pydantic"]
F["🔒 Sandbox\nDocker · Subprocess"]
G["🔄 Retry + Timeout\nExponential Backoff"]
H["📦 ToolResult\nID · Latency · Error"]
E --> F --> G --> H
end
subgraph Registry["Tool Registry"]
direction TB
I["📚 26 Standard Tools\nSemver · Categories"]
J["🔎 Semantic Search\nFull-Text + Category"]
I --- J
end
subgraph Infra["Infrastructure"]
K[("🐘 PostgreSQL\nAudit Logs")]
L[("⚡ Redis\nRate Limits · Cache")]
M["📈 Prometheus\n/metrics exporter"]
end
Clients --> Gateway
Gateway --> Runtime
Runtime --> Registry
Runtime --> Infra
Quick Start
Install:
git clone https://github.com/your-username/toolforge
cd toolforge
pip install -e .
Register and execute a custom tool:
import asyncio
from toolforge import ToolForge
tf = ToolForge()
@tf.tool(
name="calculate_tax",
description="Calculate income tax for a given gross income and rate.",
version="1.0.0",
category="finance",
)
def calculate_tax(income: float, rate: float = 0.2) -> float:
return income * rate
async def main():
result = await tf.execute("calculate_tax", {"income": 120_000.0, "rate": 0.28})
print(result.status.value) # 'success'
print(f"${result.result:,.2f}") # '$33,600.00'
print(result.execution_id) # UUID for tracing
asyncio.run(main())
Load all 26 standard tools in one line:
from toolforge import ToolForge
tf = ToolForge.with_standard_tools()
Standard Tool Ecosystem — 26 Tools
| Category | Tools |
|---|---|
| 🌐 Web | web_search, web_fetch, http_request |
| 📁 Files | file_read, file_write, file_search, directory_list |
| 📊 Data | pdf_extract, csv_read, json_transform |
| 🗄️ Database | sql_query (read-only guard), sql_schema, redis_get, redis_set |
| 🌿 Git | git_status, git_diff, git_log |
| 🐙 GitHub | github_search, github_file, github_issue, github_actions |
| 🐍 Code | python_execute, shell_execute |
| 🧠 AI | embedding_generate, vector_search, rerank |
Core Features
Custom Tool Registration
from toolforge import tool, RetryPolicy, StandardCapability
@tool(
name="fetch_stock_price",
description="Fetch live stock price from market API.",
version="1.0.0",
category="finance",
capabilities=[StandardCapability.NETWORK.value],
timeout=10.0,
retry_policy=RetryPolicy(max_retries=3, initial_delay_sec=0.5),
)
async def fetch_stock_price(ticker: str) -> dict:
"""Fetch stock data for given ticker symbol."""
return {"ticker": ticker, "price": 185.42}
Structured Tool Results
Every execution returns a fully typed ToolResult — no raw dicts, no guessing:
result = await tf.execute("web_search", {"query": "python asyncio"})
result.execution_id # UUID — for distributed tracing
result.tool_name # "web_search"
result.tool_version # "1.0.0"
result.status # SUCCESS | FAILED | TIMEOUT | PERMISSION_DENIED
result.result # structured output
result.error # ToolErrorInfo(code, message, retryable)
result.duration_ms # wall-clock latency
result.retry_count # retries attempted before success
result.unwrap() # raises RuntimeError on failure, else returns result
Capability-Based Permissions
from toolforge import PermissionContext
ctx = PermissionContext(
caller_id="research_agent",
granted_capabilities={"network", "filesystem_read"},
)
# ✅ Tool requires 'network' — succeeds
result = await tf.execute("web_search", {"query": "AI"}, context=ctx)
# ❌ Tool requires 'code_execution' — returns PERMISSION_DENIED, never raises
result = await tf.execute("python_execute", {"code": "..."}, context=ctx)
print(result.status.value) # 'permission_denied'
MCP Protocol Integration
Expose all 26 tools to Claude Desktop, Cursor, or any MCP-compatible client with zero configuration.
Stdio transport (for Claude Desktop / Cursor):
import asyncio
from toolforge import ToolForge
tf = ToolForge.with_standard_tools()
mcp = tf.create_mcp_server(server_name="my-toolforge")
asyncio.run(mcp.run_stdio())
Or directly via CLI:
python -m toolforge.integrations.mcp
HTTP transport (for remote agents):
# POST /mcp — JSON-RPC 2.0
curl -X POST http://localhost:8000/mcp \
-H "X-API-Key: tf-..." \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
MCP methods implemented:
| Method | Description |
|---|---|
initialize |
Capability handshake with client info |
tools/list |
Dynamic schema discovery for all tools |
tools/call |
Executes tool through ToolForge runtime (RBAC + retry) |
resources/list |
Spec-compliant resource listing |
prompts/list |
Spec-compliant prompt listing |
ping |
Heartbeat keepalive |
LangGraph & OpenAI Adapters
LangGraph:
tools = tf.to_langgraph_tools() # List[LangGraphToolWrapper]
# graph = create_react_agent(llm, tools)
Each wrapper supports .invoke() (sync) and .ainvoke() (async) with args_schema for type validation.
OpenAI function calling:
openai_tools = tf.to_openai_tools(category="web")
# → [{"type": "function", "function": {"name": ..., "parameters": {...}}}]
from toolforge import parse_openai_tool_call
tool_name, args = parse_openai_tool_call(tool_call_obj)
result = await tf.execute(tool_name, args)
Enterprise Security
Initialize the hardened SecuredToolForge client:
from toolforge import SecuredToolForge, Role, RateLimitConfig
tf = SecuredToolForge.with_standard_tools(
rate_limit_config=RateLimitConfig(requests_per_minute=60),
circuit_failure_threshold=5,
circuit_recovery_timeout_sec=30.0,
)
# Issue scoped API keys per role
dev_key = tf.api_keys.issue_key(user_id="alice", role=Role.DEVELOPER)
agent_key = tf.api_keys.issue_key(user_id="bob", role=Role.AGENT)
# Execute with key-based auth and RBAC enforcement
result = await tf.execute("python_execute", {"code": "print('hello')"}, api_key=dev_key)
# Query auto-redacted audit events
events = tf.audit.get_events(caller_id="alice")
Security features at a glance:
- 🔑 Authentication — SHA-256 hashed API key store + signed JWT bearer tokens
- 👥 RBAC Hierarchy —
admin > developer > agent > viewerwith per-category capability enforcement - ⏱️ Rate Limiting — Sliding-window per-user and per-tool limits (in-memory or Redis)
- 🔌 Circuit Breaker —
CLOSED → OPEN → HALF_OPENprotecting against cascading failures - 🐳 Sandboxed Execution — Ephemeral Docker containers (CPU/memory caps, read-only rootfs, no networking); graceful subprocess fallback
- 📋 Audit Logging — Structured JSON events with automatic secret/token redaction
FastAPI Platform Server
Start the server:
uvicorn toolforge.server:app --host 0.0.0.0 --port 8000 --workers 4
Complete REST API:
| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Liveness probe |
GET |
/ready |
Readiness probe with DB verification |
GET |
/api/v1/tools |
List tools with search & category filters |
GET |
/api/v1/tools/{name} |
Full JSON Schema for a specific tool |
POST |
/api/v1/tools/{name}/execute |
Synchronous tool execution |
POST |
/api/v1/tools/batch |
Batch execute multiple tools |
POST |
/api/v1/jobs |
Enqueue async background job |
GET |
/api/v1/jobs/{id} |
Poll job status & result |
POST |
/api/v1/auth/keys |
Issue scoped API keys |
POST |
/api/v1/auth/token |
Issue JWT access tokens |
GET |
/api/v1/audit |
Query historical execution records |
GET |
/api/v1/metrics |
JSON metrics summary |
GET |
/metrics |
Prometheus text exposition format |
POST |
/mcp |
Remote MCP JSON-RPC 2.0 |
GET |
/mcp/sse |
MCP Server-Sent Events stream |
Observability Dashboard
Navigate to http://localhost:8000/ for a live glassmorphism SPA dashboard:
- 📊 KPI Metrics Hub — Real-time execution volume, success rates, and p95 latencies
- 🔎 Tool Explorer — Searchable grid across all 8 categories with full JSON schema inspector
- 🧪 Interactive Playground — Execute any tool live with formatted response and latency benchmarks
- 📋 Live Audit Stream — Execution history with expandable sanitized request/response inspection
- 🔗 MCP Protocol Hub — Connection guides for Cursor & Claude Desktop + JSON-RPC 2.0 tester
- 🔑 Security Panel — Issue API keys and JWT tokens directly from the browser
Autonomous Agent Demo
ToolForge ships a built-in Autonomous Repository Debugger that resolves bugs entirely on its own through the secured runtime:
python examples/repo_debugger_agent.py
The agent executes a 6-step autonomous loop:
Step 1 discover → directory_list + file_search find project structure
Step 2 reproduce → python_execute run failing tests
Step 3 inspect → file_read read buggy source
Step 4 patch → file_write apply autonomous fix
Step 5 verify → python_execute re-run tests → green
Step 6 report → structured summary root cause + timing
Output:
[REPORT] AGENT RESOLUTION SUMMARY
Total Steps: 6
Total Runtime: 249.6ms
Root Cause: Unhandled division by zero in calculator.py
Resolution: Patched divide() with explicit ValueError guard
Final Status: RESOLVED
Production Deployment
Docker Compose (FastAPI Server + PostgreSQL 16 + Redis 7):
docker compose up --build
| URL | Description |
|---|---|
http://localhost:8000/ |
Web dashboard |
http://localhost:8000/docs |
Swagger / OpenAPI docs |
http://localhost:8000/health |
Health check |
http://localhost:8000/metrics |
Prometheus metrics |
CI/CD — GitHub Actions
Every push runs a 4-job pipeline:
Lint (ruff) → Test Matrix (3 OS × 3 Python) → Agent Smoke Test → Docker Build
| Job | Details |
|---|---|
| Lint | ruff check + ruff format --check — fast-fail gate |
| Test Matrix | Ubuntu · Windows · macOS × Python 3.11 · 3.12 · 3.13 = 9 environments |
| Agent Demo | Full autonomous agent run end-to-end |
| Docker Build | Multi-stage image build with GHA layer cache |
Testing
# Run all 86 tests (unit, integration, API, agent)
python -m pytest tests/ -v
86 passed in 6.89s
Project Phases
| Phase | Status | What Was Built |
|---|---|---|
| 1 — Core SDK | ✅ Complete | BaseTool, @tool, ToolRegistry, ToolRuntime, Pydantic v2 schemas, retries |
| 2 — Tool Ecosystem | ✅ Complete | 26 standard tools, MCP stdio server, LangGraph + OpenAI adapters |
| 3 — Security | ✅ Complete | Docker sandbox, RBAC, API keys, JWT, rate limiter, circuit breaker, audit log |
| 4 — FastAPI + Workers | ✅ Complete | REST API, PostgreSQL/SQLite, async job queue, MCP HTTP/SSE |
| 5 — Dashboard | ✅ Complete | Glassmorphism SPA, Prometheus metrics, live playground |
| 6 — Demos & CI/CD | ✅ Complete | Autonomous agent, Docker Compose stack, GitHub Actions 9-env matrix |
Tech Stack
<div align="center">
Python 3.11+ · FastAPI · Pydantic v2 · SQLAlchemy 2.0 · PostgreSQL · Redis · Docker · MCP Protocol · LangGraph · Prometheus
</div>
Contributing
Contributions are welcome. See CONTRIBUTING.md for guidelines.
Built with Python 3.11+, Pydantic v2, FastAPI, SQLAlchemy 2.0, and async-first patterns throughout.
<div align="center"> <sub>Made with ⚒️ by the ToolForge team · Apache 2.0 License</sub> </div>
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。