educator-toolkit-mcp
An MCP toolkit exposing six AI-powered educator-craft tools for reviewing, critiquing, and improving learning artifacts like lessons, curricula, AI tools, learner experiences, assessments, and content accuracy. Deployable via hosted HTTPS, local stdio MCP, or an importable Python kernel.
README
educator-toolkit-mcp
A portable MCP toolkit exposing six educator-craft AI tools for reviewing, critiquing, and improving learning artifacts. Wraps structured AI-powered review tools and exposes them three ways:
- Hosted on GCP — call directly over HTTPS (REST) or connect Claude Desktop / Cursor via MCP Streamable HTTP. No install needed.
- Local MCP server — clone and run as a subprocess of Claude Desktop over stdio.
- Importable Python kernel — embed
run_tooldirectly in your own application.
Option 1: Use the hosted service (no install)
Base URL: https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app
Prototype, no auth. The URL is unauthenticated for prototyping. Don't paste it into public places — anyone with it can call Anthropic on the shared API budget. Real auth arrives with the e-commerce milestone.
REST — POST /v1/{tool_name}
curl -X POST https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/v1/ai_tool_critique \
-H "Content-Type: application/json" \
-d '{
"question": "Is the input schema asking too much upfront?",
"artifact": "A form with 8 required fields for a first-time user, no examples.",
"artifact_kind": "ai_tool_prompt"
}'
Response shape:
{
"status": "complete" | "needs_input" | "out_of_lane" | "refused" | "error",
"assessment": "...",
"key_points": ["..."],
"follow_ups": ["..."] | null,
"suggested_tool": "..." | null,
"caveats": ["..."] | null,
"error": null
}
Replace ai_tool_critique with any of the six tool IDs listed in The six tools section below.
Claude Desktop (remote MCP)
In claude_desktop_config.json:
{
"mcpServers": {
"educator-toolkit-remote": {
"url": "https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/mcp/",
"transport": "streamable-http"
}
}
}
Note: trailing slash on /mcp/ is required (Starlette mount convention).
Restart Claude Desktop. All six tools will be available in the tool picker.
Cursor / Cline / other MCP clients
Same shape — point them at https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/mcp/ over the Streamable HTTP transport. No headers, no auth.
Health check
curl https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/health
# {"status":"ok"}
OpenAPI / docs
Auto-generated by FastAPI:
- Swagger UI: https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/docs
- OpenAPI JSON: https://educator-toolkit-server-h4ldfaoucq-uw.a.run.app/openapi.json
Option 2: Run locally as a Claude Desktop subprocess
For when you want to use your own Anthropic API key or run without Cloud Run.
git clone https://github.com/TyRobbins/educator-toolkit-mcp.git
cd educator-toolkit-mcp
uv sync --all-packages --all-groups
export ANTHROPIC_API_KEY=sk-ant-...
In claude_desktop_config.json (Windows path; macOS path differs):
{
"mcpServers": {
"educator-toolkit-local": {
"command": "uv",
"args": [
"run",
"--directory", "C:\\Users\\YOU\\path\\to\\educator-toolkit-mcp",
"educator-toolkit-mcp"
],
"env": {
"ANTHROPIC_API_KEY": "sk-ant-..."
}
}
}
}
The local server uses MCP over stdio (Claude Desktop launches it as a child process). The hosted version uses MCP over HTTP.
Option 3: Import the kernel directly
import asyncio
from educator_toolkit_kernel import AnthropicClient, run_tool
from educator_toolkit_kernel.schema import ToolInput
client = AnthropicClient() # reads ANTHROPIC_API_KEY from env
response = asyncio.run(run_tool(
tool_id="ai_tool_critique",
input=ToolInput(
question="Is the input schema asking too much upfront?",
artifact="A form with 8 required fields for a first-time user, no examples.",
artifact_kind="ai_tool_prompt",
),
client=client,
))
print(response.status)
print(response.assessment)
print(response.key_points)
The six tools
| Tool ID | What it does | Use when |
|---|---|---|
educational_design_review |
Review a learning artifact for instructional-design quality — objectives, scaffolding, Bloom's taxonomy, cognitive load | You want a learning-science review of a lesson, module, or activity |
curriculum_coherence_review |
Review a curriculum or module sequence for coherence, redundancy, and gaps across phases | You have 2+ modules and want sequencing feedback |
ai_tool_critique |
Critique an AI tool's prompt, input schema, and output usefulness for its intended learner | You have one specific AI tool (prompt + schema) and want a critique |
student_experience_simulation |
Simulate what a realistic learner would do, feel, and produce when using an artifact | You want a persona-grounded prediction of learner behavior |
assessment_design_review |
Review or design assessments: cases, rubrics, formative checks, capstone scenarios | You have an assessment artifact (rubric, case, exam) or want one designed |
sme_review |
Review content for accuracy, terminology consistency, and fidelity to a supplied source | You supply BOTH a content artifact AND source material for comparison |
Worked example
An instructional designer has built an AI tool that helps learners write stakeholder communication plans. The tool has a long intake form. They want to know if the schema is too heavy.
ToolInput sent:
{
"question": "Is the input schema asking too much upfront?",
"artifact": "A form with 8 required fields for a first-time user, no examples.",
"artifact_kind": "ai_tool_prompt"
}
ToolResponse returned:
{
"status": "complete",
"assessment": "Eight required fields with no examples is excessive for first-time users and will cause high abandonment. Core fixes: reduce required fields to 3-4, add placeholder examples to every field, use progressive disclosure for optional context, and sequence by cognitive load.",
"key_points": [
"8 required fields exceeds the friction threshold for first-time users — expect abandonment",
"Zero examples or placeholder text creates blank-canvas paralysis across all fields",
"All-required structure signals rigidity and prevents progressive onboarding",
"Fix: cut to 3-4 required fields, make the rest optional with a clear rationale",
"Fix: add one inline example per field showing what a good answer looks like",
"Fix: consider a two-step form — generate output first, refine with more input second"
],
"follow_ups": [
"Can you share the actual field labels? Some fields may be combinable or eliminable once we see them.",
"Does the underlying AI prompt use all 8 fields, or are some fields only used for logging/routing?"
],
"caveats": [
"Artifact was described, not shown — specific field label critique requires the actual schema",
"Whether the 8 fields are pedagogically correct is out of lane for this review (see educational_design_review)"
],
"error": null
}
How it works
Each tool makes one Sonnet call — no coordinator, no sub-tool loop. The caller (Claude Desktop, a REST client, or your own app) decides which tool to route to and composes results from multiple tools if needed.
- Stateless: each call is independent. The caller maintains context across follow-up rounds.
- Schema gate: every response is validated against
ToolResponsebefore returning. If the model omits the JSON envelope, the prose is wrapped verbatim. - One model per tool: all six tools currently use
claude-sonnet-4-6.
Quality
Validated against a 12-fixture golden eval set (2 per tool, LLM-judge rubric, 4 dimensions x 1-5 scale). Current average: 4.90/5.00. See evals/.
CI gate: score must stay >= 3.5.
Schema validity gate: uv run python -m evals.runners.schema_validity — all 12 fixtures must return parseable ToolResponse JSON.
Architecture
┌─────────────────────────────────────────┐
│ Hosts: Claude Desktop, Cursor, scripts │
└────────────┬──────────────┬─────────────┘
│ │
stdio MCP HTTP / MCP
│ │
┌────────────▼──┐ ┌───────▼───────────────┐
│ packages/mcp │ │ packages/server │
│ (local sub- │ │ (Cloud Run) │
│ process) │ │ /v1/{tool_name} │
│ │ │ /mcp/ │
└────────────┬──┘ └───────┬───────────────┘
│ │
└──────┬───────┘
│
┌───────▼────────────────────┐
│ packages/kernel │
│ - schema (ToolInput, │
│ ToolResponse) │
│ - registry (TOOLS) │
│ - runner (run_tool) │
│ - prompts (6 tool prompts) │
└────────────────────────────┘
Development
git clone https://github.com/TyRobbins/educator-toolkit-mcp.git
cd educator-toolkit-mcp
uv sync --all-packages --all-groups
uv run pytest packages/
# Lint + type-check
uv run ruff check .
uv run mypy packages/kernel/src packages/mcp/src packages/server/src
# Evals (require ANTHROPIC_API_KEY)
uv run python -m evals.runners.schema_validity
uv run python -m evals.runners.judge_rubric
# Run the HTTP server locally
ANTHROPIC_API_KEY=sk-ant-... uv run educator-toolkit-server
# -> http://localhost:8080/health
Deploy
gcloud builds submit --config cloudbuild.yaml --project ou-executive-persuasion
Builds the Dockerfile, pushes to GCR, deploys to Cloud Run (region us-west1, public, no auth). The ANTHROPIC_API_KEY is read from Secret Manager (entry anthropic-api-key).
License
TBD pending IP review.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。