support-agent-mcp
Exposes order status lookup and knowledge base search tools from the Support Agent AI over MCP, enabling MCP clients to handle customer support queries with grounded, citation-backed answers.
README
Support Agent + MCP
Live demo: support-agent-mcp.onrender.com/docs — interactive Swagger UI. Health:
/healthz. Hosted on Render's free tier, which sleeps after ~15 min idle, so the first request may take 30–60s to wake.
A production-shaped customer-support AI agent: a FastAPI service where a LangGraph agent answers customer questions by calling tools — order lookup, OAuth-scoped refunds, and a grounded knowledge base — with token streaming, conversation memory, OpenTelemetry traces and structured JSON logs. The same tools are re-exposed over the Model Context Protocol (MCP) for any MCP client.
Runs entirely on free infrastructure (Gemini free tier, a free Render/Spaces dyno) and the whole test suite runs offline, with no API key.
Flagship features → what they demonstrate
| Capability | Where | AI-Engineer JD bullet it answers |
|---|---|---|
| Secure REST API + validation | FastAPI + Pydantic v2 (app/main.py, app/schemas.py) |
Build and expose secure APIs |
| Agent workflow / tool routing | LangGraph ReAct agent (app/agent.py) |
Design agentic workflows |
| Tool authorization | OAuth2 password + JWT scopes; refunds gated on refund:write (app/auth.py) |
Guardrails; least-privilege tool access |
| LLM integration + bounded retry | Gemini via LangChain, exponential backoff (app/llm.py) |
Integrate LLM providers reliably |
| Retrieval grounding + citations | Pluggable KB: offline keyword / semantic Chroma (app/vectorstore.py) |
RAG, grounded answers, anti-hallucination |
| Token streaming (SSE) | POST /chat/stream (app/streaming.py) |
Async Python; responsive UX |
| Conversation memory | LangGraph checkpointer keyed by session_id (app/agent.py) |
Stateful, multi-turn agents |
| Distributed tracing | OpenTelemetry spans on agent run, every tool call, the LLM call (app/tracing.py) |
Observability for LLM systems |
| Structured logging | JSON logs + X-Request-ID correlation (app/logging_config.py) |
Production operability |
| Offline eval harness | Scored behavioural evals (evals/) |
Measure agent quality, not vibes |
| MCP service | mcp_server/server.py (FastMCP, stdio) |
Interoperable tool servers |
| Containers + free deploy | Dockerfile, docker-compose.yml, render.yaml, deploy/huggingface/ |
Containerization and deployment |
| Tests + CI gate | 56 tests, .github/workflows/ci.yml |
Testing discipline |
What makes this production-shaped
It is not the feature list — it is the failure behaviour:
- Authorization is enforced at the tool, not in the prompt. The model can decide
to refund; without
refund:writeon the caller's token the tool still refuses, and the denial is recorded as a span attribute for audit. - Degrades instead of dying. No API key →
/healthzand/tokenstill serve and chat returns an actionable503. No vector service → the keyword retriever takes over. No collector → traces go to stdout. A broken stream closes with anerrorevent rather than a half-written response. - Every answer is attributable. Policy replies carry
citationsrecovered from the retriever's own output, and every log line carries the request id also returned inX-Request-ID. - The offline/online split is deliberate. Pure logic (SSE translation, turn slicing, eval graders, exporter selection) is separated from I/O, so CI proves behaviour with no network and no key — and the same graders score a live Gemini run locally.
Architecture
flowchart LR
subgraph Client
U[HTTP client / UI]
MCPC[MCP client<br/>Claude, IDEs]
end
U -->|POST /chat<br/>POST /chat/stream<br/>Bearer token| API[FastAPI<br/>JSON logs + request id]
API -->|JWT scopes| A["LangGraph ReAct agent<br/>(Gemini)"]
A <-->|thread_id| M[(MemorySaver<br/>checkpointer)]
A --> T1[get_order_status]
A --> T2["request_refund<br/>needs refund:write"]
A --> T3[search_kb]
T3 --> KB[(KB: keyword / Chroma)]
API -->|reply · tool_calls · citations<br/>or SSE token stream| U
MCPC --> S[FastMCP server] --> T1 & T3
A -.spans.-> OTEL{{OpenTelemetry<br/>console / OTLP}}
T1 & T2 & T3 -.spans.-> OTEL
Quickstart
python -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
cp .env.example .env # add a free key from https://aistudio.google.com/apikey
pytest -q # runs fully offline, no key needed
uvicorn app.main:app --reload
Or with Docker:
docker compose up --build # spans stream to the container log
Interactive API docs at http://localhost:8000/docs.
Demo
# 1) Authenticate as an agent (gets refund:write). Try 'customer/customer' to see a denial.
TOKEN=$(curl -s -X POST localhost:8000/token -d 'username=agent&password=agent' | python -c 'import sys,json;print(json.load(sys.stdin)["access_token"])')
# 2) Grounded policy answer (returns citations)
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"message":"How long do refunds take?"}'
# 3) Authorized action
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"message":"Refund order A1001, it was defective."}'
# 4) Streamed answer — tokens and tool steps as they happen
curl -N -X POST localhost:8000/chat/stream -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"message":"How long do refunds take?"}'
# 5) Multi-turn memory — reuse the session_id and the agent remembers the order
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"message":"Where is order A1002?","session_id":"demo-1"}'
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"message":"Can I refund it?","session_id":"demo-1"}'
Endpoints
| Method | Path | Notes |
|---|---|---|
GET |
/healthz |
liveness + which backends are live |
POST |
/token |
OAuth2 password flow → JWT with scopes |
POST |
/chat |
JSON reply with tool_calls + citations |
POST |
/chat/stream |
Server-Sent Events: start, tool_call, tool_result, token, done, error |
Observability
Tracing is controlled by OTEL_TRACES_EXPORTER: console (default — spans to stdout,
zero infrastructure), otlp (any OTLP/HTTP collector via OTEL_EXPORTER_OTLP_ENDPOINT),
or none. Spans cover the agent run, each tool call and the LLM call:
{"name": "tool.request_refund", "attributes": {"tool.name": "request_refund", "authz.allowed": false}}
Logs are one JSON object per line, each carrying the request_id that is also returned
in the X-Request-ID response header (LOG_FORMAT=text for a human-readable dev view):
{"ts": "…", "level": "INFO", "logger": "support-agent.access", "message": "request",
"request_id": "9e8d65214aae4af4", "method": "GET", "path": "/healthz", "status": 200, "duration_ms": 0.75}
Evals
evals/ scores the three behaviours the agent is actually hired for: picking the right
tool, refusing an unauthorized refund, and answering policy questions from the KB
with a citation.
python -m evals.runner --min-pass-rate 0.8 # live Gemini run; exits non-zero below the bar
The graders (evals/scorers.py) are pure functions, so CI exercises them against fixtures
with no key; the live end-to-end run is pytest.mark.skipif-ed off when GOOGLE_API_KEY
is unset. That is what keeps CI hermetic while the same rubric grades a real model locally.
MCP
python -m mcp_server.server # exposes order_status + knowledge_base over stdio
Refunds are intentionally not exposed over MCP: that action requires a scope-bearing session, which the local stdio transport does not carry.
Deploy (free tiers)
Render — render.yaml is a ready blueprint (free plan, Docker runtime, health check
on /healthz). Push the repo, then Render → New → Blueprint → select it. GOOGLE_API_KEY
is declared sync: false, so Render prompts for it in the dashboard and it never enters git;
JWT_SECRET is generated per environment. Free instances sleep when idle, so the first
request after a nap is slow.
Hugging Face Spaces — deploy/huggingface/ holds a Spaces-ready Dockerfile (port 7860,
non-root user) and the Space README.md with the required front matter. Copy both to a
Docker Space along with requirements.txt, app/, mcp_server/ and evals/, then add
GOOGLE_API_KEY under Settings → Variables and secrets. Step-by-step:
deploy/huggingface/README.md.
Secrets are always injected as environment variables. .env is gitignored;
.env.example documents every variable.
Configuration
| Variable | Default | Purpose |
|---|---|---|
GOOGLE_API_KEY |
— | Gemini key (free tier). Absent → chat returns 503, everything else works. |
MODEL_NAME |
gemini-2.5-flash |
Gemini model id |
JWT_SECRET / JWT_ALGORITHM |
dev value / HS256 |
Token signing — override in any deployment |
VECTOR_BACKEND |
keyword |
keyword (offline) or chroma (semantic) |
CHECKPOINTER |
memory |
memory for multi-turn, none for stateless |
LOG_LEVEL / LOG_FORMAT |
INFO / json |
json for aggregators, text for humans |
OTEL_TRACES_EXPORTER |
console |
console, otlp, or none |
OTEL_EXPORTER_OTLP_ENDPOINT |
http://localhost:4318 |
Collector base URL for otlp |
Tech
Python · FastAPI · Pydantic · LangChain · LangGraph · MCP · SSE · OAuth2/JWT · OpenTelemetry · Chroma/pgvector · Docker · GitHub Actions
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。