support-agent-mcp

support-agent-mcp

Exposes order status lookup and knowledge base search tools from the Support Agent AI over MCP, enabling MCP clients to handle customer support queries with grounded, citation-backed answers.

Category
访问服务器

README

Support Agent + MCP

live demo ci python license

Live demo: support-agent-mcp.onrender.com/docs — interactive Swagger UI. Health: /healthz. Hosted on Render's free tier, which sleeps after ~15 min idle, so the first request may take 30–60s to wake.

A production-shaped customer-support AI agent: a FastAPI service where a LangGraph agent answers customer questions by calling tools — order lookup, OAuth-scoped refunds, and a grounded knowledge base — with token streaming, conversation memory, OpenTelemetry traces and structured JSON logs. The same tools are re-exposed over the Model Context Protocol (MCP) for any MCP client.

Runs entirely on free infrastructure (Gemini free tier, a free Render/Spaces dyno) and the whole test suite runs offline, with no API key.

Flagship features → what they demonstrate

Capability Where AI-Engineer JD bullet it answers
Secure REST API + validation FastAPI + Pydantic v2 (app/main.py, app/schemas.py) Build and expose secure APIs
Agent workflow / tool routing LangGraph ReAct agent (app/agent.py) Design agentic workflows
Tool authorization OAuth2 password + JWT scopes; refunds gated on refund:write (app/auth.py) Guardrails; least-privilege tool access
LLM integration + bounded retry Gemini via LangChain, exponential backoff (app/llm.py) Integrate LLM providers reliably
Retrieval grounding + citations Pluggable KB: offline keyword / semantic Chroma (app/vectorstore.py) RAG, grounded answers, anti-hallucination
Token streaming (SSE) POST /chat/stream (app/streaming.py) Async Python; responsive UX
Conversation memory LangGraph checkpointer keyed by session_id (app/agent.py) Stateful, multi-turn agents
Distributed tracing OpenTelemetry spans on agent run, every tool call, the LLM call (app/tracing.py) Observability for LLM systems
Structured logging JSON logs + X-Request-ID correlation (app/logging_config.py) Production operability
Offline eval harness Scored behavioural evals (evals/) Measure agent quality, not vibes
MCP service mcp_server/server.py (FastMCP, stdio) Interoperable tool servers
Containers + free deploy Dockerfile, docker-compose.yml, render.yaml, deploy/huggingface/ Containerization and deployment
Tests + CI gate 56 tests, .github/workflows/ci.yml Testing discipline

What makes this production-shaped

It is not the feature list — it is the failure behaviour:

  • Authorization is enforced at the tool, not in the prompt. The model can decide to refund; without refund:write on the caller's token the tool still refuses, and the denial is recorded as a span attribute for audit.
  • Degrades instead of dying. No API key → /healthz and /token still serve and chat returns an actionable 503. No vector service → the keyword retriever takes over. No collector → traces go to stdout. A broken stream closes with an error event rather than a half-written response.
  • Every answer is attributable. Policy replies carry citations recovered from the retriever's own output, and every log line carries the request id also returned in X-Request-ID.
  • The offline/online split is deliberate. Pure logic (SSE translation, turn slicing, eval graders, exporter selection) is separated from I/O, so CI proves behaviour with no network and no key — and the same graders score a live Gemini run locally.

Architecture

flowchart LR
    subgraph Client
      U[HTTP client / UI]
      MCPC[MCP client<br/>Claude, IDEs]
    end

    U -->|POST /chat<br/>POST /chat/stream<br/>Bearer token| API[FastAPI<br/>JSON logs + request id]
    API -->|JWT scopes| A["LangGraph ReAct agent<br/>(Gemini)"]
    A <-->|thread_id| M[(MemorySaver<br/>checkpointer)]
    A --> T1[get_order_status]
    A --> T2["request_refund<br/>needs refund:write"]
    A --> T3[search_kb]
    T3 --> KB[(KB: keyword / Chroma)]
    API -->|reply · tool_calls · citations<br/>or SSE token stream| U
    MCPC --> S[FastMCP server] --> T1 & T3

    A -.spans.-> OTEL{{OpenTelemetry<br/>console / OTLP}}
    T1 & T2 & T3 -.spans.-> OTEL

Quickstart

python -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
cp .env.example .env          # add a free key from https://aistudio.google.com/apikey
pytest -q                     # runs fully offline, no key needed
uvicorn app.main:app --reload

Or with Docker:

docker compose up --build     # spans stream to the container log

Interactive API docs at http://localhost:8000/docs.

Demo

# 1) Authenticate as an agent (gets refund:write). Try 'customer/customer' to see a denial.
TOKEN=$(curl -s -X POST localhost:8000/token -d 'username=agent&password=agent' | python -c 'import sys,json;print(json.load(sys.stdin)["access_token"])')

# 2) Grounded policy answer (returns citations)
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{"message":"How long do refunds take?"}'

# 3) Authorized action
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{"message":"Refund order A1001, it was defective."}'

# 4) Streamed answer — tokens and tool steps as they happen
curl -N -X POST localhost:8000/chat/stream -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{"message":"How long do refunds take?"}'

# 5) Multi-turn memory — reuse the session_id and the agent remembers the order
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{"message":"Where is order A1002?","session_id":"demo-1"}'
curl -s -X POST localhost:8000/chat -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{"message":"Can I refund it?","session_id":"demo-1"}'

Endpoints

Method Path Notes
GET /healthz liveness + which backends are live
POST /token OAuth2 password flow → JWT with scopes
POST /chat JSON reply with tool_calls + citations
POST /chat/stream Server-Sent Events: start, tool_call, tool_result, token, done, error

Observability

Tracing is controlled by OTEL_TRACES_EXPORTER: console (default — spans to stdout, zero infrastructure), otlp (any OTLP/HTTP collector via OTEL_EXPORTER_OTLP_ENDPOINT), or none. Spans cover the agent run, each tool call and the LLM call:

{"name": "tool.request_refund", "attributes": {"tool.name": "request_refund", "authz.allowed": false}}

Logs are one JSON object per line, each carrying the request_id that is also returned in the X-Request-ID response header (LOG_FORMAT=text for a human-readable dev view):

{"ts": "…", "level": "INFO", "logger": "support-agent.access", "message": "request",
 "request_id": "9e8d65214aae4af4", "method": "GET", "path": "/healthz", "status": 200, "duration_ms": 0.75}

Evals

evals/ scores the three behaviours the agent is actually hired for: picking the right tool, refusing an unauthorized refund, and answering policy questions from the KB with a citation.

python -m evals.runner --min-pass-rate 0.8   # live Gemini run; exits non-zero below the bar

The graders (evals/scorers.py) are pure functions, so CI exercises them against fixtures with no key; the live end-to-end run is pytest.mark.skipif-ed off when GOOGLE_API_KEY is unset. That is what keeps CI hermetic while the same rubric grades a real model locally.

MCP

python -m mcp_server.server   # exposes order_status + knowledge_base over stdio

Refunds are intentionally not exposed over MCP: that action requires a scope-bearing session, which the local stdio transport does not carry.

Deploy (free tiers)

Render — render.yaml is a ready blueprint (free plan, Docker runtime, health check on /healthz). Push the repo, then Render → New → Blueprint → select it. GOOGLE_API_KEY is declared sync: false, so Render prompts for it in the dashboard and it never enters git; JWT_SECRET is generated per environment. Free instances sleep when idle, so the first request after a nap is slow.

Hugging Face Spaces — deploy/huggingface/ holds a Spaces-ready Dockerfile (port 7860, non-root user) and the Space README.md with the required front matter. Copy both to a Docker Space along with requirements.txt, app/, mcp_server/ and evals/, then add GOOGLE_API_KEY under Settings → Variables and secrets. Step-by-step: deploy/huggingface/README.md.

Secrets are always injected as environment variables. .env is gitignored; .env.example documents every variable.

Configuration

Variable Default Purpose
GOOGLE_API_KEY — Gemini key (free tier). Absent → chat returns 503, everything else works.
MODEL_NAME gemini-2.5-flash Gemini model id
JWT_SECRET / JWT_ALGORITHM dev value / HS256 Token signing — override in any deployment
VECTOR_BACKEND keyword keyword (offline) or chroma (semantic)
CHECKPOINTER memory memory for multi-turn, none for stateless
LOG_LEVEL / LOG_FORMAT INFO / json json for aggregators, text for humans
OTEL_TRACES_EXPORTER console console, otlp, or none
OTEL_EXPORTER_OTLP_ENDPOINT http://localhost:4318 Collector base URL for otlp

Tech

Python · FastAPI · Pydantic · LangChain · LangGraph · MCP · SSE · OAuth2/JWT · OpenTelemetry · Chroma/pgvector · Docker · GitHub Actions

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选