Travel Assistant MCP Server
Exposes travel planning data and constraint checking as MCP tools over stdio, enabling AI agents to query reference information and evaluate plan validity.
README
AI Agent Based Travel Assistant Claude API + MCP
What exists here
| Piece | File | What it does |
|---|---|---|
| Sandbox data layer | src/tp_mcp/sandbox.py |
Loads TravelPlanner queries, parses reference_information into per-city sections |
| MCP server | src/tp_mcp/server.py |
Exposes the sandbox as MCP tools over stdio |
| Agent loop | src/agent/loop.py |
Claude tool-use loop driving the MCP session |
| Prompts | src/agent/prompts.py |
Versioned system prompts (planner, sole-planning, chat) |
| Plan schema | src/agent/schema.py |
Pydantic contract matching the benchmark's plan format |
| Evaluator | src/eval/constraints.py |
Commonsense + hard constraint checks, benchmark-style metrics |
| Batch runner | src/eval/run_eval.py |
Runs N queries, writes traces + a metrics summary |
| UI | src/ui/app.py |
Streamlit chat for participant testing |
| Tests | tests/test_constraints.py |
9 offline tests, no API key needed |
Setup (30 minutes)
python -m venv .venv
source .venv/bin/activate # Windows cmd: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # Windows cmd: copy .env.example .env
# then paste your key in. The file must be
# named exactly .env -- not claude.env.
python scripts/prepare_data.py
pytest -q # 9 tests, no API calls
python scripts/check_setup.py # key, dataset, MCP server, one API call
Then, first real run. Use run.py and everything works the same on every
platform -- no PYTHONPATH to set:
python run.py eval.run_eval --mode sole-planning --n 5 --level easy # cheap, ~1 min
python run.py agent.loop 0 # one agentic run, prints the full trace
python run.py eval.run_eval --mode tool-use --n 5 --level easy
streamlit run src/ui/app.py
<details> <summary>Windows command equivalents</summary>
| Unix | cmd.exe | PowerShell |
|---|---|---|
ls -a |
dir /a |
ls -Force |
cat f |
type f |
cat f |
cp a b |
copy a b |
cp a b |
mv a b |
ren a b |
mv a b |
export X=y |
set X=y |
$env:X="y" |
pytest, check_setup.py and run.py all resolve src/ themselves, so none
of them need PYTHONPATH.
</details>
Debug the MCP server on its own - do this before blaming the agent:
TP_SPLIT=validation TP_QUERY_ID=0 npx @modelcontextprotocol/inspector \
python -m tp_mcp.server
Architecture
Streamlit UI ──┐
├──► agent/loop.py ──── Anthropic Messages API (Claude)
eval runner ───┘ │ ▲
│ tool_use blocks │ tool_result blocks
▼ │
MCP ClientSession ────────────┘
│ stdio (JSON-RPC)
▼
tp_mcp/server.py (FastMCP)
│
▼
TravelPlanner sandbox (reference_information / CSV database)
The whole point is the seam at ClientSession. The model side knows nothing
about TravelPlanner; the tool side knows nothing about Claude. Swapping the
sandbox for live weather and maps APIs, or Claude for another model, touches one
side only. Say that in your design chapter - it is the architectural claim your
dissertation is actually testing.
The build plan
Do these in order. Each stage ends with something you can show your supervisor.
Stage 1 - get a number on the board (1st week).
Run --mode sole-planning on 20 easy validation queries. No tools, all reference
information pasted into the prompt. You now have a delivery rate and a pass rate.
This is your baseline and it de-risks everything: prompts, parsing, evaluation and
reporting are all proven before agentic complexity enters.
Stage 2 - make the agent work for its information (weeks 2–3).
Run --mode tool-use on the same queries. It will be worse. Read the traces in
results/, find the top three failure modes, fix them, re-run. That loop -
measure, diagnose, fix, re-measure - is the technical contribution, and the
traces are your evidence.
Stage 3 - sharpen the comparison (weeks 4–5).
Add a condition or two: with vs without the notebook tools, with vs without
check_budget, one model vs another, easy vs medium vs hard. Ablations are what
turn "I built a thing" into "I tested a claim". Three runs per condition, since
LLM output varies.
Stage 4 - extend the tool surface (weeks 5–6). Add a second MCP server (weather, or a real maps API). This is where you demonstrate that MCP composition works: two servers, one agent, no changes to the loop. Your draft promises live external tools - this is where you deliver it.
Stage 5 - human evaluation (weeks 7–8). Streamlit + 5–10 participants + the questionnaire your ethics form covers. Log every session. Report it alongside the benchmark numbers, not instead of them.
Stage 6 - write up (weeks 8+).
Methodology chapter = this README plus the design rationale in the docstrings.
Results chapter = results/*_summary.json plus failure-mode analysis. The
literature review you already have becomes the framing, which is exactly the
proportion your supervisor is asking for.
Honest limitations
- The sandbox is built from each query's
reference_information, not the full 4M-record database. The tool interface is faithful; the search space is narrower. Swapping in the full database is a contained change tosandbox.py. within_sandboxandreasonable_city_routeare therefore narrower than the benchmark's own versions. Mark them as such in your results table.- Budget checking parses prices from free text and returns "not scored" when it cannot. That is deliberate - an unparseable field must never count as a pass.
- The data is fixed at 2022. The prototype is not bookable and the UI says so.
mcpis pinned to>=1.27,<2. The v2 SDK rework lands around the 2026-07-28 spec release with breaking changes. Record the exact version you ran with (pip freeze > results/environment.txt) so your results stay reproducible.
Cost control
Each --mode tool-use query is roughly 10–25 model calls with growing context.
Start with --n 5, use a cheaper model for iteration and your best model for the
final reported runs, and keep --concurrency low while debugging so the traces
stay readable. Report which model produced which table.
�
�
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。