Travel Assistant MCP Server

Travel Assistant MCP Server

Exposes travel planning data and constraint checking as MCP tools over stdio, enabling AI agents to query reference information and evaluate plan validity.

Category
访问服务器

README

AI Agent Based Travel Assistant Claude API + MCP

What exists here

Piece File What it does
Sandbox data layer src/tp_mcp/sandbox.py Loads TravelPlanner queries, parses reference_information into per-city sections
MCP server src/tp_mcp/server.py Exposes the sandbox as MCP tools over stdio
Agent loop src/agent/loop.py Claude tool-use loop driving the MCP session
Prompts src/agent/prompts.py Versioned system prompts (planner, sole-planning, chat)
Plan schema src/agent/schema.py Pydantic contract matching the benchmark's plan format
Evaluator src/eval/constraints.py Commonsense + hard constraint checks, benchmark-style metrics
Batch runner src/eval/run_eval.py Runs N queries, writes traces + a metrics summary
UI src/ui/app.py Streamlit chat for participant testing
Tests tests/test_constraints.py 9 offline tests, no API key needed

Setup (30 minutes)

python -m venv .venv
source .venv/bin/activate          # Windows cmd:  .venv\Scripts\activate
pip install -r requirements.txt

cp .env.example .env               # Windows cmd:  copy .env.example .env
                                   # then paste your key in. The file must be
                                   # named exactly .env -- not claude.env.
python scripts/prepare_data.py

pytest -q                          # 9 tests, no API calls
python scripts/check_setup.py      # key, dataset, MCP server, one API call

Then, first real run. Use run.py and everything works the same on every platform -- no PYTHONPATH to set:

python run.py eval.run_eval --mode sole-planning --n 5 --level easy   # cheap, ~1 min
python run.py agent.loop 0                    # one agentic run, prints the full trace
python run.py eval.run_eval --mode tool-use --n 5 --level easy
streamlit run src/ui/app.py

<details> <summary>Windows command equivalents</summary>

Unix cmd.exe PowerShell
ls -a dir /a ls -Force
cat f type f cat f
cp a b copy a b cp a b
mv a b ren a b mv a b
export X=y set X=y $env:X="y"

pytest, check_setup.py and run.py all resolve src/ themselves, so none of them need PYTHONPATH. </details>

Debug the MCP server on its own - do this before blaming the agent:

TP_SPLIT=validation TP_QUERY_ID=0 npx @modelcontextprotocol/inspector \
  python -m tp_mcp.server

Architecture

  Streamlit UI ──┐
                 ├──► agent/loop.py ──── Anthropic Messages API (Claude)
  eval runner ───┘         │                      ▲
                           │  tool_use blocks     │  tool_result blocks
                           ▼                      │
                    MCP ClientSession ────────────┘
                           │  stdio (JSON-RPC)
                           ▼
                    tp_mcp/server.py  (FastMCP)
                           │
                           ▼
              TravelPlanner sandbox (reference_information / CSV database)

The whole point is the seam at ClientSession. The model side knows nothing about TravelPlanner; the tool side knows nothing about Claude. Swapping the sandbox for live weather and maps APIs, or Claude for another model, touches one side only. Say that in your design chapter - it is the architectural claim your dissertation is actually testing.


The build plan

Do these in order. Each stage ends with something you can show your supervisor.

Stage 1 - get a number on the board (1st week). Run --mode sole-planning on 20 easy validation queries. No tools, all reference information pasted into the prompt. You now have a delivery rate and a pass rate. This is your baseline and it de-risks everything: prompts, parsing, evaluation and reporting are all proven before agentic complexity enters.

Stage 2 - make the agent work for its information (weeks 2–3). Run --mode tool-use on the same queries. It will be worse. Read the traces in results/, find the top three failure modes, fix them, re-run. That loop - measure, diagnose, fix, re-measure - is the technical contribution, and the traces are your evidence.

Stage 3 - sharpen the comparison (weeks 4–5). Add a condition or two: with vs without the notebook tools, with vs without check_budget, one model vs another, easy vs medium vs hard. Ablations are what turn "I built a thing" into "I tested a claim". Three runs per condition, since LLM output varies.

Stage 4 - extend the tool surface (weeks 5–6). Add a second MCP server (weather, or a real maps API). This is where you demonstrate that MCP composition works: two servers, one agent, no changes to the loop. Your draft promises live external tools - this is where you deliver it.

Stage 5 - human evaluation (weeks 7–8). Streamlit + 5–10 participants + the questionnaire your ethics form covers. Log every session. Report it alongside the benchmark numbers, not instead of them.

Stage 6 - write up (weeks 8+). Methodology chapter = this README plus the design rationale in the docstrings. Results chapter = results/*_summary.json plus failure-mode analysis. The literature review you already have becomes the framing, which is exactly the proportion your supervisor is asking for.


Honest limitations

  • The sandbox is built from each query's reference_information, not the full 4M-record database. The tool interface is faithful; the search space is narrower. Swapping in the full database is a contained change to sandbox.py.
  • within_sandbox and reasonable_city_route are therefore narrower than the benchmark's own versions. Mark them as such in your results table.
  • Budget checking parses prices from free text and returns "not scored" when it cannot. That is deliberate - an unparseable field must never count as a pass.
  • The data is fixed at 2022. The prototype is not bookable and the UI says so.
  • mcp is pinned to >=1.27,<2. The v2 SDK rework lands around the 2026-07-28 spec release with breaking changes. Record the exact version you ran with (pip freeze > results/environment.txt) so your results stay reproducible.

Cost control

Each --mode tool-use query is roughly 10–25 model calls with growing context. Start with --n 5, use a cheaper model for iteration and your best model for the final reported runs, and keep --concurrency low while debugging so the traces stay readable. Report which model produced which table. � �

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选