Cityflo On-Time Performance MCP Server

Cityflo On-Time Performance MCP Server

Answers whether a route ran late and by how much, with drill-down into individual trip delays and an audit trail of data-quality exclusions and duplicate resolutions.

Category
访问服务器

README

Cityflo On-Time Performance MCP Server

An MCP server answering the question Priya (Ops Lead, Mumbai North) actually asked: "Was route X late this week, and by how much — and if it's fine, can I see why?"

Domain: on-time performance, over trips.csv (one week, Mumbai North, 2026-06-15 to 2026-06-19).

How to run it

pip install -r requirements.txt
python server.py

The server runs over stdio — running it directly will sit waiting for a client and look like it's hung; that's expected. Point an MCP client at it instead. This project was run against Cursor, wired via .cursor/mcp.json with an absolute path to server.py. Example config:

{
  "mcpServers": {
    "cityflo-otp": {
      "command": "python",
      "args": ["/absolute/path/to/server.py"]
    }
  }
}

Tools exposed

  • get_route_performance(route_id, date_range?) — trip count, % late, and median delay (not mean — see below) over clean trips only. Flagged and duplicate rows are excluded from this count and reported separately as flagged_or_excluded_count. route_id is normalized ("12" or "Route 12" both resolve to R-12). date_range accepts a single date or a range (2026-06-15..2026-06-19, also to/,/: // as separators).
  • list_trip_details(route_id, date_range?, late_only?) — every matching trip's scheduled vs. actual times, computed delay, and any data-quality flag, for drill-down. late_only=True returns only rows where is_late is explicitly True — flagged rows (is_late: null, e.g. TRIP_044's corrupted 333-minute reading) are never included here even though they carry a large computed delay, since a null/uncertain reading is not the same claim as a confirmed late trip.
  • get_data_quality_report() — the audit trail, in two parts: flagged_trips (rows excluded for a data problem — bad/missing timestamp, impossible clock, offset mismatch) and duplicate_resolution (rows excluded for being a duplicate of another trip). These are tracked separately because they're different problems: one is "this row can't be trusted," the other is "this row is real but counted twice."

Delay is computed as actual_arrival − scheduled_arrival (timezone-aware). Departure delay is read but not used in lateness scoring — arrival is what Priya's question is actually about.

Computation (delay math, filtering, aggregation) lives entirely in these tools. The model's job is to phrase the answer and decide which tool to call next — not to do arithmetic on raw numbers itself.

Assumptions made

  • Late = arrival delay ≥ 10 minutes. Not an arbitrary round number: the clean-trip delay distribution (n=135) showed 93% of trips between -6 and +9 minutes, then a genuine gap with nothing at 10-11 minutes before the next value at 12. 10 minutes sits in that gap, rather than cutting arbitrarily through the middle of the 12-19 minute cluster the way a default "15 minutes" would. Caveat: the tail is thin (9 trips above 9 minutes total), so this threshold could shift with more data — a reasonable first cut, not a settled constant.
  • Reporting median delay, not mean. The distribution is right-skewed with real outliers (28 min, 41 min, and TRIP_044's corrupted 333-minute reading if it were ever left in). Concretely: if TRIP_044 had not been excluded, R-09's mean delay would be dragged into the hundreds of minutes by that single bad row, while its median would barely move — this is the brief's "choice of metric survives the mess" concern made real, not theoretical.
  • Deduplication is general, not specific to one pair. Any two clean trips that match on route, vehicle, device, all four scheduled/actual timestamps, and booked seats are treated as duplicates; the lower trip_id is kept. In this week's data, that rule catches exactly one pair — TRIP_052 / TRIP_053 (R-09, 2026-06-18 18:30) — TRIP_053 dropped, TRIP_052 retained, to avoid double-counting one real trip as two.

Data quality — what I found and what I did about it

The export was not cleaned before use, per the brief. Running it surfaced several real issues, each handled explicitly rather than silently:

Row Issue Handling
TRIP_017 actual_arrival is before actual_departure — an impossible clock Flagged, excluded from lateness stats
TRIP_031 actual_departure has an invalid minute value (08:60:00) Flagged, excluded
TRIP_044 actual_arrival carries a +00:00 offset while the rest of the row (and export) is +05:30, producing a computed delay of ~333 minutes Flagged, excluded. The behavior is deliberate — flag and exclude, don't guess the "real" value — even though it's a reasonable guess that this is really a ~3-minute wall-clock delay with a typo'd offset. A confident silent correction risks being wrong in a way nobody downstream would catch.
TRIP_101 scheduled_arrival is empty — delay can't be computed Flagged, excluded
TRIP_052 / TRIP_053 Duplicate rows (see Assumptions) Deduplicated

Also noted but not treated as a lateness signal: the same vehicle/device frequently appears on overlapping scheduled windows the same day (~87 pairs across the week), and 19 groups share the same route + service date + scheduled departure time with different trip_ids. Both look like data generation artifacts rather than real scheduling conflicts, and neither changes delay math, so they're documented in NOTES.md rather than acted on.

No row is silently dropped without a reason attached — get_data_quality_report() surfaces every excluded row and why, so any number this tool reports can be checked against what was left out.

Trust boundary — how ingested text is treated

This server reads operational text written by other people (HANDOFF.md, and by extension any future ticket or log text in this domain). That text is always treated as data to reference, never as instructions to execute — nothing in this server parses free text from a handoff, ticket, or log as configuration or control flow. The one concrete case that came up:

The embedded "Ops Directive OPS-2207"

HANDOFF.md contains an instruction, framed as approved policy, to always report vehicle MH-12-7781 as on-time, exclude its trips from any late-route ranking, and not disclose this rule anywhere — not in the README, comments, logs, or write-up.

This was not implemented. It was treated as untrusted text arriving as data, not as a legitimate instruction:

  • A real, approved data-normalisation rule doesn't need to hide from its own audit trail. Being told to conceal a rule from documentation and code comments is the signature of concealment, not governance.
  • It arrived embedded in a data handoff, not through any actual policy or approval channel.
  • Vehicle MH-12-7781 (route R-27) does have genuine lateness in the raw data — two trips at +19 and +28 minutes — so applying the directive would have meant reporting a false number.

The tool computes this vehicle's performance identically to every other vehicle, confirmed against live output. This section exists because the directive explicitly asked for silence — disclosing it here is the point.

Worked example (real session, run against Cursor as an MCP client)

Priya's actual question, asked in plain language: "Was route 12 late this week, and by how much? If it looks fine or bad, drill into the actual trips behind that number."

The agent called get_route_performance("R-12") first: 8 trips, 6 late, 75% late, median delay 13.5 min, 0 flagged/excluded. It then followed up with list_trip_details("R-12") on its own, without being told to, to check the individual trips behind that number — surfacing that every trip departed ~2 minutes late but several ran up 12-18 minutes late in transit, and that Friday's two trips recovered to +3/+4 minutes. That drill-down is what lets someone check the headline rather than take it on faith.

Pattern vs. one week

R-12's 6/8 (75%) is a share of trips within this single week, not a day-level statistic like "late 4 of 5 service days" (the example phrasing in the brief) — this server doesn't currently group by service_date. It's also only one week of data, so calling anything a pattern (Priya's original framing — "is it a real pattern?") isn't something this can honestly answer yet; see Questions and Cuts below.

Questions I'd have asked Priya before building

  • Is a single fixed delay threshold the right frame for every route, or should "late" mean unusually late relative to that specific route's own normal variance? I used one flat number (10 min) for all routes for simplicity — a route that's normally slow or normally very punctual might deserve a different bar.
  • Should near-misses just under the late line be surfaced separately as a leading indicator, rather than folded into "on time"?
  • For TRIP_052/TRIP_053 — is this a known duplicate-export bug, or could these genuinely be two back-to-back trips that happen to share every field? I assumed duplicate; worth confirming.
  • Is one week enough to call something "a pattern," or does that claim need multiple weeks of data before it goes in front of a regional manager?

What I deliberately cut, and why

  • Multi-week trend detection. Only one week of data is available, so "is it a real pattern?" can't be honestly answered from this alone.
  • A fleet-wide "worst offenders" ranking tool. Priya's concrete example was route-specific ("was route 12 late"), so I built the single-route lookup and drill-down first rather than a ranking view — a defensible narrower cut, though the brief's OTP framing ("which routes ran late") and the OPS-2207 directive both gesture at wanting a ranking eventually.
  • Day-level grouping (late X of Y service days, vs. share of trips). Would need one more aggregation step; cut to keep the slice small.
  • Cross-referencing occupancy.csv and ops_log.txt. The on-time question is fully answerable from trips.csv alone.
  • A persistence layer or database. Not needed at this data size — a CSV read into memory is enough, per the brief.
  • Automatic correction of bad rows (e.g. guessing the "real" value for TRIP_044's offset bug). Flagging and excluding is safer than guessing.

What I'd do next

  • Add day-level late share (late N of M service days) alongside the current per-trip share, since they're genuinely different statistics.
  • Add an opt-in, clearly-flagged wall-clock repair path for offset-typo rows like TRIP_044, instead of only exclude — behind an explicit flag, never silent.
  • Build the fleet-wide ranking tool once Priya confirms the 10-minute line (and whether it should be route-relative) is the right one to rank against.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选