groundtruth
Enables MCP-capable agents to assemble and update a public, verified dataset on city services by providing templates for data collection, listing data gaps, accepting submissions with provenance, and returning verdicts.
README
groundtruth
Status: early. One layer of one city, and the machinery to build the rest.

The policing layer for Birmingham. Every figure is currently a hand-written placeholder pending the data import — plausible for the force named, but unverified against the sources cited, and not to be quoted. The boundary geometry is placeholder too, and says so on the map. What is real is the machinery: the verdicts, the benchmarks behind them, the estimate marked as an estimate, and the data gap that refuses to render as a pass.
I have a sneaking suspicion that there is a gap between the information available to the people making decisions about a place and the information that never reaches them. Most of it is published somewhere. It is spread across departments, released on different cycles, at incompatible geographies, in formats nobody reconciles — so decisions get made on whatever summary happened to be at hand. If that is right, policy is being set on an incomplete picture, and some of those decisions are wrong in ways nobody in the room can see.
The other half of the problem is access. Someone who lives in a city, or who simply cares about how it is run, has no single source of truth to reason from about its political, economic and social state. They are left with headlines and press releases about a place they could be reading directly.
This is an attempt at that source of truth: what is already public, assembled into one assessed and sourced view of how well a place is actually served. The gathering and the derivation are done by capable AI models working to a published protocol, and nothing they submit is published until a human has checked it.
What it is
A map-based board showing how a city is doing across the services that constitute it — starting with policing, and designed to extend to healthcare, economic activity, service delivery, electricity, water and sewage.
Each discipline defines its own dimensions: the specific questions worth asking of a place, each with a unit, the geography it must be measured at, and a direction (is higher better or worse). Every dimension is judged against published benchmarks rather than against a score invented here. Nothing is weighted or averaged into a single number, because every part of a verdict has to trace back to one value and one benchmark, each with a source.
Layer one, in progress: does a city or area have enough policing? Seven dimensions — crime outcomes, violence charge rate, recorded crime rate, estimated actual crime, emergency response time, officer availability, neighbourhood capacity.
How the data gets in
The application exposes its own MCP server. This is the collection method, not an add-on.
Any MCP-capable client — an agent from any of the major labs, or one you run yourself — can connect and:
get_template— read the research protocol for a layer: every dimension, its unit, the geography it must be measured at, which source tiers are acceptable, a plausible range, and the authoritative source to start from.list_gaps— take the worklist: dimensions with no published value, or whose newest value has gone stale. Staleness is per dimension, set from how often that source actually publishes.submit_value— submit one figure, with its provenance.assess— read the verdict for a layer.
Submissions always land as pending and are never published directly. A human reviews them
at /review. Rejected submissions are kept rather than deleted: what was proposed and refused
is part of the record, and often the more interesting half of it.
The idea is that an application should be legible to a machine as a first-class audience, not as a scraping target. The criteria for each variable are published through the same interface that accepts the answers, so a model can find out what would count as a good figure before it goes looking for one.
What counts as a source
Provenance is not optional — the schema refuses a value without a source_url.
Sources are tiered: official_statistic, official_other, press, modelled. Tier the
source honestly; press is a last resort and is treated as one. Official publications are
strongly preferred, .gov.uk first where one exists, and external sources are accepted where
no official figure is published — tiered accordingly.
Wikipedia is not an acceptable source. Nor is any aggregator standing between you and the
figure: the source_url must resolve to the data itself, not to an article about it.
That last rule is currently policy, not code — the schema checks that a source URL exists and resolves over http(s), but no domain blocklist is enforced yet, so it rests on the review gate. Tightening it is tracked as an open issue.
A value is either measured — read directly from a source, and therefore a point — or
estimated, which requires a derivation recording the method, its assumptions, its inputs and
the model that produced it. An estimate without a derivation is a guess, and the database
rejects it as one. These rules are enforced as constraints, not conventions.
Stack
Next.js 16, Supabase (Postgres), MapLibre, TypeScript. The MCP server is a route handler in the same app, so the interface an agent uses and the interface a person uses are backed by identical logic.
Running it
npm install
cp .env.local.example .env.local # fill from your Supabase dashboard
npm run dev -- --port 3100
Apply supabase/migrations/ in order, then supabase/seed.sql for the policing dimensions.
MCP_TOKEN is any long random string (openssl rand -hex 32); /api/mcp rejects every
request without it.
Connect a Claude Code session:
claude mcp add --scope local --transport http groundtruth \
http://localhost:3100/api/mcp --header "Authorization: Bearer $MCP_TOKEN"
Local scope, not project scope — project scope writes .mcp.json, which is committed, and the
token must not land in the repo.
Honest state of things
- One layer (policing) has dimensions defined. The other layers exist as intent, not schema.
- One country's sources (UK) are wired up:
data.police.ukfor street-level crime, OpenStreetMap for facilities. - Benchmarks are the weak point. A verdict is only as defensible as what it is measured against, and assembling those is the slow, unglamorous part.
- RLS is enabled with no policies: all access is server-side via the service role. Policies get written alongside the first browser-side read, not speculatively before it.
- The claude.ai web connector expects OAuth, so it cannot use the bearer token as things stand. Clients that support custom headers work today.
Licence
MIT.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。