datagovin-mcp
Enables natural-language access to India's Open Government Data platform, letting users discover datasets, inspect schemas, and query live data from data.gov.in.
README
datagovin-mcp
Natural-language access to India's Open Government Data platform, data.gov.in — 235,000+ public datasets covering air quality, agriculture, health, fuel prices, census, education, rainfall, railways, crime, budgets and more.
Two faces, one codebase:
- An MCP server — over stdio for local clients, or over Streamable HTTP so any MCP client connects by URL: Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, ChatGPT connectors, or anything built on an MCP SDK.
- A website — instant catalog search plus natural-language answers, backed by the exact same tools.
The server ships no data of its own. Discovery runs against a local full-text index built from data.gov.in's own catalog endpoint; every row you actually read is a live call to data.gov.in using your own free API key.
Why this exists
data.gov.in has an enormous catalog but no full-text search API a program can call — the normal workflow is to browse the website and copy a dataset's resource ID off its "API" button. That's a poor fit for a language model.
This project closes the gap. It harvests the platform's /lists endpoint into a
local SQLite FTS5 index — 235,241 datasets in about 150 seconds, no API key
required — so a model can go from "what's the AQI in Delhi right now?" to real
rows without anyone hunting for a UUID. It also absorbs the upstream API's rough
edges (case-sensitive filters, occasional CSV responses, a last-page pagination
quirk) so the model doesn't have to.
Tools
| Tool | What it does |
|---|---|
search_datasets(query, limit, sector) |
BM25-ranked full-text search across the whole catalog. |
list_sectors(limit) |
Sectors present in the catalog, with dataset counts. |
get_dataset_info(resource_id) |
Live schema: title, description, row count, exact field names + types. |
query_dataset(resource_id, filters, fields, sort, limit, offset) |
Pull actual filtered rows, live. |
catalog_status() |
How many datasets are indexed — distinguishes "no matches" from "not harvested yet". |
Setup
1. Install.
git clone https://github.com/<your-username>/datagovin-mcp.git
cd datagovin-mcp
python -m venv .venv && source .venv/bin/activate
pip install -e . # MCP server only
pip install -e ".[web]" # + the website
2. Build the search index. No API key needed for this step.
datagovin-harvest
Harvesting the data.gov.in catalog (no API key required)...
indexed 235,000/235,241 datasets (100%, 1,566/s)
Indexed 235,241 datasets in 150.2s -> ~/Library/Caches/datagovin-mcp/catalog.sqlite3 (346.2 MB)
Until you run this, search falls back to a small bundled seed catalog — the server still works, it just knows about three datasets. Re-run it any time to refresh; hand-curated entries are preserved.
3. Get a free API key — needed to read rows, not to search. Register at data.gov.in and generate one from your profile page.
cp .env.example .env # then paste your key into it
The .env file is read automatically. Exporting DATA_GOV_IN_API_KEY works too,
and a real environment variable always wins over the file.
Connecting a client
Local (stdio)
Add to your MCP client config (claude_desktop_config.json or equivalent):
{
"mcpServers": {
"datagovin": {
"command": "/absolute/path/to/datagovin-mcp/.venv/bin/python",
"args": ["/absolute/path/to/datagovin-mcp/server.py"],
"env": { "DATA_GOV_IN_API_KEY": "your_key_here" }
}
}
}
Remote (Streamable HTTP) — connects from anywhere
python server.py --transport http --host 0.0.0.0 --port 8000
# MCP endpoint: http://<host>:8000/mcp
Then point any MCP client at the URL:
{
"mcpServers": {
"datagovin": { "url": "https://your-host.example.com/mcp" }
}
}
Add --stateless to run several replicas behind a load balancer.
Before exposing this publicly, put it behind TLS and authentication. The server has no auth of its own, and it spends your data.gov.in API key on every request it serves.
The website
export ANTHROPIC_API_KEY=sk-ant-... # optional — enables the "Ask" button
datagovin-web # http://127.0.0.1:8000
One process serves everything:
| Route | |
|---|---|
/ |
search UI — type-ahead catalog search, click a dataset for its live schema and sample rows |
/api/search?q= |
BM25 search as JSON, no LLM involved |
/api/dataset/{id} |
live schema |
/api/dataset/{id}/rows |
live rows; any extra query param becomes an upstream filter |
/api/ask |
streaming natural-language answer (Server-Sent Events) |
/mcp |
the MCP endpoint — so the same deployment serves browsers and MCP clients |
Search works with no keys at all. DATA_GOV_IN_API_KEY unlocks rows;
ANTHROPIC_API_KEY unlocks answers. The UI tells you which are missing.
How answers work. /api/ask runs a streaming Claude tool-use loop over the
same five tools, narrating each step ("Searching the catalog for…", "Fetching 100
rows where city=Delhi") before the answer streams in. Claude is instructed to
answer only from rows it actually fetched, to name the dataset it used, and to say
so plainly when the data doesn't answer the question rather than filling the gap
from memory.
The tool definitions the website gives Claude are read directly off the MCP
server via list_tools() — there is exactly one description and one schema per
tool in this project, so the two surfaces cannot drift apart.
Curating a dataset
Harvesting brings in every dataset automatically. Use this to improve one — attach search keywords, a worked example filter, or a corrected sector, and pin it above harvested results:
python scripts/add_dataset.py <resource_id> \
--sector Agriculture \
--keywords "wheat,crop,production" \
--example-filters '{"State":"Punjab"}'
Curated fields survive later harvests.
Notes on the upstream API
Behaviours this server handles for you:
- Filter field names are case-sensitive (
filters[State]≠filters[state]) and this is undocumented. Always use the exact fieldidfromget_dataset_info. - Some legacy datasets return CSV regardless of
format=json; the client detects this by Content-Type and parses it anyway. CSV carries no row total, sototal_recordscomes backnullrather than a misleading page count. - Last-page pagination can return an empty
recordsarray withstatus: ok;returned: 0means you're done. - Max ~100 rows per request on
/resource— page withoffset. /listsneeds no API key and pages up to 1000 records at a time. It is slow and occasionally times out, so the harvester retries every page with backoff.- The API key travels in the query string (upstream's design). Every error this package raises is passed through a redactor first, so a key can never reach a log line, a tool result, or the model's context.
Project layout
datagovin-mcp/
├── server.py # entry point (kept for existing client configs)
├── datagovin/
│ ├── config.py # .env loading, cache paths
│ ├── client.py # async data.gov.in API wrapper (quirk handling)
│ ├── catalog.py # SQLite FTS5 index: search, sectors, stats
│ ├── harvest.py # builds the index from /lists
│ ├── mcp_server.py # the five tools; stdio + Streamable HTTP
│ ├── data/seed_catalog.json # bundled fallback, works before a harvest
│ └── web/
│ ├── app.py # FastAPI: search API, /api/ask, mounts /mcp
│ ├── agent.py # streaming Claude tool-use loop
│ └── static/index.html # the UI (no build step, no CDN)
├── scripts/add_dataset.py # curate/pin one dataset
└── tests/ # 87 tests, no network required
The index lives in your platform cache directory, not in the package — it is
generated data, it is ~350 MB, and an installed package directory is often
read-only. Override with DATAGOVIN_INDEX_PATH.
Development
pip install -e ".[web,dev]"
pytest # 87 tests, all offline
License
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。